@thinkingai/ae-cli 6.1.14 → 6.1.16
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/{auth-GBMV6TEJ.js → auth-2WTQOP77.js} +3 -3
- package/dist/{auth-NDSXE54J.js → auth-77BUFLGC.js} +57 -13
- package/dist/{capability-TAMDRZYV.js → capability-72DTW5M2.js} +9 -9
- package/dist/{capability-U7TDEEEG.js → capability-PJHNI4GJ.js} +9 -9
- package/dist/{chunk-AFXA7BRK.js → chunk-4KVPKXFX.js} +5 -5
- package/dist/chunk-4SGZG4XY.js +311 -0
- package/dist/chunk-6EIJSNBD.js +8 -0
- package/dist/chunk-C4MGVGJW.js +13 -0
- package/dist/{chunk-WZRX4KOH.js → chunk-GT46FPXN.js} +4 -6
- package/dist/{chunk-5XUSIK27.js → chunk-ILIU36SU.js} +10 -7
- package/dist/{chunk-JHENBQ5B.js → chunk-LYVNONC4.js} +0 -35
- package/dist/{chunk-JUW4AJXN.js → chunk-P3FGXJTU.js} +12 -7
- package/dist/chunk-QGM4M3NI.js +37 -0
- package/dist/{chunk-753BUTNZ.js → chunk-RGKJGKT7.js} +2 -2
- package/dist/{chunk-UIHQJK5E.js → chunk-SAU3QFIQ.js} +3 -3
- package/dist/{chunk-VLWOLBGZ.js → chunk-UOUS37JQ.js} +149 -161
- package/dist/{chunk-AXDXJTPC.js → chunk-UW5UN47B.js} +19 -2
- package/dist/chunk-VPKZ7I72.js +509 -0
- package/dist/{chunk-QATA32VR.js → chunk-VR3LCBHW.js} +2 -2
- package/dist/{chunk-IBH3LDAH.js → chunk-YA6SMTXG.js} +3 -3
- package/dist/chunk-ZZUOD757.js +598 -0
- package/dist/{client-L2YDMHQ6.js → client-TKG4WBHN.js} +4 -5
- package/dist/{community-report-client-M2RW4MXD.js → community-report-client-FI4LNVYS.js} +3 -2
- package/dist/{config-OL2LWGBV.js → config-RE6CMGPK.js} +7 -7
- package/dist/data-integration-2MYMANJI.js +3383 -0
- package/dist/index.js +68 -34
- package/dist/local-data-upload-client-BWHSUQQK.js +167 -0
- package/dist/{memory-MUP7PPL7.js → memory-CHRU2F7W.js} +7 -6
- package/dist/{memory-U4O5PMXH.js → memory-YK33G4T7.js} +7 -6
- package/dist/{metadata-5MIMNIMT.js → metadata-UILXHBWF.js} +10 -9
- package/dist/{metadata-LERKDJN6.js → metadata-XXR34N5P.js} +10 -9
- package/dist/{model-JTUEO5M4.js → model-K3KLWIW6.js} +7 -3
- package/dist/model-NR3JHFSJ.js +139 -0
- package/dist/{sync-MOSFNBVR.js → sync-DAVKYVMW.js} +8 -6
- package/dist/sync-FCKOVWWS.js +10261 -0
- package/dist/{te-agent-IFKZDHZI.js → te-agent-4BKBODMF.js} +686 -71
- package/dist/te-agent-HLW4VTQK.js +3893 -0
- package/dist/{te-analysis-FCHRTNIY.js → te-analysis-O6DCO6BS.js} +36 -13
- package/dist/{te-analysis-QG7UKGDA.js → te-analysis-ZMNGOVNW.js} +36 -13
- package/dist/{te-community-HNKVTERD.js → te-community-6HPBWJUZ.js} +24 -19
- package/dist/{te-community-IWE5B7W6.js → te-community-HLC43QKH.js} +24 -20
- package/dist/{te-dataops-KXCEB4CS.js → te-dataops-EJP56W3K.js} +57 -34
- package/dist/{te-dataops-KQPYNAE3.js → te-dataops-HDRUXY4K.js} +59 -34
- package/dist/{te-engage-D6EG3NOR.js → te-engage-FGBGQ4IY.js} +9 -8
- package/dist/{te-engage-ZSMIJLUW.js → te-engage-RAK5PESW.js} +9 -8
- package/dist/{te-experiment-K5US7RMG.js → te-experiment-SO5MPDMJ.js} +9 -8
- package/dist/{te-experiment-WA7TFMEL.js → te-experiment-VZF7BT6G.js} +9 -8
- package/dist/{te-kb-OIH3T6CS.js → te-kb-APXBWBDY.js} +6 -6
- package/dist/{te-system-AZ3URMUO.js → te-system-YARIK4S5.js} +7 -4
- package/dist/te-system-Z77IKZFN.js +2213 -0
- package/dist/{te-team-GZPU6UWA.js → te-team-EFKWYKMK.js} +7 -7
- package/dist/{update-TOBFXF2V.js → update-OGPSZM5A.js} +7 -7
- package/package.json +13 -3
- package/skills/ae-agent/SKILL.md +21 -11
- package/skills/ae-agent/references/approval-effect.md +44 -0
- package/skills/ae-agent/references/approval-request.md +46 -0
- package/skills/ae-agent/references/approval-task.md +45 -0
- package/skills/ae-agent/references/approval-type.md +32 -0
- package/skills/ae-agent/references/command_index.md +19 -0
- package/skills/ae-analysis/SKILL.md +11 -7
- package/skills/ae-analysis/references/adhoc_run.md +2 -0
- package/skills/ae-analysis/references/ai_models.md +27 -0
- package/skills/ae-analysis/references/analysis_gateway_assets.md +3 -2
- package/skills/ae-analysis/references/command_index.md +4 -1
- package/skills/ae-analysis/references/dashboard_get.md +8 -2
- package/skills/ae-analysis/references/dashboard_report_data_run.md +3 -1
- package/skills/ae-analysis/references/dashboard_update.md +4 -1
- package/skills/ae-analysis/references/project_space_business_filter_upsert.md +17 -0
- package/skills/ae-analysis/references/user_tag_models.md +4 -4
- package/skills/ae-data-integration/SKILL.md +54 -0
- package/skills/ae-data-integration/references/handoff.md +43 -0
- package/skills/ae-data-integration/references/local-analysis.md +27 -0
- package/skills/ae-data-integration/references/reuse.md +41 -0
- package/skills/ae-data-integration/references/sink-upload.md +58 -0
- package/skills/ae-data-integration/references/source-inspect.md +55 -0
- package/skills/ae-data-integration/references/sync-json-upload.md +60 -0
- package/skills/ae-data-integration/references/tracking-plan.md +35 -0
- package/skills/ae-data-integration/references/transform.md +65 -0
- package/skills/ae-data-integration/references/ue-mapping.md +142 -0
- package/skills/ae-data-integration/references/ue-routing.md +35 -0
- package/skills/ae-data-integration-helper/SKILL.md +1 -1
- package/skills/ae-dataops/SKILL.md +1 -1
- package/skills/ae-dataops/references/dataops-query.md +4 -4
- package/skills/ae-generate-tracking-plan/SKILL.md +54 -3
- package/skills/ae-metadata/SKILL.md +0 -1
- package/dist/chunk-3FY3RJ26.js +0 -293
- package/dist/chunk-S5NTSDBS.js +0 -198
- package/dist/chunk-ZQKDZXDO.js +0 -317
- package/dist/cli-token-4UPER74P.js +0 -21
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# analysis dashboard update
|
|
2
2
|
|
|
3
|
-
Use when the user wants to update dashboard settings in batch
|
|
3
|
+
Use when the user wants to update dashboard settings in batch, upsert a dashboard note, or replace the dashboard-level business filter.
|
|
4
4
|
|
|
5
5
|
Do not use for moving/copying dashboards. Use `dashboard copy`, `dashboard handover`, or the relevant project-space/folder command.
|
|
6
6
|
|
|
@@ -10,6 +10,7 @@ Command:
|
|
|
10
10
|
ae-cli analysis dashboard update --project-id <project_id> --operation settings --dashboard-ids '[1001,1002]' [--zone-offset 8] [--payload '{...}']
|
|
11
11
|
ae-cli analysis dashboard update --project-id <project_id> --operation settings --dashboard-id <dashboard_id> --refresh-type 1 --dashboard-status normal --payload '{"dashboard_job_schedule":"0 0 8 * * ?","time_config_open":true,"time_config":{},"cache_config":{},"schedule_ui_config":{}}'
|
|
12
12
|
ae-cli analysis dashboard update --project-id <project_id> --operation note-upsert --dashboard-id <dashboard_id> [--note-id <note_id>] [--note-title <title>] [--description <text>]
|
|
13
|
+
ae-cli analysis dashboard update --project-id <project_id> --operation business-filter --dashboard-id <dashboard_id> --filter '{"junction_kind":"and","ta_filters":[...]}'
|
|
13
14
|
```
|
|
14
15
|
|
|
15
16
|
For `operation=settings`:
|
|
@@ -22,4 +23,6 @@ For `operation=settings`:
|
|
|
22
23
|
|
|
23
24
|
For `operation=note-upsert`, pass `dashboard_id`; omit `note_id` to create a note or pass it to update an existing note. Do not mix note fields with batch settings fields.
|
|
24
25
|
|
|
26
|
+
For `operation=business-filter`, pass one `dashboard_id` and a `filter` object in snake_case QP form. This replaces the dashboard-level business filter saved in `ta_dashboard_business_filter`; it is not a condition-filter favorite or a space-level filter. Each simple condition uses fields such as `filter_type`, `column_name`, `table_type`, `column_type`, `select_type`, `calcu_symbol`, `ftv`, and `lack_value`. A condition with `lack_value=true` is saved as a selectable field but does not restrict query data. Pass `{"junction_kind":"and","ta_filters":[]}` to clear all dashboard-level business-filter conditions.
|
|
27
|
+
|
|
25
28
|
Output is the gateway envelope. `data` contains the update result.
|
|
@@ -0,0 +1,17 @@
|
|
|
1
|
+
# analysis project-space business-filter-upsert
|
|
2
|
+
|
|
3
|
+
Use when the user wants to create, replace, or clear the space-level business filter inherited by dashboards in one project space.
|
|
4
|
+
|
|
5
|
+
Command:
|
|
6
|
+
|
|
7
|
+
```bash
|
|
8
|
+
ae-cli analysis project-space business-filter-upsert --project-id <project_id> --space-id <space_id> --filter '{"junction_kind":"and","ta_filters":[...]}'
|
|
9
|
+
```
|
|
10
|
+
|
|
11
|
+
`filter` uses snake_case QP form. Each simple condition uses fields such as `filter_type`, `column_name`, `table_type`, `column_type`, `select_type`, `calcu_symbol`, `ftv`, and `lack_value`. A condition with `lack_value=true` is stored but does not restrict query data.
|
|
12
|
+
|
|
13
|
+
The backend queries the current space filter first, then creates it when absent or updates it with the current optimistic-lock ID and timestamp. Callers do not pass `id` or `update_time`.
|
|
14
|
+
|
|
15
|
+
Space-level and dashboard-level business-filter conditions are combined with `AND` when dashboard report data is queried. Pass `{"junction_kind":"and","ta_filters":[]}` to clear the space-level conditions.
|
|
16
|
+
|
|
17
|
+
Output is the gateway envelope. `data.created` tells whether a new space filter row was created, and `data.business_filter` contains the saved filter.
|
|
@@ -24,18 +24,18 @@ Top-level `type` is exactly one of `condition`, `metric`, `first_last`, or `sql`
|
|
|
24
24
|
|
|
25
25
|
## Metric tag
|
|
26
26
|
|
|
27
|
-
Required: `event`, `aggregation`. `property` and `
|
|
27
|
+
Required: `event`, `aggregation`. `property`, `time_range`, and `filters` are optional. Filters support only event properties and user properties. A string `field` is an event property; use `{name,type:"user_property"}` for a user property.
|
|
28
28
|
|
|
29
29
|
```json
|
|
30
|
-
{"type":"metric","metric":{"event":"pay","aggregation":"sum","property":"amount","time_range":{"mode":"previous","unit":"day","value":30}}}
|
|
30
|
+
{"type":"metric","metric":{"event":"pay","aggregation":"sum","property":"amount","time_range":{"mode":"previous","unit":"day","value":30},"filters":{"relation":"and","items":[{"field":"channel","operator":"eq","values":["app"]},{"field":{"name":"country","type":"user_property"},"operator":"eq","values":["US"]}]}}}
|
|
31
31
|
```
|
|
32
32
|
|
|
33
33
|
## First/last tag
|
|
34
34
|
|
|
35
|
-
Required: `event`, `occurrence=first|last`, and exactly one value source: `calculation` or `property`. `time_range` and `filters` are optional. Supplying neither or both value sources is rejected before execution.
|
|
35
|
+
Required: `event`, `occurrence=first|last`, and exactly one value source: `calculation` or `property`. `time_range` and `filters` are optional. Filters support only event properties and user properties. A string `field` is an event property; use `{name,type:"user_property"}` for a user property. Supplying neither or both value sources is rejected before execution.
|
|
36
36
|
|
|
37
37
|
```json
|
|
38
|
-
{"type":"first_last","first_last":{"event":"login","occurrence":"last","property":"platform","time_range":{"mode":"recent","unit":"day","value":30}}}
|
|
38
|
+
{"type":"first_last","first_last":{"event":"login","occurrence":"last","property":"platform","time_range":{"mode":"recent","unit":"day","value":30},"filters":{"relation":"and","items":[{"field":{"name":"country","type":"user_property"},"operator":"eq","values":["US"]}]}}}
|
|
39
39
|
```
|
|
40
40
|
|
|
41
41
|
## SQL tag
|
|
@@ -0,0 +1,54 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: ae-data-integration
|
|
3
|
+
description: "Bring local CSV, TSV, TXT, JSON, JSONL (NDJSON), XLS, and XLSX files into AE end-to-end: identify the source's business meaning, generate and confirm a tracking plan, transform rows into UE records, and upload. Also supports privacy-preserving local analysis and handing a small file to AE Agent. Use whenever a user wants to import offline/local data into AE or analyze a file without uploading it."
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# AE Data Integration
|
|
7
|
+
|
|
8
|
+
Turn local/offline files into AE data through one fixed pipeline of four submodules: **Source → Tracking plan → Transform → Sink**. A file is never uploaded merely because it is present: its business meaning is understood, confirmed once by a human, and only then ingested. The tracking plan is generated and confirmed **before** ingestion (data governance shift-left) — see [references/tracking-plan.md](references/tracking-plan.md).
|
|
9
|
+
|
|
10
|
+
Two entrances lead here: the AE Agent dialog (attach / plus-button upload) and `ae-cli`. Two sink paths exist: RESTful API for one-time/small loads (current phase), and LogBus / DataX for recurring/high-volume loads (next phase). Source and Sink are pluggable — adding one does not change the main pipeline.
|
|
11
|
+
|
|
12
|
+
## Mandatory safety rules
|
|
13
|
+
|
|
14
|
+
- Treat file paths, receiver endpoints, APPIDs, mappings, generated artifacts, and raw rows as sensitive.
|
|
15
|
+
- Do not print source values while inspecting. Summarize types, ratios, counts, warnings, and fingerprints only.
|
|
16
|
+
- Inspect samples are bounded but still sensitive. Summarize them; never paste raw sample values into a chat summary.
|
|
17
|
+
- Do not invent account IDs, distinct IDs, event times, event names, projects, APPIDs, receivers, or timezones.
|
|
18
|
+
- `value_mapping` and `random_pool` are explicit user decisions. Never invent them.
|
|
19
|
+
- Never auto-fill a missing time for `track`/`track_*` rows. A missing time on user-profile rows may be filled with the current time only by setting `missing_time: 'now'` and only after the user explicitly confirms it.
|
|
20
|
+
- Do not read or send an AE access token or CLI token to `/sync_json`. The receiver request uses only APPID and UE data.
|
|
21
|
+
- Never execute `data-integration upload` until the user has seen the target, mapping, valid/quarantined counts, batches, and dry-run and has explicitly confirmed that upload.
|
|
22
|
+
- A blocked manifest requires a second, explicit clean-subset decision. Never add `--allow-clean-subset` implicitly.
|
|
23
|
+
- If a batch times out or loses the network, treat that batch as unknown. Stop. Ask the user to verify receiver/AE data before the user chooses `--resume-from`; never resume automatically.
|
|
24
|
+
- Local analysis stays local. AE Agent attachment is a separate, confirmed branch with a 50 MB per-file limit.
|
|
25
|
+
|
|
26
|
+
## Workflow
|
|
27
|
+
|
|
28
|
+
Walk the four submodules in order. Each submodule is its own reference; follow it and come back here for the next step.
|
|
29
|
+
|
|
30
|
+
1. **Source — business identification.** Read [references/source-inspect.md](references/source-inspect.md). Profile every file fully, infer its business meaning using business-doc / user-prompt priors, then pick a branch via [references/ue-routing.md](references/ue-routing.md).
|
|
31
|
+
2. **Reuse check.** If a `.ae-data-integration/index.json` exists and the profile is `ue_eligible`, read [references/reuse.md](references/reuse.md) and match the recommended mapping against the handoff index. A match proposes a frozen package; after one explicit confirmation, run the returned `transform.mjs` command and jump to Sink (step 5). No match → continue.
|
|
32
|
+
3. **Tracking plan.** Read [references/tracking-plan.md](references/tracking-plan.md). Generate the event/property plan from the profile and get a single explicit confirmation from the user before touching data. The plan is a separate, required deliverable from the transform mapping: a user who supplies a column→field mapping directly has **not** completed this step, so build the plan from the confirmed mapping anyway. `user_set` still requires a plan (no events; every property becomes a user property). This step runs for **every** file: a second or later file merges its new events and properties into the existing project plan (tracking-plan.md step 4) — an existing plan is never a reason to skip it.
|
|
33
|
+
4. **Transform.** Read [references/transform.md](references/transform.md). Map columns to AE system fields and properties, convert, and quarantine dirty rows per [references/ue-mapping.md](references/ue-mapping.md).
|
|
34
|
+
5. **Sink — upload.** Read [references/sink-upload.md](references/sink-upload.md). Resolve the destination, dry-run, confirm, then upload per [references/sync-json-upload.md](references/sync-json-upload.md). `receiver_accepted` is not persistence: after a ~1-minute ingestion delay, verify the data landed with ae-cli (`tracking live-data list` / `tracking ingest summary` / `tracking ingest-error list`) rather than telling the user to check the console.
|
|
35
|
+
6. **Handoff.** Read [references/handoff.md](references/handoff.md). Export the reusable package (frozen mapping + transform script + plan reference) so the next same-shape file skips the full pipeline.
|
|
36
|
+
|
|
37
|
+
## Local analysis branch
|
|
38
|
+
|
|
39
|
+
When UE prerequisites fail, the file is an aggregate/analytical table, or the user wants analysis rather than ingestion, use [references/local-analysis.md](references/local-analysis.md) instead of the ingest pipeline.
|
|
40
|
+
|
|
41
|
+
## Optional AE Agent attachment handoff
|
|
42
|
+
|
|
43
|
+
Offer this only when the user asks to continue in AE Agent. Explain that the file leaves the local machine and ask for explicit privacy confirmation.
|
|
44
|
+
|
|
45
|
+
- Reject files over 50 MB; suggest local analysis or user-controlled splitting.
|
|
46
|
+
- Read the `ae-agent` `+add-attachment` reference before calling it.
|
|
47
|
+
- Dry-run first, show file name/type/size, and wait for confirmation.
|
|
48
|
+
- Then run `ae-cli agent +add-attachment --file '<path>'`.
|
|
49
|
+
- Return the attachment result, a copyable analysis prompt, and directions to open AE Agent.
|
|
50
|
+
- Do not create or execute an Agent conversation.
|
|
51
|
+
|
|
52
|
+
## Completion response
|
|
53
|
+
|
|
54
|
+
State which submodules ran, source fingerprint and selected data set, the tracking plan status, generated artifact paths, mapping confidence, valid/quarantined counts, and upload/attachment status. Keep facts separate from recommendations and clearly state whether persistence was verified.
|
|
@@ -0,0 +1,43 @@
|
|
|
1
|
+
# Handoff (reusable package)
|
|
2
|
+
|
|
3
|
+
After a successful run, export a **handoff package** so the next file of the same shape skips the full pipeline. The package freezes the transform logic (the confirmed mapping), the runnable transform script, and the tracking-plan reference. It does **not** re-freeze the raw data or any upload secrets.
|
|
4
|
+
|
|
5
|
+
## When
|
|
6
|
+
|
|
7
|
+
Run handoff after Transform — or after Sink — once the mapping is confirmed. It is local-only and idempotent: re-running it for the same table refreshes the package in place.
|
|
8
|
+
|
|
9
|
+
## CLI
|
|
10
|
+
|
|
11
|
+
```
|
|
12
|
+
ae-cli data-integration handoff --mapping <mapping> [--plan-file <draft.json>] [--out-dir .ae-data-integration]
|
|
13
|
+
```
|
|
14
|
+
|
|
15
|
+
- `--mapping` — the confirmed `ae-local-data-mapping/v1` mapping (the frozen transform logic, including `value_mapping` and `flatten_rules`).
|
|
16
|
+
- `--plan-file` — optional tracking-plan `draft.json` to reference inside the package (`plan.json`).
|
|
17
|
+
- `--out-dir` — handoff root. Default `.ae-data-integration/` (travels with the project).
|
|
18
|
+
- `--dry-run` previews the fingerprint, files, and index path without writing.
|
|
19
|
+
|
|
20
|
+
## Package layout
|
|
21
|
+
|
|
22
|
+
```
|
|
23
|
+
.ae-data-integration/
|
|
24
|
+
index.json ← shared fingerprint index (reuse detection)
|
|
25
|
+
<fingerprint[:16]>/
|
|
26
|
+
mapping.json ← frozen mapping (re-runnable by convert)
|
|
27
|
+
transform.mjs ← node transform.mjs <new-file> [<output-dir>]
|
|
28
|
+
plan.json ← optional tracking-plan reference
|
|
29
|
+
```
|
|
30
|
+
|
|
31
|
+
The `transform.mjs` wrapper shells out to `ae-cli data-integration convert`; it takes the new file path as its first argument, never bakes the source path, and re-stamps `source.sha256` with the new file's fingerprint before converting. Reuse it only for a file of the **same shape** (same header/schema and format) — the transform logic is frozen, but the content guard is re-bound to each specific file.
|
|
32
|
+
|
|
33
|
+
## Structure fingerprint and index
|
|
34
|
+
|
|
35
|
+
Every handoff records one entry in `index.json` (`ae-data-integration-index/v1`) keyed by a **structure fingerprint** — a SHA-256 over the table shape: columns (source name + type), event model (`mode`, `event_name_field`, `record_type_field`), identity fields, and excluded columns. Business logic (`value_mapping`, transforms, `time_format`, and the fixed `default_event_name`) is excluded, so re-handing off the same table with new business rules refreshes the existing entry instead of forking a new one.
|
|
36
|
+
|
|
37
|
+
Reuse matching is its own step — see [references/reuse.md](reuse.md). It compares a new file's profile structure against this index and proposes the matching package, which the user confirms before skipping the full pipeline.
|
|
38
|
+
|
|
39
|
+
## Safety rules
|
|
40
|
+
|
|
41
|
+
- Treat the mapping, plan, generated artifacts, and the `.ae-data-integration/` directory as sensitive.
|
|
42
|
+
- Never write APPID, endpoints, tokens, or raw data values into a handoff package. The package references the plan and the mapping; uploads still require an explicit, confirmed `upload` call.
|
|
43
|
+
- Do not invent a mapping or plan. Handoff only packages what the user already confirmed.
|
|
@@ -0,0 +1,27 @@
|
|
|
1
|
+
# Local analysis
|
|
2
|
+
|
|
3
|
+
Keep the source on the local machine. Generated scripts and reports belong under `.ae-cli/data-integration/<run-id>/` with restrictive permissions.
|
|
4
|
+
Set the directory to `0700` and generated scripts/reports to `0600`.
|
|
5
|
+
|
|
6
|
+
## Default report when no question is supplied
|
|
7
|
+
|
|
8
|
+
Include only analyses supported by the fields:
|
|
9
|
+
|
|
10
|
+
- Schema, row count, missingness, parse failures, and duplicate rows/keys.
|
|
11
|
+
- Numeric and categorical distributions without dumping raw values.
|
|
12
|
+
- Time coverage, frequency, gaps, and trends when a reliable time field exists.
|
|
13
|
+
- Outliers with a declared method and threshold.
|
|
14
|
+
- Correlations or segment comparisons only when sample size and field semantics support them.
|
|
15
|
+
|
|
16
|
+
## Targeted report
|
|
17
|
+
|
|
18
|
+
Restate the user's question, identify the relevant population/time/metrics, run reproducible calculations, then report:
|
|
19
|
+
|
|
20
|
+
1. Observed facts.
|
|
21
|
+
2. Inferences and alternative explanations.
|
|
22
|
+
3. Confidence.
|
|
23
|
+
4. Data limitations and excluded rows.
|
|
24
|
+
|
|
25
|
+
Do not claim causality from correlation. Do not hide missingness or filtering. Charts must have titles, units, time ranges, and definitions for derived measures.
|
|
26
|
+
|
|
27
|
+
`analyze.mjs` must accept the source path and selected data-set as parameters or documented constants, avoid network calls, and write `analysis-report.md`. It must never embed APPID, receiver, or tokens.
|
|
@@ -0,0 +1,41 @@
|
|
|
1
|
+
# Reuse (match a historical handoff package)
|
|
2
|
+
|
|
3
|
+
Before walking the full pipeline for a new file, check whether a **handoff package** already exists for the same table shape. A match lets the file skip Source re-identification and Transform — the frozen transform is run directly, then Sink proceeds with its own confirmation. It does **not** skip the Tracking plan step unconditionally: skip it only when the matched package's `plan.json` already covers every event and property the new file will produce, including flattened object sub-properties. New properties the plan does not cover still go through the Tracking plan gate and merge into the existing project plan before Sink.
|
|
4
|
+
|
|
5
|
+
## When
|
|
6
|
+
|
|
7
|
+
Run reuse right after `inspect`, only when the profile is `ue_eligible` and a `.ae-data-integration/index.json` exists. It is read-only and makes no writes.
|
|
8
|
+
|
|
9
|
+
## CLI
|
|
10
|
+
|
|
11
|
+
```
|
|
12
|
+
ae-cli data-integration reuse --mapping <recommended_mapping> [--out-dir .ae-data-integration]
|
|
13
|
+
```
|
|
14
|
+
|
|
15
|
+
- `--mapping` — the candidate mapping, typically `inspect`'s `recommended_mapping` (the new file's structure, inferred exactly as Source would infer it).
|
|
16
|
+
- `--out-dir` — handoff root. Default `.ae-data-integration/`.
|
|
17
|
+
- `--dry-run` previews the fingerprint and match verdict without reading frozen packages.
|
|
18
|
+
|
|
19
|
+
## Matching
|
|
20
|
+
|
|
21
|
+
The command computes the same **structure fingerprint** the handoff index is keyed on — columns (source name + type), event model, identity fields, and excluded columns — and looks it up in `index.json`. Business logic (the frozen event name, `value_mapping`, transforms, `time_format`) is not part of the key, so a same-shape file with different content or a different file name still matches.
|
|
22
|
+
|
|
23
|
+
## Result
|
|
24
|
+
|
|
25
|
+
- **No match** (`matched: false`): continue the full pipeline from Tracking plan.
|
|
26
|
+
- **Match** (`matched: true`): the result carries the matched package — `mapping_file`, optional `plan_file`, the frozen `default_event_name` (so the user sees which event name will be reused), and a `run` command:
|
|
27
|
+
|
|
28
|
+
```
|
|
29
|
+
node .ae-data-integration/<fingerprint[:16]>/transform.mjs <new-input-file> [<output-dir>]
|
|
30
|
+
```
|
|
31
|
+
|
|
32
|
+
## Confirmation gate
|
|
33
|
+
|
|
34
|
+
Do **not** run the returned command on your own. Show the user the proposed package — the frozen event name, the property mapping it implies, and the fact that the confirmed business logic (event name, `value_mapping`, flatten rules) is reused unchanged — and wait for one explicit confirmation. Before confirming, diff the new file's flattened properties against the package's `plan.json`: if the package has no `plan_file`, or the new file produces properties the plan does not cover, run the Tracking plan step to add them (merge into the existing project plan) first. Only then run `transform.mjs` on the new file and continue to Sink.
|
|
35
|
+
|
|
36
|
+
## Safety rules
|
|
37
|
+
|
|
38
|
+
- Reuse only for a file of the **same shape** (same header/schema and format). The `transform.mjs` wrapper re-binds the content fingerprint to the new file, so the content guard still applies per run.
|
|
39
|
+
- A match does not authorize an upload. Sink still requires its own explicit, confirmed `upload` call.
|
|
40
|
+
- A mismatch is not an error — it means the new file's shape differs and needs the full pipeline.
|
|
41
|
+
- Never invent or edit the frozen mapping to force a match; if the shape genuinely differs, walk the pipeline again.
|
|
@@ -0,0 +1,58 @@
|
|
|
1
|
+
# Sink — resolve destination and upload
|
|
2
|
+
|
|
3
|
+
## Resolve the destination
|
|
4
|
+
|
|
5
|
+
Use the current AE CLI project commands; never guess IDs:
|
|
6
|
+
|
|
7
|
+
The logical `analysis_common +list_projects` lookup is represented by `project info list`; the
|
|
8
|
+
logical `analysis_meta +get_project_config` lookup is represented by `project info get` plus
|
|
9
|
+
`project timezone get`. Do not invent or shell-execute the legacy tool labels.
|
|
10
|
+
|
|
11
|
+
```bash
|
|
12
|
+
ae-cli project info list
|
|
13
|
+
ae-cli project info get --project-id <project-id>
|
|
14
|
+
ae-cli project timezone get --project-id <project-id>
|
|
15
|
+
ae-cli config current
|
|
16
|
+
```
|
|
17
|
+
|
|
18
|
+
Resolve and show the project ID/name, APPID, receiver `/sync_json` endpoint, and project/source timezone. Prefer receiver configuration returned by the project capability. If it is missing, replacing `-web-` with `-receiver-` in the active SaaS host creates only a candidate; label it unverified and ask the user to confirm it. Never upload to that candidate automatically.
|
|
19
|
+
|
|
20
|
+
## Upload
|
|
21
|
+
|
|
22
|
+
Read [sync_json upload](sync-json-upload.md). Run dry-run with the exact final flags:
|
|
23
|
+
|
|
24
|
+
```bash
|
|
25
|
+
ae-cli data-integration upload \
|
|
26
|
+
--ue-file '<run-dir>/valid.ue.jsonl' \
|
|
27
|
+
--manifest '<run-dir>/manifest.json' \
|
|
28
|
+
--endpoint '<https://.../sync_json>' \
|
|
29
|
+
--appid '<appid>' \
|
|
30
|
+
--batch-size 500 \
|
|
31
|
+
--dry-run
|
|
32
|
+
```
|
|
33
|
+
|
|
34
|
+
Show the masked target, project, file fingerprint, record count, quarantined count, batch count, and persistence limitation. Re-state the system-field mapping first (`#type`/mode, `#account_id`, `#distinct_id`, `#time` + source timezone, `#event_name`), then re-list the final property mapping for every file or sheet being uploaded — source column → target AE name → type, grouped into event properties (`track`) and user properties (profile modes) — never a counts-only summary. Wait for explicit confirmation. Execute the same command without `--dry-run` only after confirmation. For a blocked manifest, first show the quarantine statistics and separately ask whether the user accepts uploading only valid rows; add `--allow-clean-subset` only after a clear yes.
|
|
35
|
+
|
|
36
|
+
`status=receiver_accepted` means receiver acceptance only, not durable storage. Say that persistence remains unverified. Never report success on this status alone.
|
|
37
|
+
|
|
38
|
+
## Verify the data landed
|
|
39
|
+
|
|
40
|
+
Ingestion is asynchronous — expect about a 1-minute delay between `receiver_accepted` and the data being queryable. Instead of telling the user to check in the console, offer to verify with ae-cli (control-plane commands; require the CLI to be logged in to the AE host and a project-id):
|
|
41
|
+
|
|
42
|
+
```bash
|
|
43
|
+
# Recent received data (closest to "did it land"):
|
|
44
|
+
ae-cli tracking live-data list -p <project-id>
|
|
45
|
+
ae-cli tracking live-data list -p <project-id> --data-type error # only errored records
|
|
46
|
+
|
|
47
|
+
# Ingestion counts over a time window (better for bulk uploads):
|
|
48
|
+
ae-cli tracking ingest summary -p <project-id> --start-time '<YYYY-MM-DD HH:mm:ss>' --end-time '<YYYY-MM-DD HH:mm:ss>'
|
|
49
|
+
|
|
50
|
+
# Ingest errors for one event/property name (the silent-drop detective):
|
|
51
|
+
ae-cli tracking ingest-error list -p <project-id> --data-name <event-or-property> --start-time '<...>' --end-time '<...>'
|
|
52
|
+
```
|
|
53
|
+
|
|
54
|
+
Wait about a minute after `receiver_accepted` before the first check, then re-check once if the data has not appeared. `live-data list` only returns recent data, so for large or bulk uploads prefer `ingest summary` (counts over the upload window) or an `analysis query` on the event. Report what was verified and what was not; persistence is never established by `receiver_accepted` alone.
|
|
55
|
+
|
|
56
|
+
## Continue with the next file
|
|
57
|
+
|
|
58
|
+
After the verification confirms the data landed in AE, ask whether every file has been imported. If more files remain, batch them through one inspect and one convert — repeat `--input-file` and use a wildcard mapping plus `--type-resolutions` when conflicts were reported — which produces one manifest (and one `valid.ue.jsonl`) per file. Upload is always one manifest at a time: run a separate confirmed upload for each file's `valid.ue.jsonl`. Each upload keeps accumulating into the project's event/property metadata.
|
|
@@ -0,0 +1,55 @@
|
|
|
1
|
+
# Source — business identification
|
|
2
|
+
|
|
3
|
+
## Establish intent and file
|
|
4
|
+
|
|
5
|
+
Confirm:
|
|
6
|
+
|
|
7
|
+
1. The exact file path or paths.
|
|
8
|
+
2. The desired outcome: AE ingestion, local analysis, or help deciding.
|
|
9
|
+
3. For ingestion, the target environment/project if already known.
|
|
10
|
+
|
|
11
|
+
Accept one or more CSV, TSV, TXT, JSON, JSONL (NDJSON), XLS, or XLSX files. CSV/TSV/TXT/JSON/JSONL/XLSX may be at most 200 MB each; XLS may be at most 50 MB each.
|
|
12
|
+
|
|
13
|
+
## Inspect without exposing raw values
|
|
14
|
+
|
|
15
|
+
Run:
|
|
16
|
+
|
|
17
|
+
```bash
|
|
18
|
+
ae-cli data-integration inspect --input-file '<path>'
|
|
19
|
+
```
|
|
20
|
+
|
|
21
|
+
Repeat `--input-file` to inspect several files at once; the output then adds cross-file type conflicts and a per-file column union.
|
|
22
|
+
|
|
23
|
+
If `selection_required=true`, show only the Sheet/JSON Path candidates and ask the user to choose. Then rerun:
|
|
24
|
+
|
|
25
|
+
```bash
|
|
26
|
+
ae-cli data-integration inspect --input-file '<path>' --data-set '<candidate-id>' --source-timezone '<iana-timezone>'
|
|
27
|
+
```
|
|
28
|
+
|
|
29
|
+
Report row/column counts, field types, missing/unique/time-parse ratios, UE eligibility, mapping confidence, and warnings. Samples are bounded (up to 5 distinct, truncated) — summarize, never paste them. ID-like columns (`id`, `*_id`, `*_key`, `*_code`, `*_no`, `*_num`) stay `string` even when every value is numeric; JSON-encoded object/array values inside CSV cells are recognized as `object`/`list`, not `string`. Read [UE routing](ue-routing.md) before choosing a branch.
|
|
30
|
+
|
|
31
|
+
## Advanced input
|
|
32
|
+
|
|
33
|
+
- **Headerless files** — inspect auto-detects a missing header row on CSV/TSV and reports `no_headers: true` with a `header_detection` verdict and `auto_headers` placeholders (`col_1..col_N`); the first row is already treated as data. `--headerless` forces the same behavior without detection. Never keep the `col_1..col_N` placeholders — they carry no business meaning. For each column, read its bounded samples and inferred type and propose a meaningful name, present every proposal to the user (column position, sample summary, suggested name), and let the user confirm or rename each one; record the confirmed names in the mapping's `headers` field. When the user already knows the names, re-run inspect with `--headers 'col1,col2,...'` so the recommended mapping carries them.
|
|
34
|
+
- **TSV / TXT** — `.tsv` and `.tab` use a tab delimiter with no quoting convention; `.txt` and unknown extensions are content-sniffed into CSV, TSV, or NDJSON.
|
|
35
|
+
- **Encoding** — text files are auto-detected (UTF-8, GBK, GB2312, Big5, and others); no flag is needed.
|
|
36
|
+
- **Excel sheets** — `--merge-sheets` streams every worksheet in file order instead of a single selected sheet; otherwise ask which sheet/`--data-set` to use. Inspect also reports `header_consistency` (`all_same` or `different`) across a workbook's sheets, with `header_details` listing each sheet's header row when they differ; prefer `--merge-sheets` only when headers match.
|
|
37
|
+
- **Multi-file type conflicts** — when the same column has different inferred types across files, present each conflict and resolve with `--type-resolutions` on `convert` (see [transform](transform.md)).
|
|
38
|
+
|
|
39
|
+
## Nested flattening (NDJSON/JSON records and JSON-encoded CSV/TSV cells)
|
|
40
|
+
|
|
41
|
+
Nested data is flattened one level per user decision. This flow is agent-driven and non-interactive: never pipe pre-filled answers into any prompt, and never silently pick `Flatten` for a node.
|
|
42
|
+
|
|
43
|
+
1. **Locate the nested structure.**
|
|
44
|
+
- NDJSON/JSON: read the top-level `nested_tree` from the inspect result (record-root paths).
|
|
45
|
+
- CSV/TSV: a JSON-encoded object column carries its own tree at `columns[].nested_tree` (paths relative to that cell). Array cells (`items`-style) are `list` columns and carry no tree.
|
|
46
|
+
Object nodes list child keys, array nodes carry an `elementKind`, primitive leaves carry an inferred type and bounded samples. Summarize node kinds, never paste samples.
|
|
47
|
+
2. **Judge first, then confirm.** Read the node names and inferred types and propose, per container, whether to flatten (and to what depth), keep as JSON, or keep as a list. Present the proposal as a table and default to it; ask the user to confirm or adjust. Only ask node-by-node for the nodes you cannot judge from the names.
|
|
48
|
+
- Business-entity names with stable scalar children (`user_info`, `address`, `order`) → propose flattening to the leaves.
|
|
49
|
+
- Generic container names (`payload`, `data`, `config`, `meta`, `extra`, `attributes`) → propose keeping as JSON, unless the user wants a specific sub-field.
|
|
50
|
+
- Arrays are always kept as a list property and never split.
|
|
51
|
+
- Nesting deeper than one level with no clear intent → propose keeping as JSON or flattening only the first level.
|
|
52
|
+
3. **Record the answers in `flatten_rules` (`{ "out_column": "dot.path" }`).**
|
|
53
|
+
- NDJSON/JSON: the path is from the record root (`user_info.name`).
|
|
54
|
+
- CSV/TSV: the path is `<column>.<cell-relative path>` (`user_profile.name`), and add the source column to `exclude_columns` so the whole object is not also mapped.
|
|
55
|
+
A leaf path becomes a snake_case out column when not explicitly named (`user_info.address.geo.lat` → `user_info_address_geo_lat`; CSV prefixes the column: `user_profile.level` → `user_profile_level`). A string that looks like a number (phone, zip, ID) is kept as a string unless the user says otherwise. Containers kept whole are declared `type: 'object'`/`'list'` **with `transform: 'json'`** so conversion restores the native structure.
|
|
@@ -0,0 +1,60 @@
|
|
|
1
|
+
# `/sync_json` upload
|
|
2
|
+
|
|
3
|
+
The command sends a JSON array of receiver envelopes:
|
|
4
|
+
|
|
5
|
+
```json
|
|
6
|
+
[
|
|
7
|
+
{ "appid": "<appid>", "debug": 0, "data": { "#type": "track", "#time": "2026-08-10 10:00:00.000", "#account_id": "u-1", "#event_name": "open", "properties": { "channel": "web" } } }
|
|
8
|
+
]
|
|
9
|
+
```
|
|
10
|
+
|
|
11
|
+
The CLI assembles each envelope around the original UE JSONL text, preserving JSON number serialization from the generated file. System fields (`#type`, `#time`, `#account_id`/`#distinct_id`, `#event_name`) sit at the top level; all properties are nested under `properties`.
|
|
12
|
+
|
|
13
|
+
## Transport guarantees
|
|
14
|
+
|
|
15
|
+
- Explicit HTTP(S) endpoint ending in `/sync_json`.
|
|
16
|
+
- No AE access token, CLI token, Authorization header, or redirect following.
|
|
17
|
+
- Each request is retried 3 times with 2000ms backoff for network failures and HTTP 5xx. HTTP 4xx, non-zero receiver codes, redirects, and invalid responses are never retried.
|
|
18
|
+
- No compression unless the user enables it (see below).
|
|
19
|
+
- Sequential batches; default 500, maximum 1000 records.
|
|
20
|
+
- Fixed 30-second timeout.
|
|
21
|
+
- `code=0` means receiver accepted the batch, not durable storage.
|
|
22
|
+
|
|
23
|
+
## Optional uncertain-batch salvage and compression
|
|
24
|
+
|
|
25
|
+
Compression is opt-in and only changes transport, never the receiver contract.
|
|
26
|
+
|
|
27
|
+
- `--retry` — after the 3 transport retries, split a still-uncertain batch into single-record retries and report the offsets that keep failing, instead of stopping the whole upload. Data-level errors (HTTP 4xx, `code != 0`) are never retried. A final result still reports `persistence_verified=false`.
|
|
28
|
+
- `--compress gzip` — gzip the request body with `Content-Encoding: gzip`. The request still carries no auth header and does not follow redirects.
|
|
29
|
+
- `--debug` — mark every record `debug=1` so AE returns per-record validation details for the batch. Testing only; never use for production uploads. Default `debug=0`.
|
|
30
|
+
|
|
31
|
+
## Confirmation gate
|
|
32
|
+
|
|
33
|
+
Before execution, show:
|
|
34
|
+
|
|
35
|
+
- Resolved project ID/name.
|
|
36
|
+
- Masked receiver target and APPID.
|
|
37
|
+
- UE file SHA-256.
|
|
38
|
+
- Record count, quarantined count, resume offset, batch size/count, and bytes.
|
|
39
|
+
- Mapping mode/confidence and timezone.
|
|
40
|
+
- `persistence_verified=false`.
|
|
41
|
+
|
|
42
|
+
Wait for a clear confirmation. If the manifest is blocked, separately confirm clean-subset upload before using `--allow-clean-subset`.
|
|
43
|
+
|
|
44
|
+
## Interrupted delivery
|
|
45
|
+
|
|
46
|
+
Timeout or network failure makes the current batch `delivery_state=unknown`. The error's `meta` reports `completed_records`/`completed_batches`, `uncertain_batch_start`/`uncertain_batch_size`, and a candidate offset `resume_from_after_verification` — the zero-based record position in `valid.ue.jsonl` where the uncertain batch starts. Do not use that offset until the user verifies what receiver/AE accepted. Never retry or resume automatically.
|
|
47
|
+
|
|
48
|
+
To continue after verification, re-run the same upload with that offset; records before it are not re-sent:
|
|
49
|
+
|
|
50
|
+
```bash
|
|
51
|
+
ae-cli data-integration upload \
|
|
52
|
+
--ue-file '<run-dir>/valid.ue.jsonl' \
|
|
53
|
+
--manifest '<run-dir>/manifest.json' \
|
|
54
|
+
--endpoint '<https://.../sync_json>' \
|
|
55
|
+
--appid '<appid>' \
|
|
56
|
+
--batch-size 500 \
|
|
57
|
+
--resume-from <verified-offset>
|
|
58
|
+
```
|
|
59
|
+
|
|
60
|
+
The resumed run keeps the same confirmation gate as a normal upload.
|
|
@@ -0,0 +1,35 @@
|
|
|
1
|
+
# Tracking plan (plan before ingest)
|
|
2
|
+
|
|
3
|
+
Generate and confirm the event/property plan **before** any transform or upload. This is the data-governance shift-left rule: the plan carries business meaning into AE, so dashboards and reports are usable immediately after ingestion.
|
|
4
|
+
|
|
5
|
+
## Input
|
|
6
|
+
|
|
7
|
+
- The `inspect` profile: columns, types, samples, UE eligibility, and mapping confidence.
|
|
8
|
+
- Optional business documentation and a one-line user goal.
|
|
9
|
+
|
|
10
|
+
## Sub-steps
|
|
11
|
+
|
|
12
|
+
1. **Event-model decision** — reuse the UE routing result: single-table single-event `track`, single-table multi-event (event-name column), `user_set`, or `mixed`. The agent may propose splitting one table into several events (for example an ad table into `ad_show`/`ad_click` by `campaign_type`); that proposal must be confirmed by the user in the confirmation gate.
|
|
13
|
+
2. **Column → property draft** — identify system columns (time, `distinct_id`/`account_id`, event-name column, user-property-name column); map the remaining columns to event / user / common properties. Name events and properties in snake_case and fill **every** `display_name`, `desc`, and `event_tag` (events also carry `event_desc`). Infer all three from field names, value distribution, samples, and business-doc / prompt priors — never leave them empty: `desc`/`event_desc` state what the item means in plain language (language follows the user), and `event_tag` picks the closest category from the canonical tag list (see the `event_tag` appendix in `../../ae-generate-tracking-plan/references/business-dimension-mapping.md`). When you cannot infer a `desc` or `event_tag`, mark it pending and ask the user for it inside the confirmation gate. Infer types (`number` / `bool` / `datetime` / enum) the same way; CSV defaults to `string`. Columns that stay uncertain or conflicting are marked pending and asked only inside the confirmation gate.
|
|
14
|
+
3. **Confirmation gate (single, one pass)** — present one summary table: event list + property list (uncertain types highlighted) + field scope (default: plan fields only, with a full-import switch) + unrecognized/dirty-data handling. The user answers once with ok or edits (renames, types, add/drop columns).
|
|
15
|
+
4. **Merge with the existing plan** — fetch the project's current tracking plan; same-name property type conflicts are severe, same-name events are advisory; decide append vs replace. This runs for **every** file, not just the first: when the project already has a plan (an earlier file or run), diff this file's events and properties against it and put the additions — new events, new properties, new object sub-properties from flattening — in the confirmation gate. An existing plan is never a reason to skip this step; only when every addition is already present may you skip the merge, and even then state and confirm that fact with the user.
|
|
16
|
+
5. **Persist the plan** — draft.json → xlsx → upload with `sdk_integration_mode=none`.
|
|
17
|
+
6. **Hand off the field mapping** — the confirmed plan plus column→property mapping, `value_mapping`, and `flatten_rules` feed the Transform submodule.
|
|
18
|
+
|
|
19
|
+
## CLI
|
|
20
|
+
|
|
21
|
+
`ae-cli data-integration plan --mapping <mapping> [--event-name <name>...] [--plan-name <name>] [--out <draft.json>] [--dry-run]` converts the confirmed `ae-local-data-mapping/v1` mapping into a tracking-plan `draft.json` with `sdk_integration_mode=none` and `source_type=data`.
|
|
22
|
+
|
|
23
|
+
- `user_set` mode → no events; every mapping property becomes a user property.
|
|
24
|
+
- `track` mode → one event per `--event-name` (or the mapping `default_event_name`); every property becomes an event property.
|
|
25
|
+
- `mixed` mode → the track event(s) above, plus the same properties mirrored as user properties.
|
|
26
|
+
- `exclude_columns` are dropped from the draft.
|
|
27
|
+
- `desc` and `event_tag` flow from the mapping: each property's `desc` (falling back to its source column name) and each event's `event_meta.<name>.desc` / `event_meta.<name>.tag` (falling back to the source event name) are written into the draft. Fill them in the mapping so the plan is never empty.
|
|
28
|
+
- The per-row event-name column cannot be enumerated without a full scan, so when `default_event_name` is absent the CLI requires `--event-name` for each concrete event name.
|
|
29
|
+
- `--dry-run` previews events and properties without writing `draft.json`.
|
|
30
|
+
|
|
31
|
+
Afterward reuse `ae-cli tracking plan draft --in draft.json --out draft.xlsx` and `ae-cli tracking plan validate / upload` for xlsx generation and ingestion.
|
|
32
|
+
|
|
33
|
+
## Reuse
|
|
34
|
+
|
|
35
|
+
The data path owns plan drafting and upload through the commands above: `ae-cli data-integration plan` (draft.json) → `ae-cli tracking plan draft` (xlsx) → `validate` → `upload`; do not reimplement these steps. Plan drafting, validation, and upload semantics otherwise follow the `ae-generate-tracking-plan` skill. The data path adds a dry-run / condensed single-gate confirmation mode and a "data sample as input" path; both are in progress and should not be reimplemented here.
|
|
@@ -0,0 +1,65 @@
|
|
|
1
|
+
# Transform — column → UE mapping
|
|
2
|
+
|
|
3
|
+
Precondition: the tracking plan ([tracking-plan.md](tracking-plan.md)) has been generated and confirmed. A field mapping is not a tracking plan; if the plan is missing, return to tracking-plan.md first.
|
|
4
|
+
|
|
5
|
+
Read [UE mapping](ue-mapping.md) and apply these gates:
|
|
6
|
+
|
|
7
|
+
- At least one real account/distinct ID source exists, or an explicitly user-approved fixed placeholder (`account_id_value`/`distinct_id_value`) or `random_pool`.
|
|
8
|
+
- A real event/profile time source exists and is within the supported window.
|
|
9
|
+
- Track data has a verified event column or user-approved default event name.
|
|
10
|
+
- Low-confidence or mixed mappings are explicitly reviewed.
|
|
11
|
+
|
|
12
|
+
Confirm the system fields with the user before touching properties. These are the top-level `#` fields, and users may not understand them, so explain each in plain language before asking; never infer a decision from a column name alone.
|
|
13
|
+
|
|
14
|
+
1. **`#type` / mode.** Confirm the recommendation's `mode` (`track` / `user_set` / `mixed`). If `record_type_field` is set, confirm each distinct value normalizes to one of the eight record types. Ask whether the data reports events (→ `track`) or modifies user profiles (→ `user_set` or another profile type).
|
|
15
|
+
2. **`#account_id` and `#distinct_id` — ask together; at least one is required.** Explain both: `#account_id` identifies logged-in users (database `user_id`, phone, member ID); `#distinct_id` identifies anonymous visitors (device/cookie/visitor ID). Present every `identity_candidates` entry from the inspect result alongside the recommendation's `account_id_field`/`distinct_id_field` picks, and ask which column maps to `#account_id`, which to `#distinct_id`, and whether both apply. A column named `user_id` can be either an anonymous or a login ID — only the user knows. If a required identity column is missing, offer an explicit `account_id_value`/`distinct_id_value` placeholder or `random_pool`; never invent one.
|
|
16
|
+
3. **`#time`.** Confirm the time column and the source timezone (an IANA name from the user/project context). Leave `time_format` unset unless auto-detection failed (US/EU ambiguous dates). Also confirm the data's `#zone_offset` — the whole-hour UTC offset AE applies to those times, emitted inside `properties` (not top-level). The property's value is a number, so ask the user in numeric terms ("is the offset 8, i.e. UTC+8?") and set `zone_offset_value` to that integer; when a column carries the offset per row, set `zone_offset_field` instead. The two are mutually exclusive.
|
|
17
|
+
4. **`#event_name` (track only).** Confirm the event-name column, or a reviewed `default_event_name` when no column exists.
|
|
18
|
+
5. **`#ip` / `#uuid` (optional).** Ask only when the data has an IP- or UUID-like column; map it via `ip_field`/`uuid_field`, otherwise skip.
|
|
19
|
+
|
|
20
|
+
When an `#event_name`, `#account_id`, or `#distinct_id` column's values do not satisfy AE naming rules (pure Chinese, uppercase, spaces), do not stop — scan the distinct values, list them to the user, and ask for one AE-name replacement each; record the pairs in `value_mapping` (see [UE mapping](ue-mapping.md)). A value with no matching key keeps its original text and fails validation, so confirm every distinct value is covered or excluded. The same mechanism applies to a property column via that entry's own `value_mapping`.
|
|
21
|
+
|
|
22
|
+
Then confirm the property set with the user before saving the mapping. Present every property as a readable table — one row per property with source column, target AE name, and type — never dump the raw mapping JSON at the user. Whether the rows are event or user properties follows `mode`: `track` → **event properties**, `user_set` (or another profile mode) → **user properties**, `mixed` → the same set applies to both event and user rows.
|
|
23
|
+
|
|
24
|
+
Walk the table item by item:
|
|
25
|
+
|
|
26
|
+
1. **Auto-renamed targets.** Recommendation sanitizes column names (camelCase split, illegal chars → `_`, digit-leading → `field_`, nothing recognizable → `field_N`). Show each source → target; ask for a manual name for every `field_N` fallback and rewrite the `target`.
|
|
27
|
+
2. **Types.** Confirm each inferred type (`number`/`string`/`boolean`/`datetime`/`list`/`object`). ID-like columns stay `string` even when numeric; `object`/`list` columns carry `transform: 'json'`. A type is locked on first receipt, so review it before committing. If the user changes a type, warn how many rows would fail coercion — `convert` reports each failing row with its error code, so run it with the tentative mapping and confirm the quarantined count before finalizing.
|
|
28
|
+
3. **Exclusions.** Ask which columns to drop and record them as `exclude_columns`.
|
|
29
|
+
|
|
30
|
+
Before saving, present the complete mapping for a final sign-off on one grouped page — system fields (`#type`/mode, `#account_id`, `#distinct_id`, `#time` + source timezone, `#zone_offset`, `#event_name`, `#ip`/`#uuid`), then the property table (source column → target AE name → type), then excluded columns — and wait for an explicit yes. Never dump the raw mapping JSON at the user.
|
|
31
|
+
|
|
32
|
+
Save the reviewed `ae-local-data-mapping/v1` JSON locally. Convert into a new output directory:
|
|
33
|
+
|
|
34
|
+
```bash
|
|
35
|
+
ae-cli data-integration convert \
|
|
36
|
+
--input-file '<path>' \
|
|
37
|
+
--mapping '<mapping.json>' \
|
|
38
|
+
--output-dir '.ae-cli/data-integration/<run-id>'
|
|
39
|
+
```
|
|
40
|
+
|
|
41
|
+
For multiple files, repeat `--input-file` and pass a wildcard mapping plus `--type-resolutions` when inspect reported conflicts:
|
|
42
|
+
|
|
43
|
+
```bash
|
|
44
|
+
ae-cli data-integration convert \
|
|
45
|
+
--input-file '<a.csv>' --input-file '<b.csv>' \
|
|
46
|
+
--mapping '<wildcard-mapping.json>' \
|
|
47
|
+
--type-resolutions '<resolutions.json>' \
|
|
48
|
+
--output-dir '.ae-cli/data-integration/<run-id>'
|
|
49
|
+
```
|
|
50
|
+
|
|
51
|
+
The command never modifies the source. Inspect `manifest.json`; summarize valid and quarantined counts and the block reason. Do not expose rows from `invalid.rows.jsonl` unless the user specifically asks to inspect the local quarantine.
|
|
52
|
+
|
|
53
|
+
## Re-report only the failed rows (salvage loop)
|
|
54
|
+
|
|
55
|
+
Re-uploading the whole `valid.ue.jsonl` would re-send rows that already succeeded (track events would duplicate, `user_add` would double-count). Instead, re-process only the quarantined rows against the fixed mapping, anchored to the same source:
|
|
56
|
+
|
|
57
|
+
```bash
|
|
58
|
+
ae-cli data-integration convert \
|
|
59
|
+
--input-file '<same-source>' \
|
|
60
|
+
--mapping '<fixed-mapping.json>' \
|
|
61
|
+
--salvage-from '<run-dir>/invalid.rows.jsonl' \
|
|
62
|
+
--output-dir '.ae-cli/data-integration/<run-id>-salvage'
|
|
63
|
+
```
|
|
64
|
+
|
|
65
|
+
This emits a `valid.ue.jsonl` containing only the rows that now pass, plus a new `invalid.rows.jsonl` with whatever still fails. Fixes are rarely one-shot, so repeat: feed each round's `invalid.rows.jsonl` into the next `--salvage-from` until no rows fail or the user stops. Each round's `valid.ue.jsonl` is disjoint from earlier rounds', so upload each round independently (same confirmation and `--allow-clean-subset` gates as a normal upload). `--salvage-from` is single-file only and the source must be the same file that produced the quarantine.
|