@thinkingai/ae-cli 6.1.13 → 6.1.16

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (110) hide show
  1. package/dist/{auth-GBMV6TEJ.js → auth-2WTQOP77.js} +3 -3
  2. package/dist/{auth-NDSXE54J.js → auth-77BUFLGC.js} +57 -13
  3. package/dist/{capability-TAMDRZYV.js → capability-72DTW5M2.js} +9 -9
  4. package/dist/{capability-U7TDEEEG.js → capability-PJHNI4GJ.js} +9 -9
  5. package/dist/{chunk-AFXA7BRK.js → chunk-4KVPKXFX.js} +5 -5
  6. package/dist/chunk-4SGZG4XY.js +311 -0
  7. package/dist/chunk-6EIJSNBD.js +8 -0
  8. package/dist/chunk-C4MGVGJW.js +13 -0
  9. package/dist/{chunk-WZRX4KOH.js → chunk-GT46FPXN.js} +4 -6
  10. package/dist/{chunk-5XUSIK27.js → chunk-ILIU36SU.js} +10 -7
  11. package/dist/{chunk-JHENBQ5B.js → chunk-LYVNONC4.js} +0 -35
  12. package/dist/{chunk-JUW4AJXN.js → chunk-P3FGXJTU.js} +12 -7
  13. package/dist/chunk-QGM4M3NI.js +37 -0
  14. package/dist/{chunk-753BUTNZ.js → chunk-RGKJGKT7.js} +2 -2
  15. package/dist/{chunk-UIHQJK5E.js → chunk-SAU3QFIQ.js} +3 -3
  16. package/dist/{chunk-VLWOLBGZ.js → chunk-UOUS37JQ.js} +149 -161
  17. package/dist/{chunk-AXDXJTPC.js → chunk-UW5UN47B.js} +19 -2
  18. package/dist/chunk-VPKZ7I72.js +509 -0
  19. package/dist/{chunk-QATA32VR.js → chunk-VR3LCBHW.js} +2 -2
  20. package/dist/{chunk-IBH3LDAH.js → chunk-YA6SMTXG.js} +3 -3
  21. package/dist/chunk-ZZUOD757.js +598 -0
  22. package/dist/{client-L2YDMHQ6.js → client-TKG4WBHN.js} +4 -5
  23. package/dist/{community-report-client-M2RW4MXD.js → community-report-client-FI4LNVYS.js} +3 -2
  24. package/dist/{config-OL2LWGBV.js → config-RE6CMGPK.js} +7 -7
  25. package/dist/data-integration-2MYMANJI.js +3383 -0
  26. package/dist/index.js +68 -34
  27. package/dist/local-data-upload-client-BWHSUQQK.js +167 -0
  28. package/dist/{memory-MUP7PPL7.js → memory-CHRU2F7W.js} +7 -6
  29. package/dist/{memory-U4O5PMXH.js → memory-YK33G4T7.js} +7 -6
  30. package/dist/{metadata-UORZTW3T.js → metadata-UILXHBWF.js} +22 -18
  31. package/dist/{metadata-VVQ3AQT5.js → metadata-XXR34N5P.js} +22 -18
  32. package/dist/{model-JTUEO5M4.js → model-K3KLWIW6.js} +7 -3
  33. package/dist/model-NR3JHFSJ.js +139 -0
  34. package/dist/{sync-MOSFNBVR.js → sync-DAVKYVMW.js} +8 -6
  35. package/dist/sync-FCKOVWWS.js +10261 -0
  36. package/dist/{te-agent-IFKZDHZI.js → te-agent-4BKBODMF.js} +686 -71
  37. package/dist/te-agent-HLW4VTQK.js +3893 -0
  38. package/dist/{te-analysis-KOKHJJPE.js → te-analysis-O6DCO6BS.js} +84 -61
  39. package/dist/{te-analysis-LLOCCBJ4.js → te-analysis-ZMNGOVNW.js} +84 -61
  40. package/dist/{te-community-HNKVTERD.js → te-community-6HPBWJUZ.js} +24 -19
  41. package/dist/{te-community-IWE5B7W6.js → te-community-HLC43QKH.js} +24 -20
  42. package/dist/{te-dataops-KXCEB4CS.js → te-dataops-EJP56W3K.js} +57 -34
  43. package/dist/{te-dataops-KQPYNAE3.js → te-dataops-HDRUXY4K.js} +59 -34
  44. package/dist/{te-engage-BRLSRVQT.js → te-engage-FGBGQ4IY.js} +398 -12
  45. package/dist/{te-engage-HE4OPZIT.js → te-engage-RAK5PESW.js} +398 -12
  46. package/dist/{te-experiment-K5US7RMG.js → te-experiment-SO5MPDMJ.js} +9 -8
  47. package/dist/{te-experiment-WA7TFMEL.js → te-experiment-VZF7BT6G.js} +9 -8
  48. package/dist/{te-kb-OIH3T6CS.js → te-kb-APXBWBDY.js} +6 -6
  49. package/dist/{te-system-AZ3URMUO.js → te-system-YARIK4S5.js} +7 -4
  50. package/dist/te-system-Z77IKZFN.js +2213 -0
  51. package/dist/{te-team-GZPU6UWA.js → te-team-EFKWYKMK.js} +7 -7
  52. package/dist/{update-TOBFXF2V.js → update-OGPSZM5A.js} +7 -7
  53. package/package.json +13 -3
  54. package/skills/ae-agent/SKILL.md +21 -11
  55. package/skills/ae-agent/references/approval-effect.md +44 -0
  56. package/skills/ae-agent/references/approval-request.md +46 -0
  57. package/skills/ae-agent/references/approval-task.md +45 -0
  58. package/skills/ae-agent/references/approval-type.md +32 -0
  59. package/skills/ae-agent/references/command_index.md +19 -0
  60. package/skills/ae-analysis/SKILL.md +11 -7
  61. package/skills/ae-analysis/references/adhoc_run.md +2 -0
  62. package/skills/ae-analysis/references/ai_models.md +31 -1
  63. package/skills/ae-analysis/references/alert_list.md +2 -3
  64. package/skills/ae-analysis/references/analysis_gateway_assets.md +8 -7
  65. package/skills/ae-analysis/references/bi_panel_list.md +2 -2
  66. package/skills/ae-analysis/references/command_index.md +18 -15
  67. package/skills/ae-analysis/references/dashboard_get.md +8 -2
  68. package/skills/ae-analysis/references/dashboard_list.md +2 -2
  69. package/skills/ae-analysis/references/dashboard_report_data_run.md +3 -1
  70. package/skills/ae-analysis/references/dashboard_update.md +4 -1
  71. package/skills/ae-analysis/references/favorite_add.md +2 -2
  72. package/skills/ae-analysis/references/favorite_remove.md +2 -2
  73. package/skills/ae-analysis/references/metric_create.md +7 -8
  74. package/skills/ae-analysis/references/project_space_business_filter_upsert.md +17 -0
  75. package/skills/ae-analysis/references/project_space_list.md +2 -2
  76. package/skills/ae-analysis/references/public_link_list.md +2 -2
  77. package/skills/ae-analysis/references/report_list.md +3 -3
  78. package/skills/ae-analysis/references/report_list_export.md +2 -2
  79. package/skills/ae-analysis/references/sql_table_list.md +3 -2
  80. package/skills/ae-analysis/references/user_cluster_update.md +4 -2
  81. package/skills/ae-analysis/references/user_tag_models.md +4 -4
  82. package/skills/ae-analysis/references/user_tag_update.md +4 -2
  83. package/skills/ae-analysis/references/virtual_property_sql_rule_update.md +3 -3
  84. package/skills/ae-data-integration/SKILL.md +54 -0
  85. package/skills/ae-data-integration/references/handoff.md +43 -0
  86. package/skills/ae-data-integration/references/local-analysis.md +27 -0
  87. package/skills/ae-data-integration/references/reuse.md +41 -0
  88. package/skills/ae-data-integration/references/sink-upload.md +58 -0
  89. package/skills/ae-data-integration/references/source-inspect.md +55 -0
  90. package/skills/ae-data-integration/references/sync-json-upload.md +60 -0
  91. package/skills/ae-data-integration/references/tracking-plan.md +35 -0
  92. package/skills/ae-data-integration/references/transform.md +65 -0
  93. package/skills/ae-data-integration/references/ue-mapping.md +142 -0
  94. package/skills/ae-data-integration/references/ue-routing.md +35 -0
  95. package/skills/ae-data-integration-helper/SKILL.md +1 -1
  96. package/skills/ae-dataops/SKILL.md +1 -1
  97. package/skills/ae-dataops/references/dataops-query.md +4 -4
  98. package/skills/ae-engage/SKILL.md +10 -2
  99. package/skills/ae-engage/references/flow-metric-user.md +2 -1
  100. package/skills/ae-engage/references/flow-node-metric-user.md +2 -1
  101. package/skills/ae-engage/references/flow-node-user.md +2 -1
  102. package/skills/ae-engage/references/task-indicator-user.md +104 -0
  103. package/skills/ae-engage/references/task-user-detail-export.md +31 -0
  104. package/skills/ae-generate-tracking-plan/SKILL.md +54 -3
  105. package/skills/ae-metadata/SKILL.md +0 -1
  106. package/skills/ae-metadata/references/metadata_property_dimension_table_bind_existing.md +1 -1
  107. package/dist/chunk-3FY3RJ26.js +0 -293
  108. package/dist/chunk-S5NTSDBS.js +0 -198
  109. package/dist/chunk-ZQKDZXDO.js +0 -317
  110. package/dist/cli-token-4UPER74P.js +0 -21
@@ -0,0 +1,142 @@
1
+ # UE mapping contract
2
+
3
+ The mapping version is `ae-local-data-mapping/v1`.
4
+
5
+ Required structure:
6
+
7
+ ```json
8
+ {
9
+ "version": "ae-local-data-mapping/v1",
10
+ "source": {
11
+ "sha256": "<source-sha256>",
12
+ "format": "csv",
13
+ "data_set": "$"
14
+ },
15
+ "mode": "track",
16
+ "confidence": "high",
17
+ "account_id_field": "user_id",
18
+ "time": {
19
+ "field": "event_time",
20
+ "format": "auto",
21
+ "source_timezone": "Asia/Shanghai"
22
+ },
23
+ "event_name_field": "event_name",
24
+ "properties": [
25
+ { "source": "amount", "target": "amount", "type": "number" }
26
+ ]
27
+ }
28
+ ```
29
+
30
+ Use `distinct_id_field` when no account field exists. `mixed` requires `record_type_field`. Track requires `event_name_field` or a reviewed `default_event_name`. `source.format` is one of `csv`, `tsv`, `json`, `jsonl`, `xls`, `xlsx`. For a multi-file template mapping, `source.sha256` may be the wildcard `*`.
31
+
32
+ ## Record types
33
+
34
+ Every record carries a `#type`. `mode: 'mixed'` resolves the type per row via `record_type_field`; otherwise the mode fixes it (`track` or `user_set`). The `record_type_field` values are normalized case- and underscore-insensitively, with `event`→`track` and `user`/`userset`→`user_set` aliases.
35
+
36
+ | `#type` | Meaning | Requires `#event_name` |
37
+ | --- | --- | :---: |
38
+ | `track` | Report an event into the event table | Yes |
39
+ | `user_set` | Overwrite user properties (create if missing) | No |
40
+ | `user_setOnce` | Initialize a property only when empty | No |
41
+ | `user_add` | Increment numeric user properties | No |
42
+ | `user_unset` | Clear user property values | No |
43
+ | `user_del` | Delete the user from the user table | No |
44
+ | `user_append` | Append elements to list properties | No |
45
+ | `user_uniq_append` | Append elements with deduplication | No |
46
+
47
+ ## System fields
48
+
49
+ Top-level `#` fields map from named source columns. The user confirms each mapping; never infer it from a column name alone — a `user_id` column can be either an anonymous or a login ID, and only the user knows which.
50
+
51
+ | Field | Plain-language meaning | Required |
52
+ | --- | --- | :---: |
53
+ | `#distinct_id` | Anonymous visitor ID (device/cookie/visitor) — identifies unauthenticated traffic | At least one of `#distinct_id` / `#account_id` |
54
+ | `#account_id` | Login/account ID (database `user_id`, phone, member ID) — identifies authenticated users | At least one of `#distinct_id` / `#account_id` |
55
+ | `#time` | Event/profile occurrence time; AE bins data by it | Yes |
56
+ | `#event_name` | What the user did (`purchase`, `login`) | Only for `track` |
57
+ | `#ip` | Client IP; AE resolves geo from it | No |
58
+ | `#uuid` | Short-window deduplication ID | No |
59
+
60
+ `#zone_offset` is a preset property that tells AE the data's UTC offset (whole hours, -12..14) so `#time` is interpreted correctly. Unlike the fields above it lives **inside `properties`**, not at the top level. Provide it via `zone_offset_value` (a whole-hour integer such as `8` for UTC+8; an IANA name is also accepted and resolved to its offset at conversion time — sub-hour zones round to the nearest whole hour and DST zones reflect the offset then in effect, so historical data crossing a DST boundary should use `zone_offset_field`) or `zone_offset_field` (a source column carrying the offset per row); the two are mutually exclusive. Rows whose `zone_offset_field` value is missing or not an integer in -12..14 are quarantined.
61
+
62
+ Mapping keys: `account_id_field`, `distinct_id_field`, `time.field`, `event_name_field` (or a reviewed `default_event_name`), `ip_field`, `uuid_field`. Inspect surfaces every identity-shaped column in `identity_candidates` (name, `account`/`distinct` kind, unique and missing ratios) so the agent can present all candidates for confirmation. When the required identity column is absent, use an explicit `account_id_value`/`distinct_id_value` placeholder or `random_pool` — a user decision, never invented.
63
+
64
+ ## Output record shape
65
+
66
+ System fields sit at the top level; every mapped property is nested under `properties`.
67
+
68
+ ```json
69
+ { "#type": "track", "#time": "2026-08-10 10:00:00.000", "#account_id": "u-1", "#event_name": "purchase", "properties": { "amount": 99.9, "channel": "app" } }
70
+ ```
71
+
72
+ ```json
73
+ { "#type": "user_set", "#time": "2026-08-10 10:00:00.000", "#distinct_id": "d-1", "properties": { "user_level": 5, "user_tag": "vip" } }
74
+ ```
75
+
76
+ `track` requires `#event_name`; user-profile types (`user_set`, `user_add`, …) do not.
77
+
78
+ ## Optional overlay fields
79
+
80
+ All fields below are optional and are explicit user decisions — never invent them.
81
+
82
+ | Field | Shape | Effect |
83
+ | --- | --- | --- |
84
+ | `time_format` | `string` (strptime pattern) | Overrides auto-detection for `time.field`; used for ambiguous US/EU dates |
85
+ | `value_mapping` | `{ account_id? / distinct_id? / event_name? / record_type?: {original: replacement} }` | Exact-key replacement per system field (business-data keys map to AE-name values). A value with no matching key keeps its original text and fails validation, so confirm every distinct value is covered or excluded |
86
+ | `account_id_value` / `distinct_id_value` | `string` | Fixed placeholder identity (≤128 chars). Applies to every row when the corresponding `*_field` is absent; when the field exists, fills only the rows whose column is empty |
87
+ | `random_pool` | `{ account_ids?: string[], distinct_ids?: string[] }` | Synthesizes a random identity when the source field is absent |
88
+ | `exclude_columns` | `string[]` | Source columns skipped when building properties |
89
+ | `flatten_rules` | `{ outColumn: 'dot.path' }` | Nested flatten map. NDJSON/JSON paths are from the record root (`user_info.name`); CSV/TSV paths are `<column>.<cell-relative path>` into a JSON-encoded object cell (`user_profile.name`) — add the source column to `exclude_columns` when flattening it |
90
+ | `headers` | `string[]` | User-confirmed column names for a headerless file; presence means the first row is data. Never use inspect's `col_1..col_N` placeholders — infer names from each column's values, confirm them with the user, then write them here |
91
+ | `missing_time` | `'now'` | Fill a missing/empty `#time` with the current time, for user-profile rows only (explicit user decision; track rows are never filled) |
92
+ | `ip_field` | `string` | Source column emitted as the top-level `#ip` system field (client IP; AE resolves geo) |
93
+ | `uuid_field` | `string` | Source column emitted as the top-level `#uuid` system field (short-window deduplication ID) |
94
+ | `zone_offset_value` | `number` (integer -12..14) or IANA `string` | Emits the fixed `#zone_offset` preset property inside `properties`. An IANA name resolves to its integer UTC offset |
95
+ | `zone_offset_field` | `string` | Source column whose per-row integer value (-12..14) is emitted as `#zone_offset`; missing/non-integer rows are quarantined. Mutually exclusive with `zone_offset_value` |
96
+ | `event_meta` | `{ <event-name>: { desc?: string, tag?: string } }` | Per-event business description and `event_tag` for the tracking plan, keyed by AE event name. Inferred from the data and user context; the user supplies anything not inferable — never leave them empty |
97
+
98
+ Each `properties` entry may also carry `value_mapping` (per-property exact-key replacement), `transform` (one of `stringify`, `number`, `boolean`, `json`), `time_format` (only meaningful for `type: 'datetime'`), and `desc` (business description for the tracking plan; inferred, or user-provided when not inferable). Container columns (`object`/`list`) whose values arrive as JSON text — JSON-encoded CSV cells, or flattened NDJSON leaves — must set `transform: 'json'` so conversion parses the text back into a native object/array; it is a safe no-op when the value already arrived native.
99
+
100
+ ## Time formats
101
+
102
+ Values with an explicit offset or `Z` are parsed by the JavaScript `Date` constructor (authoritative). Other values are matched against 21 formats; a `time_format` field overrides the match.
103
+
104
+ | Category | Example | Auto-detected |
105
+ | --- | --- | :---: |
106
+ | Standard AE | `2024-01-15 10:30:00.123` / `2024-01-15 10:30:00` | Yes |
107
+ | ISO 8601 | `2024-01-15T10:30:00` / with `.SSS` / with offset | Yes |
108
+ | Unix epoch | `1705314600` (seconds) / `1705314600000` (milliseconds) | Yes |
109
+ | Slash-separated | `2024/01/15 10:30:00` / `2024/01/15` | Yes |
110
+ | Dot-separated | `2024.01.15 10:30:00` / `2024.01.15` | Yes |
111
+ | Compact digits | `20240115103000` / `202401151030` / `20240115` | Yes |
112
+ | English month names | `15 Jan 2024 10:30:00` / `January 15 2024 10:30:00` / `Jan 15 2024 10:30:00` | Yes |
113
+ | Chinese date | `2024年1月15日` / `2024年1月15日 10:30:00` / `2024年1月15日 10时30分00秒` | Yes |
114
+ | US format | `01/15/2024 10:30:00` (MM/DD/YYYY) | No — use `time_format` |
115
+ | EU format | `15/01/2024 10:30:00` (DD/MM/YYYY) | No — use `time_format` |
116
+
117
+ ## Data rules
118
+
119
+ - **Ingestion time window.** The receiver accepts event/profile times from 3 years before to 3 days after the server time; the CLI enforces this range. Client-side reporting has a tighter window (10 days before to 3 days after); historical data beyond the window needs AE support to extend it.
120
+ - **Property type locking.** A property's type is locked on first receipt, and properties sharing a name across events are one property — later reports must use the same type. Values whose type mismatches the locked type are dropped silently, so review an inferred type before committing it.
121
+ - **Quote stripping.** Paired single or double quotes around a cell value are stripped (`'user_001'` → `user_001`); a single-sided quote is kept as data (SQL exports often quote every value).
122
+ - **Column name sanitization.** Recommended mapping auto-names properties: accents are stripped, camelCase is split to snake_case, everything is lowercased, and illegal characters become `_`; a digit-leading name is prefixed `field_`, and a name with nothing recognizable falls back to `field_N`. Review these targets — especially the `field_N` fallbacks — and rename before converting.
123
+ - **Object sub-property keys.** Object and list-of-object keys follow the same naming rules as property names (lowercase snake_case, letter-leading, at most 50 chars, no `#`). A key that does not — for example a Chinese key inside a JSON cell such as `{"等级":"金牌"}` — quarantines the whole row, so confirm object keys before converting.
124
+
125
+ ## Review checklist
126
+
127
+ - Source SHA-256 and data-set ID match the inspection result (or the mapping uses the `*` wildcard for multi-file).
128
+ - Identity values remain strings and are at most 128 characters.
129
+ - Source timezone is an IANA name derived from user/project context.
130
+ - `#zone_offset`, when set, is a whole-hour integer in -12..14 (or an IANA name/column that resolves to one) and is emitted inside `properties`, never at the top level.
131
+ - Event/property names are lowercase snake_case, begin with a letter, and are at most 50 characters.
132
+ - Target property names are unique and do not collide with UE system fields.
133
+ - Types are one of `string`, `number`, `boolean`, `datetime`, `list`, or `object`.
134
+ - Text is at most 2 KB; numbers stay within -9E15..9E15.
135
+ - Lists contain at most 500 strings (255 bytes each) or 500 objects.
136
+ - Objects contain at most 100 legal sub-properties; nested values follow the same type limits.
137
+ - A conversion rule does not hide a real type conflict.
138
+ - Event/profile times fall within the receiver window: previous 3 years through next 3 days.
139
+ - Do not fabricate UUIDs, identities, times, or events; `#ip`/`#uuid` map from named source columns only (`ip_field`/`uuid_field`).
140
+ - `value_mapping`, `random_pool`, and fixed `account_id_value`/`distinct_id_value` came from an explicit user decision and match the actual distinct values/columns.
141
+
142
+ `user_set` output for the same identity is ordered by time so receiver application order is deterministic. Conversion applies whole-row quarantine: any error on a row — identity, time, event, record type, or a single property (type coercion, size/limit) — drops the entire row. The row is written to `invalid.rows.jsonl` with its error codes, counted in `manifest.output.invalid_records`, and the manifest is blocked until reviewed. The failed rows are re-reportable without re-sending valid rows: fix the mapping, then run `convert --input-file <same-source> --mapping <fixed-mapping> --salvage-from <invalid.rows.jsonl>`. The salvage run re-processes only the listed row numbers against the same source, emits a `valid.ue.jsonl` containing only the newly fixed rows, and writes a new `invalid.rows.jsonl` with whatever still fails — so the loop repeats (feeding each round's `invalid.rows.jsonl` into the next `--salvage-from`) until no rows fail or the user stops.
@@ -0,0 +1,35 @@
1
+ # UE routing
2
+
3
+ Use this reference after `data-integration inspect`.
4
+
5
+ ## Route to UE ingestion
6
+
7
+ All of these must hold:
8
+
9
+ - A real `#account_id` or `#distinct_id` source is present.
10
+ - A real time column is parseable; no synthetic time is permitted.
11
+ - The records are event facts or user-property snapshots rather than aggregates.
12
+ - Field-level records can be represented without losing essential meaning.
13
+
14
+ Classification order:
15
+
16
+ 1. A legal explicit `#type` wins (all eight record types — `track`, `user_set`, `user_setOnce`, `user_add`, `user_unset`, `user_del`, `user_append`, `user_uniq_append`).
17
+ 2. An event/action field implies `track`.
18
+ 3. Repeated users across a time series imply `track`; without an event field, propose a normalized file/Sheet name and require review.
19
+ 4. One row per user with snapshot-like fields implies `user_set`.
20
+ 5. Rows that mix track and user-profile facts in one file use `mixed` with a `record_type_field`; require explicit review.
21
+ 6. Low-confidence output is a proposal, never silent approval.
22
+
23
+ Aggregated metrics, pivot tables, cross-tabs, model outputs, free-form documents, and records without real identity/time should normally use local analysis.
24
+
25
+ ## Route to local analysis
26
+
27
+ Choose local analysis when:
28
+
29
+ - The user wants insights, not project ingestion.
30
+ - UE identity or time prerequisites are missing.
31
+ - Each row is an aggregate rather than a user/event record.
32
+ - Conversion would invent semantics or discard important structure.
33
+ - The user declines an uncertain mapping or destination.
34
+
35
+ Explain the reason briefly. Do not frame local analysis as an error.
@@ -62,7 +62,7 @@ Trigger when a user asks about:
62
62
  - **Configuration**: Debug mode / data validation / retry strategy
63
63
  - **Troubleshooting**: Upload failures / missing data / format errors
64
64
 
65
- **Do NOT trigger**: Inline Q&A within `ae-generate-tracking-plan` or `ae-generate-tracking-code` sessions (those skills have their own workflows).
65
+ **Do NOT trigger**: Inline Q&A within `ae-generate-tracking-plan` or `ae-generate-tracking-code` sessions (those skills have their own workflows). Offline local-file import (CSV / Excel / JSONL → AE) is owned by the `ae-data-integration` skill — route those requests there instead of answering with LogBus / RESTful guidance.
66
66
 
67
67
  ## Workflow
68
68
 
@@ -173,7 +173,7 @@ Detailed metadata browsing, SQL query, async download, and cancel workflows live
173
173
  Key constraints:
174
174
  - IDE is query-only; create/modify/delete tables with `dataops_datatable`.
175
175
  - Prefer `dataops_datatable +dict_search_tables` for table discovery unless raw engine metadata or schema browsing is required.
176
- - SQL query submission creates a download task; rows are not returned inline.
176
+ - Submit exactly one read-only SQL query. It creates a platform-bounded download task; rows are not returned inline and the result is not an unlimited or full export.
177
177
 
178
178
  ---
179
179
 
@@ -61,7 +61,7 @@ ae-cli dataops_ide +search_tables --spaceCode "${spaceCode}" --searchKey "user"
61
61
 
62
62
  ## Workflow C: Execute SQL Query (Download-Centered Async Flow)
63
63
 
64
- Use this flow for SQL query execution. Result rows are not returned through MCP/CLI; submit creates a Gaia download-center task directly.
64
+ Use this flow for exactly one read-only SQL query. Result rows are not returned through MCP/CLI; submit creates a Gaia download-center task directly. The query semantics are preserved, but the result remains platform-bounded; when present, `downloadRowLimit` reports that cap. This is not an unlimited or full export.
65
65
 
66
66
  ```bash
67
67
  # Step 1: Submit SQL and create a download task. Defaults: repoCode=te_etl, engineType=TASK_ENGINE_TRINO.
@@ -72,7 +72,7 @@ ae-cli dataops_ide +submit_sql_query --spaceCode "${spaceCode}" --repoCode "te_e
72
72
  # Step 2: Poll the download task status by spaceCode/downloadTaskId. Rows are not returned through MCP/CLI.
73
73
  ae-cli dataops_ide +get_sql_query_status --spaceCode "${spaceCode}" --downloadTaskId ${downloadTaskId}
74
74
 
75
- # Step 3: CLI-only local save after downloadStatus=SUCCESS. Use a .zip suffix.
75
+ # Step 3: CLI-only streaming save after downloadStatus=SUCCESS. The target is replaced only after the stream completes.
76
76
  ae-cli dataops_ide +get_sql_query_status --spaceCode "${spaceCode}" --downloadTaskId ${downloadTaskId} --downloadTo "./result.zip"
77
77
 
78
78
  # (Optional) Cancel the download task.
@@ -104,8 +104,8 @@ ae-cli dataops_ide +cancel_sql_query --spaceCode "${spaceCode}" --downloadTaskId
104
104
  - **Table search**: `+search_tables` requires `--spaceCode` and `--searchKey`. `--connType`, `--repoCode`, and `--size` are optional and default to `SPACE`, `te_etl`, and `20`. It returns `items`, `searchKey`, `size`, `totalCount`, `tableCount`, `viewCount`, `returnedCount`, `hasMore`, and `nextAction`.
105
105
  - **Table detail**: `+ide_get_table_detail` requires `--spaceCode`, `--catalog`, `--schema`, and `--tableName`. `--connType`, `--repoCode`, `--engineType`, `--entityType`, and `--includeDdl` are optional and default to `SPACE`, `te_etl`, `TASK_ENGINE_TRINO`, auto-detect, and `false`. It returns identity, storage metadata, columns, partitions, partition keys, optional layout fields, and `tableDdl` only when requested.
106
106
  - **Schema info**: `+get_schema_info` requires `--spaceCode`, `--catalog`, and `--schema`. `--connType` and `--repoCode` are optional and default to `SPACE` and `te_etl`. It returns only `schema`, `tableNum`, and `viewNum`.
107
- - **SQL submit**: `+submit_sql_query` requires `--spaceCode` and `--sql`. `--repoCode` and `--engineType` are optional and default to `te_etl` and `TASK_ENGINE_TRINO`. On success it returns `requestId`, `spaceCode`, `repoCode`, `downloadTaskId`, `downloadStatus`, `downloadApi`, `downloadParams`, and `nextAction`; rows are never returned.
108
- - **SQL status**: `+get_sql_query_status` requires `--spaceCode` and `--downloadTaskId`. `--requestId` is optional trace-only. `--downloadTo` is CLI-only; after `downloadStatus=SUCCESS` it saves the result zip and adds `localFile`. It returns status/progress metadata, `nextAction`, `downloadApi`, and `downloadParams`; rows are never returned.
107
+ - **SQL submit**: `+submit_sql_query` requires `--spaceCode` and exactly one read-only query in `--sql`. `--repoCode` and `--engineType` are optional and default to `te_etl` and `TASK_ENGINE_TRINO`. On success it returns task metadata and, when exposed by Gaia, `downloadRowLimit`; rows are never returned and the result remains platform-bounded.
108
+ - **SQL status**: `+get_sql_query_status` requires `--spaceCode` and `--downloadTaskId`. `--requestId` is optional trace-only. `--downloadTo` is CLI-only; after `downloadStatus=SUCCESS` it streams the result zip to a temporary sibling file, publishes it only after completion, and adds `localFile`. It returns status/progress metadata, `nextAction`, `downloadApi`, and `downloadParams`; rows are never returned.
109
109
  - **SQL cancel**: `+cancel_sql_query` requires `--spaceCode` and `--downloadTaskId`. `--requestId` is optional trace-only. It returns cancellation request metadata such as `downloadCancelStatus`.
110
110
  - **engineType**: `TASK_ENGINE_TRINO` (default, interactive queries) | `TASK_ENGINE_STARROCKS` (high-concurrency analytics)
111
111
  - **entityType**: optional `TABLE` or `VIEW` hint. Omit it unless the target type must be forced.
@@ -55,7 +55,7 @@ When the user mentions a product term below (including common Chinese UI labels)
55
55
  | **Push channel** | Project-level message push channels (Webhook, FCM, APNS, etc.) | `engage-setting` | `references/channel-list.md` | `channel-detail.md`, `add-channel.md` (**Webhook vs Client differ**: `url` = HTTP vs scene key; custom params `user:` vs `user:`/`client:`), `update-channel-status.md`, `delete-channel.md`, `channel-update-config.md`, `channel-test-send.md`, `channel_touch_limits_list.md` |
56
56
  | **Config channel** | Config-center Webhook/client config channels (not the same as push channels) | `engage-scene` | `references/scene-config-channel.md` | `channel-mgmt.md` (create/enable-disable/copy/delete workflows). User params in `config.customsParamList` require `columnName` with `user:` prefix (e.g. `user:#account_id`); preflight names with ae-analysis `analysis-meta property list/get`. |
57
57
  | **Operation strategy** | Ops/delivery strategies under a config item | `engage-scene` | `references/scene-strategy.md` | Custom audience: [`scene-strategy-audience.md`](references/scene-strategy-audience.md) — semantic `definitionRequest` (Analysis condition shape); do not pass `targetClusterQp`/`qp`; preflight props (stop + list if missing); template: `scene-template.md` |
58
- | **Operation task** | Hermes push/engagement tasks (list, save, lifecycle, reports) | `engage-task` | `references/task-list.md` | `task-detail.md` (get), `save-task.md`, `build-task-save-guide.md`, `task-stats.md`, `task-delete.md`, `push-record-query.md`, `task-data-overview.md`, `task-data-detail.md`, `task-metric-detail.md`, `task-experiment-report.md` |
58
+ | **Operation task** | Hermes push/engagement tasks (list, save, lifecycle, reports) | `engage-task` | `references/task-list.md` | `task-detail.md` (get), `save-task.md`, `build-task-save-guide.md`, `task-stats.md`, `task-delete.md`, `push-record-query.md`, `task-user-detail-export.md`, `task-indicator-user.md`, `task-data-overview.md`, `task-data-detail.md`, `task-metric-detail.md`, `task-experiment-report.md` |
59
59
  | **Operation activity** | Campaign activity management and delivery trends by activity, topic, or standalone task | `engage-activity` | `references/activity-activity.md` | `activity-data-detail.md`, `activity-topic.md`, `activity-task.md`, `activity-approval.md` |
60
60
  | **Template** | Strategy templates under a config item | `engage-scene` | `references/scene-template.md` | `scene-config-param.md` (template fields reference `paramId`); enable via `template update` then `template update-status` before strategy create |
61
61
 
@@ -178,11 +178,16 @@ ae-cli engage-task task submit-approval --project-id 1 --task-id task_123
178
178
  # Query task reports through the Hermes inline task-data capabilities
179
179
  ae-cli engage-task effect query --project-id 1 --task-id task_123 --start-time 2026-04-01 --end-time 2026-04-07 --metric-id-list '["metric_1"]'
180
180
  ae-cli engage-task data-detail query --project-id 1 --task-id task_123 --detail-type time --start-time 2026-04-01 --end-time 2026-04-07
181
+ ae-cli engage-task indicator-user sql --project-id 1 --task-id task_123 --indicator main --start-time 2026-04-01 --end-time 2026-04-07
182
+ ae-cli engage-task indicator-user run --project-id 1 --task-id task_123 --indicator secondary --secondary-index 1 --start-time 2026-04-01 --end-time 2026-04-07 --limit 100
183
+ ae-cli engage-task indicator-user export --project-id 1 --task-id task_123 --indicator metric --metric-id metric_1 --source metric --start-time 2026-04-01 --end-time 2026-04-07 --artifact-format csv
181
184
 
182
185
  ```
183
186
 
184
187
  For L3 task reports, read `references/task-data-overview.md`, `references/task-data-detail.md`,
185
188
  `references/task-metric-detail.md`, or `references/task-experiment-report.md` before invocation.
189
+ Before using `engage-task indicator-user`, read `references/task-indicator-user.md`; its grouping,
190
+ indicator, summary/detail, metric, experiment, and timezone flags have conditional compatibility rules.
186
191
 
187
192
  ### 3. config
188
193
 
@@ -214,6 +219,7 @@ ae-cli engage-flow version list --project-id 1 --flow-id flow_id_123
214
219
  ae-cli engage-flow flow update-remark --project-id 1 --flow-uuid flow_uuid_123 --flow-version-desc "Second version"
215
220
  ae-cli engage-task operation-log query --project-id 1 --task-id task_id_123
216
221
  ae-cli engage-task push-record query --project-id 1 --task-id task_id_123 --page-num 1 --page-size 20
222
+ ae-cli engage-task user-detail export --project-id 1 --task-id task_id_123 --task-instance-id instance_123 --user-status fail --artifact-format csv
217
223
  ae-cli engage-task segment-list query --project-id 1 --task-id task_id_123
218
224
  ae-cli engage-task group list --project-id 1
219
225
  ae-cli engage-task metric list --project-id 1 --task-id task_id_123
@@ -502,6 +508,8 @@ More detailed single-command guidance is available in the business-oriented `ref
502
508
  - `references/version-list.md` (`engage-flow.version.list`)
503
509
  - `references/flow-update-remark.md` (`ae-cli engage-flow flow update-remark`; capability `engage-flow.version.update-remark`)
504
510
  - `references/push-record-query.md` (`engage-task.push-record.query`)
511
+ - `references/task-user-detail-export.md` (`engage-task user-detail export`; capability `engage-task.user-detail.export`)
512
+ - `references/task-indicator-user.md` (`engage-task indicator-user {sql,run,export}`; capabilities `engage-task.indicator-user.{sql,run,export}`)
505
513
  - `references/segment-list-query.md` (`engage-task.segment-list.query`)
506
514
  - `references/group-list.md` (`engage-task.group.list`)
507
515
  - `references/task-delete.md` (`engage-task.task.delete`)
@@ -520,7 +528,7 @@ This split documentation structure is easier to extend later, because commands w
520
528
 
521
529
  ### task
522
530
 
523
- `operation-log query` / `push-record query` / `segment-list *` / `ops *` / `metric *` / `race release` / `channel-ref stats` / `group *` / `task delete` / `task modify-group` / `task submit-approval` / `task get` / `task list` / `task stats` / `task build-save-guide` / `task save` / `task manage` / `effect query` / `data-detail query` (via `engage-task`), plus L3 capabilities `engage-task.task-data.{overview,detail,metric-detail,experiment-report}`
531
+ `operation-log query` / `push-record query` / `user-detail export` / `indicator-user sql` / `indicator-user run` / `indicator-user export` / `segment-list *` / `ops *` / `metric *` / `race release` / `channel-ref stats` / `group *` / `task delete` / `task modify-group` / `task submit-approval` / `task get` / `task list` / `task stats` / `task build-save-guide` / `task save` / `task manage` / `effect query` / `data-detail query` (via `engage-task`), plus L3 capabilities `engage-task.task-data.{overview,detail,metric-detail,experiment-report}`
524
532
 
525
533
  ### query
526
534
 
@@ -109,7 +109,8 @@ ae-cli engage-query query cancel --run-id <run_id>
109
109
  ## Output shape
110
110
 
111
111
  - `run` returns `rows`, `total`, and `column_meta`.
112
- - `export` writes one row per matched user. The primary column is usually `#user_id`.
112
+ - `export` writes one row per matched user with `#user_id`, `#account_id`, and `#distinct_id`.
113
+ - For cross-user data sources, `#account_id` and `#distinct_id` are empty strings.
113
114
 
114
115
  ## Common mistakes
115
116
 
@@ -113,7 +113,8 @@ ae-cli engage-query query cancel --run-id <run_id>
113
113
  ## Output shape
114
114
 
115
115
  - `run` returns `rows`, `total`, and `column_meta`.
116
- - `export` writes one row per matched user. The primary column is usually `#user_id`.
116
+ - `export` writes one row per matched user with `#user_id`, `#account_id`, and `#distinct_id`.
117
+ - For cross-user data sources, `#account_id` and `#distinct_id` are empty strings.
117
118
 
118
119
  ## Common mistakes
119
120
 
@@ -113,7 +113,8 @@ ae-cli engage-query query cancel --run-id <run_id>
113
113
  ## Output shape
114
114
 
115
115
  - `run` returns `rows`, `total`, and `column_meta`.
116
- - `export` writes one row per matched user. The primary column is usually `#user_id`.
116
+ - `export` writes one row per matched user with `#user_id`, `#account_id`, and `#distinct_id`.
117
+ - For cross-user data sources, `#account_id` and `#distinct_id` are empty strings.
117
118
 
118
119
  ## Common mistakes
119
120
 
@@ -0,0 +1,104 @@
1
+ # engage-task indicator-user
2
+
3
+ Build SQL, query inline users, or export the users behind an engagement-task report indicator.
4
+
5
+ > Capability ids: `engage-task.indicator-user.sql`, `engage-task.indicator-user.run`, and `engage-task.indicator-user.export` · Domain: `engage` · Risk: read-only.
6
+
7
+ ## Commands
8
+
9
+ ```bash
10
+ # Build and validate SQL without executing it
11
+ ae-cli engage-task indicator-user sql \
12
+ --project-id 1 \
13
+ --task-id task_123 \
14
+ --indicator main \
15
+ --start-time 2026-04-01 \
16
+ --end-time 2026-04-07
17
+
18
+ # Query a bounded inline user list
19
+ ae-cli engage-task indicator-user run \
20
+ --project-id 1 \
21
+ --task-id task_123 \
22
+ --indicator secondary \
23
+ --secondary-index 2 \
24
+ --start-time 2026-04-01 \
25
+ --end-time 2026-04-07 \
26
+ --limit 100
27
+
28
+ # Export a custom metric's users
29
+ ae-cli engage-task indicator-user export \
30
+ --project-id 1 \
31
+ --task-id task_123 \
32
+ --indicator metric \
33
+ --metric-id metric_1 \
34
+ --start-time 2026-04-01 \
35
+ --end-time 2026-04-07 \
36
+ --artifact-format csv
37
+ ```
38
+
39
+ Use `sql` to obtain the generated user-detail SQL, `run` for a small inline result, and `export` for a downloadable full result.
40
+
41
+ ## Required input
42
+
43
+ All commands require `--project-id`, `--task-id`, `--indicator`, `--start-time`, and `--end-time`.
44
+
45
+ Dates use `yyyy-MM-dd`, and the start date must not be later than the end date. The task must exist under the specified project.
46
+
47
+ `--indicator` maps to the report population as follows:
48
+
49
+ | Value | Report population | Usage note |
50
+ |---|---|---|
51
+ | `original_trigger` | Original triggered users | Triggered tasks only |
52
+ | `plan` | Planned-delivery users | Task delivery funnel |
53
+ | `actual_trigger` | Actual-trigger/actual-delivery users | Task delivery funnel |
54
+ | `trigger_success` | Successfully triggered/delivered users | Task delivery funnel |
55
+ | `view` | Users who produced the configured view event | Requires view-event data |
56
+ | `click` | Users who produced the configured click event | Requires click-event data |
57
+ | `main` | Main conversion-indicator users | Combine with `--retention-type` for converted/lost users |
58
+ | `secondary` | Nth secondary conversion-indicator users | Requires `--secondary-index` |
59
+ | `metric` | Users behind a configured custom metric | Requires `--metric-id` |
60
+ | `click_activate` | Click activation users | Used in experiment activation reports |
61
+ | `main_activate` | Main-indicator activation users | Used in experiment activation reports |
62
+ | `secondary_activate` | Nth secondary-indicator activation users | Requires `--secondary-index`; used in experiment activation reports |
63
+
64
+ ## Conditional input
65
+
66
+ - `secondary` and `secondary_activate` require `--secondary-index 1..10`. Other indicators reject this flag.
67
+ - `metric` requires `--metric-id`; obtain the ID with `ae-cli engage-task metric list`. Other indicators reject this flag.
68
+ - `main`, `secondary`, `main_activate`, and `secondary_activate` must reference a conversion indicator configured on the task.
69
+ - `click_activate`, `main_activate`, and `secondary_activate` use `source=experiment` and `group-by=experiment`; explicitly supplied values must also be `experiment`. The task must have an activation event configured. `click_activate` additionally requires a click-event experiment.
70
+ - Non-metric indicators grouped by experiment use `source=experiment`, and the task must be an experiment task. Metric indicators continue to use `source=metric`.
71
+ - Explicit `--source experiment` always requires an experiment task, even with date, batch, or trigger grouping.
72
+ - `--group-by experiment --is-summary false` requires `--exp-group-id`. The `plan` indicator never accepts `--exp-group-id`; when grouped by experiment, it supports summary queries only.
73
+ - `--group-by batch` is supported only by non-triggered tasks. A non-summary batch segment requires `--task-instance-id`.
74
+ - `--group-by trigger` and `original_trigger` are supported only by triggered tasks.
75
+ - `--retention-type lost` supports `actual_trigger`, `trigger_success`, `view`, `click`, `main`, and `secondary`; it also supports `plan` for triggered tasks. It is rejected for `original_trigger`, `metric`, all activate indicators, and `plan` on non-triggered tasks.
76
+
77
+ ## Optional report context
78
+
79
+ - `--group-by` selects `batch`, `date`, `trigger`, or `experiment`; it defaults to `experiment` for activate indicators and `date` otherwise.
80
+ - `--is-summary` defaults to `true`. Set it to `false` when reproducing a drilled-down batch or experiment row.
81
+ - `--retention-type` selects `retention` or `lost`; the default is `retention`. See the supported combinations above.
82
+ - `--source` selects `task`, `experiment`, or `metric`. It defaults to `metric` for metric indicators, `experiment` for activate indicators or experiment grouping, and `task` otherwise.
83
+ - `--task-instance-id` narrows a non-triggered task to one execution instance.
84
+ - `--exp-group-id` narrows an experiment report to one group.
85
+ - `--push-language-code`, `--user-time-zone`, and `--show-time-zone` reproduce the corresponding task-report filters. Copy their values from the report context when needed.
86
+ - `--push-language-code` accepts `all`, `default`, `ar`, `az`, `bs`, `ca`, `zh-Hans`, `zh-Hant`, `hr`, `cs`, `da`, `nl`, `en`, `et`, `fi`, `fr`, `ka`, `bg`, `de`, `el`, `hi`, `he`, `hu`, `id`, `it`, `ja`, `ko`, `lv`, `lt`, `ms`, `nb`, `fa`, `pl`, `pt`, `pa`, `ro`, `ru`, `sr`, `sk`, `es`, `sv`, `th`, `tr`, `uk`, or `vi`.
87
+ - `--user-time-zone` accepts `all` or a decimal offset from `-12` through `14` that exists in the task's execution details. An unavailable value is rejected instead of falling back to all users.
88
+ - `--show-time-zone` accepts offsets from `-12` through `14`. Omit it or pass `99` to use the backend server's default time zone.
89
+
90
+ ## Execution input and output
91
+
92
+ - `sql` returns `sql` and `indicator` and does not execute the query.
93
+ - `run` accepts `--request-id`, `--limit`, and `--timeout-seconds`. The default limit is `100` and the maximum is `1000`; the default timeout is `120` seconds and the maximum is `180` seconds. It returns `#user_id`, `#account_id`, and `#distinct_id` for each user.
94
+ - `export` accepts `--request-id`, `--artifact-format jsonl|csv`, and `--timeout-seconds`. The default format is `jsonl`; the maximum timeout is six hours. It returns `run_id` and `artifact_id`.
95
+ - For cross-user data sources, `#account_id` and `#distinct_id` are empty strings.
96
+
97
+ Use the query lifecycle commands to inspect and download an export:
98
+
99
+ ```bash
100
+ ae-cli engage-query run inspect --run-id <run_id>
101
+ ae-cli engage-query artifact download --artifact-id <artifact_id> --output ./task-indicator-users.csv.gz
102
+ ```
103
+
104
+ `--artifact-format` controls the exported file. The global `--format json|table` only controls CLI display output.
@@ -0,0 +1,31 @@
1
+ # engage-task user-detail export
2
+
3
+ Export user details for one execution instance of a non-triggered engagement task.
4
+
5
+ > Capability id: `engage-task.user-detail.export` · Domain: `engage` · Output: asynchronous artifact.
6
+
7
+ ```bash
8
+ ae-cli engage-task user-detail export \
9
+ --project-id <project-id> \
10
+ --task-id <task-id> \
11
+ --task-instance-id <task-instance-id> \
12
+ --user-status fail \
13
+ --artifact-format csv
14
+ ```
15
+
16
+ `--user-status` accepts English values only. The statuses below follow the task funnel from audience filtering to the final delivery result:
17
+
18
+ 1. `deduplicate` — users removed because the configured push ID is duplicated.
19
+ 2. `frequency_control` — users removed by the task-level frequency-control rules.
20
+ 3. `sample` — users excluded by experiment-layer sampling before planned delivery.
21
+ 4. `push_plan` — users retained in the planned-delivery population after deduplication, frequency control, and sampling.
22
+ 5. `fatigue_control` — planned users removed by channel-level fatigue-control rules.
23
+ 6. `push_actual` — users retained in the actual-delivery population after fatigue control.
24
+ 7. `exp_skip_push` — actual-delivery users intentionally not sent because their experiment control group is configured to skip delivery.
25
+ 8. `success` — users for whom delivery completed successfully.
26
+ 9. `fail` — users for whom delivery was attempted but failed. This is the default value.
27
+
28
+ Use `--task-exec-detail-id` to narrow an instance to one execution detail. Triggered tasks are not supported in this version.
29
+ Each artifact contains at most 1,000,000 rows, matching the Hermes task export safety boundary.
30
+
31
+ The command returns `run_id` and `artifact_id`. Use `ae-cli engage-query run inspect` to inspect progress and `ae-cli engage-query artifact download` to download the completed artifact.
@@ -21,7 +21,8 @@ description: "Interactive generation of an AE tracking plan and upload. Trigger
21
21
  | 埋点模板 | Tracking Plan Template | Pre-built industry/genre xlsx templates |
22
22
  | 方案名称 | Plan Name | User-facing plan identifier |
23
23
  | 应用场景 | Application Scenario | One-sentence description of what the app does |
24
- | 素材来源 | Source Material Type | prd / chat / codebase / template |
24
+ | 素材来源 | Source Material Type | prd / chat / codebase / template / data |
25
+ | 数据样本 / 文件画像 | Data Sample / File Profile | Column→property mapping source from an `ae-data-integration` inspect profile (`source_type: data`) |
25
26
  | 业务维度 | Business Dimension | Revenue model, core loop, functional entries, currency system |
26
27
  | 收入模型 | Revenue Model | IAA / IAP / mixed / subscription / commission |
27
28
  | 核心循环 | Core Loop | Core gameplay loop (e.g. "grind stages → earn coins → gacha for heroes") |
@@ -115,8 +116,9 @@ Choose your source material (up to 2):
115
116
  3 - Codebase (local project path; hidden in sandbox) — Analyze source code to extract events and properties
116
117
  4 - Pre-built template (built-in industry and game genre templates) — Select a built-in template
117
118
  5 - Modify existing tracking plan (local AE format xlsx file) — Import an existing tracking plan xlsx as baseline for modification; can be combined with Product doc / Description / Codebase, but NOT with Pre-built template
119
+ 6 - Data sample / file profile (from ae-data-integration inspect) — Map data columns (CSV/Excel/JSONL) to events & properties from an inspect profile; single-source path, not combinable with other options
118
120
 
119
- Reply with number(s), e.g. 1,5 or 4. Select up to 2.
121
+ Reply with number(s), e.g. 1,5 or 4. Select up to 2 (option 6 is single-source).
120
122
  ```
121
123
 
122
124
  If in a sandbox environment, ask exactly:
@@ -128,8 +130,9 @@ Choose your source material (up to 2):
128
130
  2 - Detailed description (conversational) — Describe app business flow, core features, user behaviors, monetization model, etc.
129
131
  3 - Pre-built template (built-in industry and game genre templates) — Select a built-in template
130
132
  4 - Modify existing tracking plan (sandbox workspace path) — Import an existing tracking plan xlsx as baseline for modification; can be combined with Product doc / Description, but NOT with Pre-built template
133
+ 5 - Data sample / file profile (from ae-data-integration inspect) — Map data columns (CSV/Excel/JSONL) to events & properties from an inspect profile; single-source path, not combinable with other options
131
134
 
132
- Reply with number(s), e.g. 1,4 or 3. Select up to 2.
135
+ Reply with number(s), e.g. 1,4 or 3. Select up to 2 (option 5 is single-source).
133
136
  ```
134
137
 
135
138
  Do not rewrite this source material list as unnumbered bullets, cards, or prose. The user must be able to reply with the visible numbers.
@@ -153,8 +156,10 @@ Based on user selection, determine source material type and record to `meta.sour
153
156
  | Codebase only | `codebase` | Scan source code, extract events/properties from business logic |
154
157
  | Template only | `template` | Provide built-in template selection |
155
158
  | Existing plan only | `existing_plan` | Import xlsx as baseline (see "Modify Existing Tracking Plan Flow" below) |
159
+ | Data sample only | `data` | Read the inspect profile (`ae-local-data-profile/v1`), map columns → events/properties (see "Data-path (source_type = data)" below) |
156
160
  | Any two-item combo | Join two types with `_` | First as baseline, second as supplement (priority: existing_plan → template → codebase → prd → chat) |
157
161
  | Existing plan + Pre-built template | **NOT allowed** | Both provide event baselines; semantic conflict |
162
+ | Data sample + any other | **NOT allowed** | Data sample is a standalone single-source path |
158
163
 
159
164
  **Follow-up questions** (ask in follow-up order defined in Multi-Source Combination Rules below):
160
165
  - Product doc → if not in a sandbox environment, ask exactly:
@@ -184,6 +189,7 @@ Based on user selection, determine source material type and record to `meta.sour
184
189
  - Codebase → ask **"What is the project directory path?"**, then scan source to extract business logic
185
190
  - Pre-built template → display matching templates for user confirmation
186
191
  - Modify existing tracking plan → ask **"Please provide the xlsx file path of your existing tracking plan"**, then follow the flow below
192
+ - Data sample → ask **"Please provide the path of the inspect profile JSON (or the run directory containing it)"**, then read the `ae-local-data-profile/v1` product and follow the "Data-path (source_type = data)" flow below
187
193
 
188
194
  **Modify Existing Tracking Plan Flow** (when user selects this option):
189
195
 
@@ -363,6 +369,49 @@ Use the baseline source's inference method as primary. Supplementary sources (es
363
369
  }
364
370
  ```
365
371
 
372
+ ### Data-path (source_type = data) — condensed single-gate flow
373
+
374
+ When the user selects the **Data sample / file profile** option (`source_type = data`), the flow diverges from the standard 5-item anchor. This path **maps table columns → events/properties** instead of inventing events from business understanding, and uses a **single confirmation gate** instead of the 5-segment Refine loop.
375
+
376
+ **What to skip** (only Item 1 — Application Scenario and Item 2 — Data sample are collected):
377
+
378
+ - **Item 3 (SDK Integration Config)** → skip. `meta.sdk_integration_mode = "none"` (data ingested via RESTful / LogBus / DataX, no SDK). No SDK auto-track events are injected (Phase 1.4 already skips `none`).
379
+ - **Item 4 (User Identity System)** → skip the visitor-ID strategy question. Derive `meta.user_identity` from the inspect profile's `identity_candidates` instead (e.g. a `distinct_id` / `account_id` column), `account_id_source: "user_account"` when an account column exists, otherwise `"none"`.
380
+ - **Business Dimension confirmation** → skip. Mapping is driven by columns, not by revenue model / core loop. `meta.business_dimension` stays empty.
381
+
382
+ **Draft construction (replaces Phase 1.1/1.2/1.3 for the data path)**:
383
+
384
+ 1. **Read the inspect profile**: the `ae-local-data-profile/v1` JSON product (columns / types / samples / UE eligibility / mapping confidence). If it is not present or is stale, re-run `ae-cli data-integration inspect` for the source file first.
385
+ 2. **Event model determination** (reuse UE routing): single-table single-event `track` / single-table multi-event (event-name column) / single-table `user_set` / mixed. The agent may propose **splitting one table into multiple events** (e.g. an ad table split by `campaign_type` into `ad_show` / `ad_click`); such proposals MUST be confirmed in the gate.
386
+ 3. **Column → property mapping draft**:
387
+ - Identify system columns first: time field, `distinct_id` / `account_id`, event-name column, user-property-name column.
388
+ - Map the remaining columns to event properties / user properties / super properties.
389
+ - Naming: `snake_case` event/property names + `display_name` + `desc` + `event_tag` (language follows the user's input).
390
+ - Type inference: CSV columns default to `string`; infer `number` / `bool` / `datetime` / enum from field name + value distribution + business doc/prompt priors. **Uncertain or conflicting columns are marked "to-confirm" and asked only inside the gate** (do not ask column-by-column beforehand).
391
+ 4. **Single confirmation gate** (replaces Phase 2, see below).
392
+ 5. **Merge with existing plan** (reuse Phase 4.1/4.2 conflict detection).
393
+ 6. **Persist**: `.ae-cli/draft.json` → xlsx → upload (`sdk_integration_mode = none`).
394
+
395
+ **Single confirmation gate (one round, one summary table)**:
396
+
397
+ Present ONE merged table covering all of the following in a single message, then wait for a single reply:
398
+
399
+ - **Event list**: `event_name` / `display_name` / `desc` / `event_tag` / `platform`.
400
+ - **Property list**: `name` / `display_name` / `type` / `desc` / `source`, with uncertain types highlighted (marked "to-confirm").
401
+ - **Field scope**: default to plan-internal fields only; include a **full-import switch** for bringing all source columns in.
402
+ - **Unrecognized / dirty data handling**: how unmapped columns, null values, and unparseable rows are treated.
403
+ - **Same-name property type conflicts**: flagged inline in the gate (see Phase 4.2 Type A).
404
+
405
+ User replies once: `ok` (accept all), or targeted edits — rename / retype / add / remove individual columns or events. After the gate is confirmed, jump to Phase 3 (project token) then Phase 4 (merge + upload); do NOT enter the 5-segment Refine loop.
406
+
407
+ **dry-run mode**:
408
+
409
+ When the user requests dry-run (or the caller passes `data-integration plan --dry-run`), produce the draft preview + column→property mapping summary ONLY:
410
+
411
+ - Show the single confirmation gate table (events + properties + field scope + unrecognized-data handling) and the mapping result, but do NOT write `.ae-cli/draft.json`.
412
+ - Do NOT generate xlsx, do NOT archive to `plans/`, do NOT upload to AE.
413
+ - State explicitly that nothing was persisted; the user can approve a real run afterwards.
414
+
366
415
  ### Item 3 — SDK Integration Config (client + server combined)
367
416
 
368
417
  Ask: **"What is your client platform? (multi-select OK, e.g. Android + iOS) Will you integrate a server-side SDK?"**
@@ -822,6 +871,8 @@ AE_LANG=<user_lang> ae-cli tracking code import-template --template-name "<templ
822
871
 
823
872
  ## Phase 2 — Refine (5-segment loop)
824
873
 
874
+ > **Data path (`source_type = data`)**: skip this 5-segment loop entirely — use the single confirmation gate defined in "Data-path (source_type = data)" instead.
875
+
825
876
  In order, one conversation round per segment:
826
877
 
827
878
  1. **sdk_config (SDK config + User identity, combined)** — Show SDK integration mode, platform/language, visitor ID strategy, account ID source, corresponding SDK calls
@@ -83,4 +83,3 @@ ae-cli analysis input-file purpose list --project-id 1
83
83
  ## Related Skills
84
84
 
85
85
  - **`ae-analysis`**: `analysis-meta event get` / `analysis-meta property get` for detail, and `analysis-meta event list` / `analysis-meta property list` to discover names.
86
- - **Dev workflow**: `.cursor/skills/te-cli-capability-gateway` — inspect gateway schema when adding new metadata capabilities.
@@ -18,7 +18,7 @@ ae-cli metadata property bind-existing-dimension-table --project-id <project_id>
18
18
  | `--property-scope` | Yes | Property owner table, for example `event` or `user`. |
19
19
  | `--data-table-id` | Yes | Existing dimension data table ID. |
20
20
  | `--timestamp-join-format` | No | Timestamp join format. |
21
- | `--dict-columns` | No | Dictionary column names JSON array. |
21
+ | `--dict-columns` | No | Dictionary column names JSON array. The CLI converts each string to the canonical `{"column_name":"..."}` object required by the capability. |
22
22
 
23
23
  ## Decision Rules
24
24