@thinkingai/ae-cli 6.1.18 → 6.1.20

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (127) hide show
  1. package/README.md +97 -38
  2. package/README.zh.md +101 -38
  3. package/dist/{auth-ROB2EDYV.js → auth-FUM37MHF.js} +241 -126
  4. package/dist/{capability-DKMYUTLC.js → capability-AXFQW5WM.js} +49 -35
  5. package/dist/{chunk-JHENBQ5B.js → chunk-4P355ZWB.js} +70 -1
  6. package/dist/chunk-6ZIQV5GW.js +26 -0
  7. package/dist/{chunk-TUKQZTMI.js → chunk-7K24F7T2.js} +2 -0
  8. package/dist/{chunk-OO6XR6LK.js → chunk-AMBFK2K3.js} +2 -0
  9. package/dist/chunk-ATSM5XAW.js +623 -0
  10. package/dist/{chunk-4XXOWOTA.js → chunk-BBEFP4SB.js} +31 -38
  11. package/dist/{chunk-BYYS3ANB.js → chunk-CZU3V3DQ.js} +5 -15
  12. package/dist/{chunk-VTWMIC5L.js → chunk-E2JKXMVO.js} +2 -0
  13. package/dist/chunk-ECBLHAZO.js +15 -0
  14. package/dist/{sync-HKIOZXQE.js → chunk-I4WQAEYB.js} +31 -540
  15. package/dist/{chunk-DWO43OIB.js → chunk-JQ3ENZZH.js} +2 -0
  16. package/dist/chunk-LNZBEQXW.js +25216 -0
  17. package/dist/{chunk-ZQ47LWTI.js → chunk-QJQZH5GQ.js} +49 -79
  18. package/dist/{chunk-QZ3AS4KK.js → chunk-RSG4ONOI.js} +9 -8
  19. package/dist/{chunk-4NN5IWVN.js → chunk-T6OSFQZD.js} +2 -0
  20. package/dist/{chunk-Y3LOALAV.js → chunk-TAL6CZH6.js} +8 -7
  21. package/dist/{chunk-3KWQYGYI.js → chunk-TKHSULJT.js} +2 -0
  22. package/dist/chunk-VXNMYUXU.js +202 -0
  23. package/dist/{chunk-3FY3RJ26.js → chunk-WZ6YVQSF.js} +15 -14
  24. package/dist/{chunk-6EIJSNBD.js → chunk-Y74WTIKJ.js} +2 -0
  25. package/dist/{community-report-client-C7WDGET3.js → community-report-client-XXWGSBSD.js} +3 -4
  26. package/dist/{config-BMYZX2UE.js → config-EXUGQN5T.js} +10 -10
  27. package/dist/{data-integration-QEKDWQDY.js → data-integration-4NZ53OMT.js} +920 -97
  28. package/dist/index.js +56 -81
  29. package/dist/{local-data-upload-client-4YYHSYD6.js → local-data-upload-client-KYOKVYZV.js} +3 -4
  30. package/dist/{memory-I2WXDTV2.js → memory-ATNKZNW5.js} +6 -7
  31. package/dist/{metadata-I4C2EWUN.js → metadata-VZVC7YMH.js} +10 -11
  32. package/dist/{model-HLHIEFMU.js → model-E4JMQ4V2.js} +8 -9
  33. package/dist/{personal-semantic-preference-LIPACBDX.js → personal-semantic-preference-7S2SQ3UV.js} +9 -10
  34. package/dist/{project-semantic-RT3R2VQD.js → project-semantic-2SZP2OPO.js} +13 -14
  35. package/dist/sync-YV3E66IF.js +520 -0
  36. package/dist/{te-agent-BR6VDBNX.js → te-agent-JB5T3PO7.js} +396 -92
  37. package/dist/{te-analysis-7VUNUYWZ.js → te-analysis-3YJAAT2D.js} +196 -43
  38. package/dist/{te-community-5DMNKJWY.js → te-community-UDBI672N.js} +12 -34
  39. package/dist/{te-dataops-6P5IKWNJ.js → te-dataops-ZLYOCXZ4.js} +480 -81
  40. package/dist/{te-engage-KZPR5R22.js → te-engage-4XG6OJML.js} +88 -16
  41. package/dist/{te-experiment-6BITX4RD.js → te-experiment-VXUWPINJ.js} +83 -12
  42. package/dist/te-kb-WYQWHFSC.js +1732 -0
  43. package/dist/{te-system-FXITO2JG.js → te-system-7G6F2LJA.js} +569 -35
  44. package/dist/{te-team-ADOC2ROP.js → te-team-E7FBBXMQ.js} +8 -7
  45. package/dist/{update-YCYCKJOO.js → update-D47BUG25.js} +8 -8
  46. package/package.json +22 -10
  47. package/skills/ae-agent/SKILL.md +30 -13
  48. package/skills/ae-agent/references/agent-distribution.md +94 -0
  49. package/skills/ae-agent/references/approval-request.md +4 -0
  50. package/skills/ae-agent/references/command_index.md +9 -2
  51. package/skills/ae-agent/references/create-automation.md +20 -2
  52. package/skills/ae-agent/references/get-agent-context.md +70 -0
  53. package/skills/ae-agent/references/list-automations.md +18 -3
  54. package/skills/ae-agent/references/list-mcp-credentials.md +1 -1
  55. package/skills/ae-agent/references/mcp-token.md +3 -3
  56. package/skills/ae-agent/references/set-mcp-credential.md +0 -1
  57. package/skills/ae-agent/references/update-automation.md +18 -0
  58. package/skills/ae-analysis/SKILL.md +11 -2
  59. package/skills/ae-analysis/references/adhoc_run.md +2 -0
  60. package/skills/ae-analysis/references/ai_models.md +23 -3
  61. package/skills/ae-analysis/references/analysis_gateway_assets.md +3 -3
  62. package/skills/ae-analysis/references/audience_models.md +18 -0
  63. package/skills/ae-analysis/references/command_index.md +9 -9
  64. package/skills/ae-analysis/references/cross_source_config.md +84 -0
  65. package/skills/ae-analysis/references/dashboard_update.md +1 -1
  66. package/skills/ae-analysis/references/project_timezone_update.md +13 -4
  67. package/skills/ae-analysis/references/property_create.md +2 -0
  68. package/skills/ae-analysis/references/super_metadata_batch_create.md +2 -0
  69. package/skills/ae-analysis/references/user_cluster_models.md +2 -0
  70. package/skills/ae-analysis/references/user_cluster_update.md +8 -4
  71. package/skills/ae-analysis/references/user_tag_create.md +30 -2
  72. package/skills/ae-analysis/references/user_tag_models.md +17 -3
  73. package/skills/ae-analysis/references/user_tag_update.md +14 -2
  74. package/skills/ae-data-integration/SKILL.md +3 -1
  75. package/skills/ae-data-integration/references/dimension-routing.md +36 -0
  76. package/skills/ae-data-integration/references/error-handling.md +54 -1
  77. package/skills/ae-data-integration/references/local-analysis.md +2 -0
  78. package/skills/ae-data-integration/references/source-inspect.md +18 -2
  79. package/skills/ae-data-integration/references/tracking-plan.md +1 -1
  80. package/skills/ae-data-integration/references/transform.md +4 -2
  81. package/skills/ae-data-integration/references/ue-mapping.md +5 -2
  82. package/skills/ae-data-integration/references/ue-routing.md +40 -1
  83. package/skills/ae-dataops/SKILL.md +11 -1
  84. package/skills/ae-dataops/references/dataops-backfill.md +135 -0
  85. package/skills/ae-engage/SKILL.md +5 -0
  86. package/skills/ae-engage/references/build-task-save-guide.md +5 -1
  87. package/skills/ae-engage/references/save-flow.md +37 -1
  88. package/skills/ae-engage/references/save-task.md +6 -0
  89. package/skills/ae-experiment/SKILL.md +6 -2
  90. package/skills/ae-experiment/references/save_metric.md +20 -8
  91. package/skills/ae-generate-tracking-plan/SKILL.md +25 -13
  92. package/skills/ae-generate-tracking-plan/references/business-dimension-mapping.md +1 -1
  93. package/skills/ae-kb/SKILL.md +218 -36
  94. package/skills/ae-kb/references/query-workflow.md +59 -29
  95. package/skills/ae-kb/references/versions.md +46 -0
  96. package/skills/ae-system/SKILL.md +29 -31
  97. package/skills/ae-system/references/channel-management.md +303 -0
  98. package/skills/ae-use-agent/SKILL.md +42 -0
  99. package/skills/ae-use-agent/references/local-agent.md +114 -0
  100. package/dist/auth-GBMV6TEJ.js +0 -14
  101. package/dist/capability-HYVVPG25.js +0 -352
  102. package/dist/chunk-EFH4XWYC.js +0 -556
  103. package/dist/chunk-J2DEBMRF.js +0 -313
  104. package/dist/chunk-JRJY5DMJ.js +0 -71
  105. package/dist/chunk-OMPRXM3V.js +0 -349
  106. package/dist/chunk-QNOLN2LJ.js +0 -509
  107. package/dist/chunk-RJDU7NYP.js +0 -1198
  108. package/dist/chunk-RNAALWJK.js +0 -98
  109. package/dist/chunk-SERWF6G5.js +0 -13
  110. package/dist/chunk-UW5UN47B.js +0 -70
  111. package/dist/chunk-ZQKDZXDO.js +0 -317
  112. package/dist/client-L2YDMHQ6.js +0 -203
  113. package/dist/memory-3ORCR7JH.js +0 -893
  114. package/dist/metadata-VUOQJE26.js +0 -339
  115. package/dist/model-UGRDX4MW.js +0 -139
  116. package/dist/personal-semantic-preference-OEISBRHM.js +0 -239
  117. package/dist/project-semantic-FFPWFPIW.js +0 -1114
  118. package/dist/sync-TFHU2UTG.js +0 -10261
  119. package/dist/te-agent-VLYOV7S4.js +0 -3894
  120. package/dist/te-analysis-4YGQL5RC.js +0 -9357
  121. package/dist/te-community-ISDQWJU7.js +0 -1859
  122. package/dist/te-dataops-CVULXNVB.js +0 -2209
  123. package/dist/te-engage-N5WI32H6.js +0 -4898
  124. package/dist/te-experiment-UVR4HLND.js +0 -988
  125. package/dist/te-kb-RCLSSH2Q.js +0 -935
  126. package/dist/te-system-K2GYMCTB.js +0 -2213
  127. package/skills/ae-agent/references/auto-provision-mcp-credentials.md +0 -57
@@ -13,6 +13,8 @@ ae-cli analysis-meta property create --dry-run
13
13
 
14
14
  Capability id: `metadata.property.create`.
15
15
 
16
+ Authorization requires the single project function permission `editSuperMeta` with the `metadata:write` scope. In the zh-CN permission UI, this permission is labeled `元数据管理 > 编辑`; the corresponding English label is `Metadata Management > Edit`. This is the same project permission used by `metadata.super_metadata.batch_create`. If either command returns `PROJECT_PERMISSION_DENIED`, do not retry with a different payload and do not describe the two capability IDs as two separate permissions. Ask a project administrator to grant this shared project permission to the current identity.
17
+
16
18
  Input sends `project_id`, `table_type`, `payload`.
17
19
 
18
20
  Output is a successful gateway envelope with no business data. Read back with `property get` using the same table type.
@@ -12,6 +12,8 @@ ae-cli analysis-meta super-metadata batch-create --project-id <project_id> --eve
12
12
 
13
13
  Capability id: `metadata.super_metadata.batch_create`.
14
14
 
15
+ Authorization requires the single project function permission `editSuperMeta` with the `metadata:write` scope. In the zh-CN permission UI, this permission is labeled `元数据管理 > 编辑`; the corresponding English label is `Metadata Management > Edit`. This is the same project permission used by `metadata.property.create`. If either command returns `PROJECT_PERMISSION_DENIED`, do not retry with a different payload and do not describe the two capability IDs as two separate permissions. Ask a project administrator to grant this shared project permission to the current identity.
16
+
15
17
  Input sends `project_id` plus any non-empty JSON arrays among `events`, `event_properties`, and `user_properties`. Use snake_case object fields exactly as documented by the common-service schema:
16
18
 
17
19
  - Event items: `event_name`, optional `event_desc`, optional `remark`, optional `super_event_prop_names`.
@@ -7,6 +7,8 @@ Top-level variants:
7
7
  - condition: `{"type":"condition","conditions":{"relation":"and|or","items":[...]}}`
8
8
  - SQL: `{"type":"sql","sql":"...","params":[{"name":"partdate","type":"part_date","recent_day":"1-7"}]}`
9
9
 
10
+ For `user-cluster update`, `type` must match the existing cluster type. The update command cannot change a cluster between condition and SQL, and it cannot change the cluster's analysis entity.
11
+
10
12
  For SQL clusters, quote Trino special identifiers with double quotes, for example `{"type":"sql","sql":"SELECT \"#user_id\" FROM v_user_1 WHERE vip_level >= 3"}`. This applies to identifiers containing `#`, `$`, `@`, spaces, or punctuation; single quotes are string literals.
11
13
 
12
14
  If a SQL cluster reads an event table, include a predicate on the quoted `"$part_date"` date-partition column; the backend rejects event-table SQL without it. This does not apply to the user-table example above.
@@ -2,18 +2,22 @@
2
2
 
3
3
  Update a condition or SQL user cluster. Discover the exact `cluster_name` first.
4
4
 
5
- Do not use it for ID-file membership replacement or to create a missing cluster. Supplying `--definition-request` automatically starts recomputation after the definition is updated; do not call `user-cluster refresh` afterward. Updating only `--display-name` or `--remark` does not recompute. `--auto-refresh-cron` changes an existing enabled auto-refresh schedule and does not enable auto refresh. A successful update means the definition was saved, not that the new result is complete. Poll `user-cluster get` until `progress=100` and `refresh_end_time` is not older than `update_time` before using `users_num` or querying members.
5
+ This is a `high-risk-write`. First run the final command with `--dry-run`, summarize the exact cluster and fields that will change, and wait for explicit user confirmation. Then run the unchanged command with `--yes`. Do not use it for ID-file membership replacement, to create a missing cluster, to change its analysis entity, or to change it between condition and SQL types.
6
+
7
+ Supplying `--definition-request` automatically starts recomputation after the definition is updated; do not call `user-cluster refresh` afterward. Updating only `--display-name` or `--remark` does not recompute. `--auto-refresh-cron` changes an existing enabled auto-refresh schedule and does not enable auto refresh. A successful update means the definition was saved, not that the new result is complete. Poll `user-cluster get` until `progress=100` and `refresh_end_time` is not older than `update_time` before using `users_num` or querying members.
6
8
 
7
9
  The response distinguishes both paths. A definition update returns `computation.triggered_automatically=true`, `result_freshness.is_stale=true`, and normally `next_action=poll_get` with an exact capability/input pair. A display-name/remark-only update returns `computation.status=not_triggered`, `result_freshness.status=fresh`, and `next_action=none`.
8
10
 
9
- Flags: `--project-id`, `--cluster-name` required. Optional: `--display-name`, `--definition-request`, `--authenticated-only`, `--remark`, `--zone-offset`, `--auto-refresh-cron`. The cluster type comes from `definition_request.type` when the definition changes.
11
+ Flags: `--project-id`, `--cluster-name` required. Optional: `--display-name`, `--definition-request`, `--authenticated-only`, `--remark`, `--zone-offset`, `--auto-refresh-cron`. There is no `--entity-id` update flag. When the definition changes, `definition_request.type` must match the existing cluster type; it selects the definition variant but does not change the saved type.
10
12
 
11
13
  `display_name` is at most 80 characters and `remark` is at most 400 characters. The CLI rejects violations before dispatch. `cluster_name` is an existing exact identifier and cannot be renamed by update.
12
14
 
13
15
  Read `user_cluster_models.md` before changing the definition. The backend validates and compiles `definition_request` inside update and refuses to modify the cluster if clarification is required. Condition definitions are saved as mixed-condition clusters.
14
16
 
15
17
  ```bash
16
- ae-cli analysis user-cluster update --project-id <project_id> --cluster-name retained_users --display-name "Retained Users v2"
18
+ ae-cli analysis user-cluster update --project-id <project_id> --cluster-name retained_users --display-name "Retained Users v2" --dry-run
19
+
20
+ ae-cli analysis user-cluster update --project-id <project_id> --cluster-name retained_users --display-name "Retained Users v2" --yes
17
21
 
18
- ae-cli analysis user-cluster update --project-id <project_id> --cluster-name retained_users --auto-refresh-cron '0 30 2 * * ? *'
22
+ ae-cli analysis user-cluster update --project-id <project_id> --cluster-name retained_users --auto-refresh-cron '0 30 2 * * ? *' --dry-run
19
23
  ```
@@ -6,14 +6,42 @@ Do not use it for uploaded-ID tags; use `user-tag create-id`. A successful creat
6
6
 
7
7
  The response reports this directly: `computation.triggered_automatically=true`, `computation.status=submitted`, and `result_freshness.is_stale=true`. Follow `next_action`; when it is `poll_get`, invoke `next_capability_id` with the exact `next_input` returned by the command.
8
8
 
9
- Flags: `--project-id`, `--tag-name`, `--display-name`, `--definition-request` required. Optional: `--authenticated-only`, `--zone-offset`, `--entity-id`. The tag type comes from `definition_request.type`.
9
+ Flags: `--project-id`, `--tag-name`, `--display-name`, `--definition-request` required. Optional: `--authenticated-only`, `--zone-offset`, `--entity-id`, `--enable-auto-refresh`, `--auto-refresh-schedule`, `--auto-refresh-cron`. The tag type comes from `definition_request.type`.
10
10
 
11
11
  `tag_name` is a machine identifier: 1-80 characters, starts with a letter, and contains only letters, digits, or underscores. `display_name` is 1-80 characters. The CLI rejects violations before dispatch.
12
12
 
13
- Read `user_tag_models.md` before constructing `--definition-request`. Create does not accept `--remark`; set it later with `user-tag update` when needed.
13
+ Read `user_tag_models.md` before constructing `--definition-request`. Dynamic first/last ranges use semantic `time_range` values such as `{"mode":"recent","unit":"month","value":1}` for this month or `{"mode":"start_to_today","start_time":"2026-07-01"}` for a fixed start date through today. Create does not accept `--remark`; set it later with `user-tag update` when needed.
14
14
 
15
15
  The backend validates and compiles the definition inside the create operation; if metadata is ambiguous or missing, creation fails without creating the tag.
16
16
 
17
17
  ```bash
18
18
  ae-cli analysis user-tag create --project-id <project_id> --tag-name high_value --display-name "High Value" --definition-request '{"type":"condition","condition_values":[{"value":"high","events":[{"event":"pay","operator":"gte","value":3,"aggregation":"count","time_range":{"mode":"recent","unit":"day","value":30}}]}]}'
19
19
  ```
20
+
21
+ First/last tag for this month:
22
+
23
+ ```bash
24
+ ae-cli analysis user-tag create --project-id <project_id> --tag-name latest_platform_this_month --display-name "Latest Platform This Month" --definition-request '{"type":"first_last","first_last":{"event":"login","occurrence":"last","property":"platform","time_range":{"mode":"recent","unit":"month","value":1}}}'
25
+ ```
26
+
27
+ Periodic refresh can be configured in this same create operation. It is separate from the initial computation. Omitted scheduling flags leave periodic refresh disabled; supplying a schedule enables it. `--enable-auto-refresh true` requires a schedule on create. Do not combine `--enable-auto-refresh false` with a schedule, or pass both schedule forms.
28
+
29
+ Append one of these options to the create command:
30
+
31
+ ```bash
32
+ # Every day at 02:30 in the tag timezone
33
+ --auto-refresh-schedule '{"frequency":"daily","time":"02:30"}'
34
+
35
+ # Monday and Sunday at 09:00 (ISO weekdays: 1=Monday, 7=Sunday)
36
+ --auto-refresh-schedule '{"frequency":"weekly","time":"09:00","weekdays":[1,7]}'
37
+
38
+ # The 1st and 15th of each month at 06:00
39
+ --auto-refresh-schedule '{"frequency":"monthly","time":"06:00","month_days":[1,15]}'
40
+
41
+ # Custom Quartz schedule, including multiple executions per day
42
+ --auto-refresh-cron '0 0/30 8-18 * * ? *'
43
+ ```
44
+
45
+ `time` uses 24-hour `HH:mm`. `month_days` accepts 1-31; dates absent from a month are skipped. Weekly/monthly schedules require their respective day array; other frequencies reject those fields. The structured form preserves the page's daily/weekly/monthly frequency selection. Cron uses the page's custom schedule mode and preserves every cron field.
46
+
47
+ Schedules use the tag timezone. Use the existing `--zone-offset` only with an offset supported by the project; omit it to use the project's default behavior. Inspect `user-tag get` for `enable_auto_refresh` (1=enabled, 0=disabled), `scheduler_ui_config`, and `cluster_zone_offset` after creation.
@@ -24,18 +24,32 @@ Top-level `type` is exactly one of `condition`, `metric`, `first_last`, or `sql`
24
24
 
25
25
  ## Metric tag
26
26
 
27
- Required: `event`, `aggregation`. `property`, `time_range`, and `filters` are optional. Filters support only event properties and user properties. A string `field` is an event property; use `{name,type:"user_property"}` for a user property.
27
+ Required: `event`, `aggregation`. `property`, `percentile`, `time_range`, and `filters` are optional. Filters support only event properties and user properties. A string `field` is an event property; use `{name,type:"user_property"}` for a user property.
28
+
29
+ Use `aggregation=percentile` with a numeric event `property` and pass `percentile`. Supported percentile values match the page controls: `5`, `10`, `20`, `25`, `30`, `40`, `60`, `70`, `75`, `80`, `90`, `95`, and `99`. The `percentile` field is required for percentile aggregation and is rejected for every other aggregation.
28
30
 
29
31
  ```json
30
32
  {"type":"metric","metric":{"event":"pay","aggregation":"sum","property":"amount","time_range":{"mode":"previous","unit":"day","value":30},"filters":{"relation":"and","items":[{"field":"channel","operator":"eq","values":["app"]},{"field":{"name":"country","type":"user_property"},"operator":"eq","values":["US"]}]}}}
31
33
  ```
32
34
 
35
+ Percentile example:
36
+
37
+ ```json
38
+ {"type":"metric","metric":{"event":"pay","aggregation":"percentile","property":"amount","percentile":90}}
39
+ ```
40
+
33
41
  ## First/last tag
34
42
 
35
- Required: `event`, `occurrence=first|last`, and exactly one value source: `calculation` or `property`. `time_range` and `filters` are optional. Filters support only event properties and user properties. A string `field` is an event property; use `{name,type:"user_property"}` for a user property. Supplying neither or both value sources is rejected before execution.
43
+ Required: `event`, `occurrence=first|last`, and exactly one value source: `calculation` or `property`. `time_range` and `filters` are optional. Filters support only event properties and user properties. A string `field` is an event property; use `{name,type:"user_property"}` for a user property. Supplying neither or both value sources is rejected before execution. Use the semantic time mappings in [`audience_models.md`](audience_models.md) for dynamic ranges such as today, this month, or a fixed start date through today.
44
+
45
+ ```json
46
+ {"type":"first_last","first_last":{"event":"login","occurrence":"last","property":"platform","time_range":{"mode":"recent","unit":"month","value":1},"filters":{"relation":"and","items":[{"field":{"name":"country","type":"user_property"},"operator":"eq","values":["US"]}]}}}
47
+ ```
48
+
49
+ From a fixed date through today:
36
50
 
37
51
  ```json
38
- {"type":"first_last","first_last":{"event":"login","occurrence":"last","property":"platform","time_range":{"mode":"recent","unit":"day","value":30},"filters":{"relation":"and","items":[{"field":{"name":"country","type":"user_property"},"operator":"eq","values":["US"]}]}}}
52
+ {"type":"first_last","first_last":{"event":"login","occurrence":"first","calculation":"specific_time","time_range":{"mode":"start_to_today","start_time":"2026-07-01"}}}
39
53
  ```
40
54
 
41
55
  ## SQL tag
@@ -2,11 +2,11 @@
2
2
 
3
3
  Update a user tag. Discover the exact `tag_name` first.
4
4
 
5
- Do not use it for ID-file value replacement or to create a missing tag. Supplying `--definition-request` automatically starts recomputation after the definition is updated; do not call `user-tag refresh` afterward. Updating only `--display-name` or `--remark` does not recompute. `--auto-refresh-cron` changes an existing enabled auto-refresh schedule and does not enable auto refresh. A successful update means the definition was saved, not that the new result is complete. Record the current `refresh_time` before updating, then poll `user-tag get` until `progress=100` and `refresh_time` advances before using `users_num` or querying members.
5
+ Do not use it for ID-file value replacement or to create a missing tag. Supplying `--definition-request` automatically starts recomputation after the definition is updated; do not call `user-tag refresh` afterward. Updating only `--display-name` or `--remark` does not recompute. A schedule-only update changes periodic refresh without starting immediate recomputation. `--auto-refresh-cron` or `--auto-refresh-schedule` can enable periodic refresh directly, including on a previously disabled tag. A successful update means the definition was saved, not that the new result is complete. Record the current `refresh_time` before updating, then poll `user-tag get` until `progress=100` and `refresh_time` advances before using `users_num` or querying members.
6
6
 
7
7
  The response distinguishes both paths. A definition update returns `computation.triggered_automatically=true`, `result_freshness.is_stale=true`, and normally `next_action=poll_get` with an exact capability/input pair. A display-name/remark-only update returns `computation.status=not_triggered`, `result_freshness.status=fresh`, and `next_action=none`.
8
8
 
9
- Flags: `--project-id`, `--tag-name` required. Optional: `--display-name`, `--definition-request`, `--authenticated-only`, `--remark`, `--zone-offset`, `--auto-refresh-cron`. The tag type comes from `definition_request.type` when the definition changes.
9
+ Flags: `--project-id`, `--tag-name` required. Optional: `--display-name`, `--definition-request`, `--authenticated-only`, `--remark`, `--zone-offset`, `--enable-auto-refresh`, `--auto-refresh-schedule`, `--auto-refresh-cron`. The tag type comes from `definition_request.type` when the definition changes.
10
10
 
11
11
  `display_name` is at most 80 characters and `remark` is at most 400 characters. The CLI rejects violations before dispatch. `tag_name` is an existing exact identifier and cannot be renamed by update.
12
12
 
@@ -17,3 +17,15 @@ ae-cli analysis user-tag update --project-id <project_id> --tag-name user_level
17
17
 
18
18
  ae-cli analysis user-tag update --project-id <project_id> --tag-name user_level --auto-refresh-cron '0 30 2 * * ? *'
19
19
  ```
20
+
21
+ Omitting scheduling flags preserves the existing enable state and schedule, including during a definition update. Omitting `--zone-offset` preserves the saved tag timezone. `--enable-auto-refresh true` can reuse an existing saved plan; if none exists, provide a new schedule. `--enable-auto-refresh false` disables periodic refresh. Do not combine false with a schedule or pass both schedule forms. Structured schedule fields and timezone rules are described in [user_tag_create.md](user_tag_create.md).
22
+
23
+ ```bash
24
+ # Enable daily refresh in one update
25
+ ae-cli analysis user-tag update --project-id <project_id> --tag-name user_level --auto-refresh-schedule '{"frequency":"daily","time":"02:30"}'
26
+
27
+ # Disable periodic refresh
28
+ ae-cli analysis user-tag update --project-id <project_id> --tag-name user_level --enable-auto-refresh false
29
+ ```
30
+
31
+ Read `user-tag get` after updating to verify `enable_auto_refresh` (1=enabled, 0=disabled), `scheduler_ui_config`, and `cluster_zone_offset`. Schedule-only updates return `computation.status=not_triggered` and `next_action=none`.
@@ -17,6 +17,7 @@ Two entrances lead here: the AE Agent dialog (attach / plus-button upload) and `
17
17
  - Do not invent account IDs, distinct IDs, event times, event names, projects, APPIDs, receivers, or timezones.
18
18
  - `value_mapping` and `random_pool` are explicit user decisions. Never invent them.
19
19
  - Never auto-fill a missing time for `track`/`track_*` rows. A missing time on user-profile rows may be filled with the current time only by setting `missing_time: 'now'` and only after the user explicitly confirms it.
20
+ - The mapping's fill-in options (`missing_time: 'now'`, `account_id_value`/`distinct_id_value`, `random_pool`, `exclude_columns`, `value_mapping`) record a decision the user has already made; they are never to make a validation failure disappear. When `convert` quarantines rows, present the failure first — the error `code`, the row count, its share of the total, and which events or time ranges are affected — and state the consequence in the user's terms ("3,000 events would carry a time that is not when they happened"). Only after the user has seen that may one of these options be set. When the evidence is insufficient, report the batch as not passed and pending customer data; never shrink the upload scope to manufacture a pass.
20
21
  - Do not read or send an AE access token or CLI token to `/sync_json`. The receiver request uses only APPID and UE data.
21
22
  - Never execute `data-integration upload` until the user has seen the target, mapping, valid/quarantined counts, batches, and dry-run and has explicitly confirmed that upload.
22
23
  - A blocked manifest requires a second, explicit clean-subset decision. Never add `--allow-clean-subset` implicitly.
@@ -35,6 +36,7 @@ Use this skill when the user wants to bring a **local data file** (CSV/TSV/TXT/J
35
36
  | Generate / upload a project-level tracking plan (source material is PRD / chat / template / code; deliverable is a real platform tracking plan) | ae-generate-tracking-plan |
36
37
  | Upload documents / URLs to a knowledge base | ae-kb |
37
38
  | Reports / dashboards / queries / governance on data already in AE | ae-analysis |
39
+ | Dimension / dictionary data (a stable-entity lookup — city / product / device) to load as a dimension table bound to a property | ae-metadata |
38
40
 
39
41
  This skill also produces a tracking-plan draft (`source_type: data`) as a governance prerequisite; that draft is an input to ae-generate-tracking-plan, not a substitute for its five-phase platform plan.
40
42
 
@@ -42,7 +44,7 @@ This skill also produces a tracking-plan draft (`source_type: data`) as a govern
42
44
 
43
45
  Walk the four submodules in order. Each submodule is its own reference; follow it and come back here for the next step.
44
46
 
45
- 1. **Source — business identification.** Read [references/source-inspect.md](references/source-inspect.md). Profile every file fully, infer its business meaning using business-doc / user-prompt priors, then pick a branch via [references/ue-routing.md](references/ue-routing.md).
47
+ 1. **Source — business identification.** Read [references/source-inspect.md](references/source-inspect.md). Profile every file fully, infer its business meaning using business-doc / user-prompt priors, then pick a branch via [references/ue-routing.md](references/ue-routing.md): UE ingestion, dimension routing ([references/dimension-routing.md](references/dimension-routing.md)), or local analysis.
46
48
  2. **Reuse check.** If the profile is `ue_eligible`, read [references/reuse.md](references/reuse.md) and match the recommended mapping against the handoff index. `reuse` searches the current directory's `.ae-cli/data-integration/` upward, then `~/.ae-cli/data-integration/`, so a package written elsewhere is still found. A match proposes a frozen package; after one explicit confirmation, run the returned `transform.mjs` command and jump to Sink (step 5). No match → continue.
47
49
  3. **Tracking plan.** Read [references/tracking-plan.md](references/tracking-plan.md). The plan is generated from the mapping (`plan --mapping`), so confirm the recommended mapping's key system fields with the user first — `mode`, `#account_id`/`#distinct_id`, `#time` + timezone, `#event_name`, `#ip`/`#uuid` (see [references/transform.md](references/transform.md) steps 1–5) — then generate the event/property plan and get a single explicit confirmation from the user before touching data. The plan is a separate, required deliverable from the transform mapping: a user who supplies a column→field mapping directly has **not** completed this step, so build the plan from the confirmed mapping anyway. `user_set` still requires a plan (no events; every property becomes a user property). This step runs for **every** file: a second or later file merges its new events and properties into the existing project plan (tracking-plan.md step 4) — an existing plan is never a reason to skip it.
48
50
  4. **Transform.** Read [references/transform.md](references/transform.md). Map columns to AE system fields and properties, convert, and quarantine dirty rows per [references/ue-mapping.md](references/ue-mapping.md).
@@ -0,0 +1,36 @@
1
+ # Dimension routing
2
+
3
+ Use this reference after [ue-routing.md](ue-routing.md) has classified the file as dimension data — a stable-entity lookup with no row-level identity or event time. It covers the handoff to ae-metadata; this skill does not ingest dimension data itself.
4
+
5
+ ## What dimension data is
6
+
7
+ Judge by content, never by file extension — CSV / TSV / TXT / JSON / JSONL / XLS / XLSX can all be dimension data. The classification signals live in [ue-routing.md](ue-routing.md): no row-level identity, no row-level event time, finite entity enumeration, and a join key shared with event data.
8
+
9
+ Dimension data is not the only file that fails UE prerequisites. Aggregates, pivot tables, and cumulative snapshots also lack identity/time, but those are local-analysis material, not dictionaries. The tell is the entity shape and the join key: a dimension table maps one entity code to its attributes (`city_code` → name / level), while an aggregate summarizes many rows into one measure.
10
+
11
+ Low confidence is a proposal, never a silent decision — ask the user instead of routing automatically.
12
+
13
+ ## Handoff to ae-metadata
14
+
15
+ Extract the dimension sheet / file to CSV, then hand the following commands to ae-metadata. The sequence lists entry points only; full flags live in ae-metadata's references.
16
+
17
+ Bind prerequisite (confirm before creating the table):
18
+
19
+ - The table binds to an existing property (`property_name` + `property_scope` = user or event). That property usually appears once event/user data is uploaded first (e.g. events carrying `city_code`), so dimension binding is a second-phase action after data lands.
20
+ - If the target property does not exist yet, create it first via ae-analysis metadata or a tracking plan. Never invent a property name.
21
+
22
+ Entry command sequence:
23
+
24
+ ```bash
25
+ # 1. Extract the dimension sheet/file to CSV (metadata upload accepts CSV only,
26
+ # purpose data_table.csv)
27
+ ae-cli analysis input-file upload --project-id <id> --purpose data_table.csv --file <dim.csv>
28
+ # 2. Create + bind in one step (or split into csv-write + bind-existing)
29
+ ae-cli metadata property create-and-bind-csv-dimension-table --project-id <id> \
30
+ --property-name <p> --property-scope user|event --input-file-id ifile_xxx
31
+ # 3. Later dictionary changes (add/update/delete):
32
+ ae-cli metadata data-table csv-write --operation incremental_update|replace_update \
33
+ --data-table-id <id> --input-file-id ifile_xxx
34
+ ```
35
+
36
+ Binding model: AE attaches a dimension table to a user/event property, turning it into a dict property whose values join through the table's key column to expand `--dict-columns`. See ae-metadata's dimension-table reference for the full flags.
@@ -52,7 +52,7 @@ the user decide — never to guess an encoding or structure and retry silently.
52
52
  | `LOCAL_DATA_INPUT_INVALID` | Generic parse failure (malformed CSV/TSV/JSON/XLS) | Verify encoding and structure, then retry without changing the source |
53
53
  | `LOCAL_DATA_JSONL_INVALID` | A JSONL line is not valid JSON (`location.record` names the line) | Point at the offending record |
54
54
  | `LOCAL_DATA_JSON_ROOT_INVALID` | JSON root is not an object or array | Check the file's top-level shape |
55
- | `LOCAL_DATA_XLSX_INVALID` | Workbook metadata missing / no readable sheets / worksheet entry missing | Re-export the workbook |
55
+ | `LOCAL_DATA_XLSX_INVALID` | Workbook metadata missing / no readable sheets / worksheet entry missing | Re-export the workbook (sheet recognition accepts `<sheet>` and namespaced `<x:sheet>` alike, so this is a real metadata gap) |
56
56
  | `LOCAL_DATA_SET_NOT_FOUND` / `LOCAL_DATA_SET_REQUIRED` | Sheet or JSON Path not found / ambiguous | Ask which `--data-set` to use |
57
57
 
58
58
  A parse error is never a reason to change the mapping or the tracking plan. Report the code and
@@ -81,6 +81,59 @@ Write failures (disk full `ENOSPC`, permission denied `EACCES`) do not hang the
81
81
  output streams fail with a clear `Failed to write "<path>"` message telling the user to check disk
82
82
  space and directory permissions. Fix the environment, then re-run into a fresh output directory.
83
83
 
84
+ ## File-level data-quality severity
85
+
86
+ The classes above are per row or per field. The user also needs one verdict for the file as a whole,
87
+ and its signals arrive scattered across the manifest and stderr. Grade them before reporting so the
88
+ same file gets the same verdict no matter who reports it.
89
+
90
+ | Severity | Signal | Handling |
91
+ | --- | --- | --- |
92
+ | Critical | Established with the user to be a cumulative snapshot or an aggregate report (see [ue-routing.md](ue-routing.md)); a required identity or time column empty for the whole file | Do not upload. Resolve with the user first |
93
+ | High | `manifest.output.invalid_records` is a large share of the row count; inspect reported `leading_title_rows` or `header_signal` and the user has not yet said whether those rows are a title/banner; inspect reported `xlsx_structure.merged_covered_cells` and the user has not yet said whether the merged label belongs on the rows the block covers; `summary_rows` was reported and the user has not yet said whether those rows are totals or real records; `duplicate_keys` was reported and the user has not yet said whether the repeated rows are separate observations | Upload is allowed, but name the item explicitly in the confirmation gate and in the completion response |
94
+ | Medium | `manifest.output.skipped_fields`; ragged delimited rows; `manifest.output.flatten_misses`; `manifest.output.unreadable_cells`; `xlsx_structure.hidden_rows` or `hidden_columns` read as data | Report the counts; the gate is unchanged |
95
+ | Low | `manifest.output.lan_ip_records` | Report once |
96
+
97
+ Rules:
98
+
99
+ - Severity is reported, never silently applied. A Critical finding stops the pipeline and is stated
100
+ as a data finding, not as a program error.
101
+ - Grade only what the run actually emitted. Every signal that belongs at High or Critical above now
102
+ has a detector, so grade it from what the run reported — never tell the user the tool checked
103
+ something it did not.
104
+ - `unreadable_cells` counts XLSX cells that carried no value the tool may use, grouped by cause:
105
+ `formula_no_cached_value` (the file stores a formula but not the result Excel last computed),
106
+ `error_value` (`#N/A`, `#DIV/0!`, …), `unreadable_object` (an unrecognized cell shape). The cells
107
+ read as missing and their rows are kept, so the record count says nothing about them — a column
108
+ that is empty in AE while the spreadsheet looks full is this. The tool never evaluates a formula
109
+ and never guesses a result; the fix is to recalculate and re-export in Excel, or to export values
110
+ instead of formulas. `inspect` reports the same counts before conversion.
111
+ - `xlsx_structure` records worksheet layout the rows themselves cannot carry: `merged_covered_cells`
112
+ counts cells that are empty only because a merged block covers them (Excel shows the value on the
113
+ block's first row), `hidden_rows` and `hidden_columns` name what the worksheet hides. The default
114
+ read changes none of it, so these counts describe what was uploaded: a column mostly missing in AE
115
+ while the spreadsheet looks full is the first of them. `merged_cells_filled` and
116
+ `excluded_hidden_rows` say what the run did about it, which is only ever what the user asked for
117
+ via `--fill-merged-cells` / `--exclude-hidden-rows` (mapping: `fill_merged_cells` /
118
+ `exclude_hidden_rows`). XLSX only — a legacy `.xls` workbook is not scanned.
119
+ - `summary_rows` names rows that read as a summary line rather than an observation, by `row` (the
120
+ data-row ordinal) and `signals`: `total_label` (a cell reads as `合计` / `总计` / `小计` / `汇总` / `Total` /
121
+ `Subtotal`) and `column_total` (a number equal to the total of its column's other rows). The rows
122
+ were converted like any other, so this is a finding about what was uploaded: one fabricated event
123
+ whose amount is the whole group's, and a column whose reported `sum` is twice its real total. There
124
+ is no flag that drops a data row, because a row labelled `合计` is sometimes a real record; the fix
125
+ is to remove it from the source file or re-export without it. Any format, `.xls` included.
126
+ - `duplicate_keys` names rows the source repeated under the same business key, by `key_columns` (the
127
+ columns compared), `duplicate_groups`, `extra_rows` (surplus records an upload would carry), and
128
+ `groups` with `count`, the data-row `rows`, and a `key_hash` prefix — never the key's own values.
129
+ Nothing was removed: a repeat is sometimes a real pair of records, two order lines in the same
130
+ checkout second, and AE appends accepted events with no way to un-send one, so ask the user whether
131
+ the rows are separate observations; if not, have them remove the rows from the source file. Values
132
+ are compared as written, so a repeat spelled two ways is missed, and `tracking_truncated` means
133
+ distinct keys outran the scan's budget and there may be more. Any format, `.xls` included.
134
+ - A large quarantine share has no fixed threshold. State the ratio and the dominant error `code`,
135
+ and let the user judge.
136
+
84
137
  ## Cross-cutting rules
85
138
 
86
139
  - Classify first, act second. Match on `code`, not on message text.
@@ -1,5 +1,7 @@
1
1
  # Local analysis
2
2
 
3
+ Dimension / dictionary data (a stable-entity lookup with a join key) is not local-analysis material — it routes to ae-metadata as a dimension table. See [ue-routing.md](ue-routing.md) and [dimension-routing.md](dimension-routing.md).
4
+
3
5
  Keep the source on the local machine. Generated scripts and reports belong under `.ae-cli/data-integration/runs/<run-id>/` with restrictive permissions.
4
6
  Set the directory to `0700` and generated scripts/reports to `0600`.
5
7
 
@@ -10,6 +10,8 @@ Confirm:
10
10
 
11
11
  Accept one or more CSV, TSV, TXT, JSON, JSONL (NDJSON), XLS, or XLSX files. CSV/TSV/TXT/JSON/JSONL/XLSX have no hard size limit; files over 1 GB print a stderr warning with an estimated processing time and suggest splitting. XLS over 100 MB prints a memory-risk warning (the legacy parser loads the whole workbook, roughly 5-10x file size); XLS over 1 GB is still rejected — convert it to XLSX or split it first.
12
12
 
13
+ XLSX worksheet recognition matches `<sheet>` and namespaced `<x:sheet>` alike: some cleaning/export tools rewrite the default OOXML namespace as a prefix, and the workbook reads the same either way. The worksheet *row* nodes themselves are still read unprefixed, so a workbook whose row data is also namespaced falls back to CSV.
14
+
13
15
  ## Inspect without exposing raw values
14
16
 
15
17
  Before the full inspection (which streams and profiles the entire file and can take
@@ -39,16 +41,30 @@ If `selection_required=true`, show only the Sheet/JSON Path candidates and ask t
39
41
  ae-cli data-integration inspect --input-file '<path>' --data-set '<candidate-id>' --source-timezone '<iana-timezone>'
40
42
  ```
41
43
 
42
- Report row/column counts, field types, missing/unique/time-parse ratios, UE eligibility, mapping confidence, and warnings. Samples are bounded (up to 5 distinct, truncated) — summarize, never paste them. ID-like columns (`id`, `*_id`, `*_key`, `*_code`, `*_no`, `*_num`) stay `string` even when every value is numeric; JSON-encoded object/array values inside CSV cells are recognized as `object`/`list`, not `string`. IP- and UUID-named columns are additionally checked against their value specs: inspect warns how many non-empty values are invalid IPv4/IPv6, private/LAN IPs, or non-UUID strings, so the user can decide whether to map them as `ip_field`/`uuid_field`. Read [UE routing](ue-routing.md) before choosing a branch.
44
+ Report row/column counts, field types, missing/unique/time-parse ratios, UE eligibility, mapping confidence, and warnings. Samples are bounded (up to 5 distinct, truncated) — summarize, never paste them. ID-like columns (`id`, `*_id`, `*_key`, `*_code`, `*_no`, `*_num`) stay `string` even when every value is numeric; JSON-encoded object/array values inside CSV cells are recognized as `object`/`list`, not `string`. IP- and UUID-named columns are additionally checked against their value specs: inspect warns how many non-empty values are invalid IPv4/IPv6, private/LAN IPs, or non-UUID strings, so the user can decide whether to map them as `ip_field`/`uuid_field`. Excel columns whose cells carry a date number format infer as `datetime` and are named in a warning (see **Excel date cells** below). Read [UE routing](ue-routing.md) before choosing a branch.
43
45
 
44
46
  A stderr `Warning: … column count different from the header row …` means the CSV/TSV has ragged rows (extra fields dropped, missing fields treated as empty); report it as a data-quality signal. For how every pipeline failure — abnormal data, parse errors, and program errors — is classified and handled, see [error handling](error-handling.md).
45
47
 
48
+ **Value frequency and numeric distribution.** A distinct count says how many different values a column holds, not whether they are worth uploading. Two per-column fields answer that, and both appear in `inspect` output only — never in the convert manifest:
49
+
50
+ - `value_frequency` — the 10 most frequent values as `{value, count, ratio}`, values truncated the way `samples` are and counted after truncation. It is reported only for a column whose distinct values all fit the tracked budget (200), so the counts are exact and complete when present; a column past the budget reports nothing here and its `unique_count` is the field to read instead. Use it to separate an enum from free text: a `渠道` column with three values is a property worth uploading, and its listed values are also what a `value_mapping` decision is made from. A single value covering every row usually means an export artifact, not data — propose `exclude_columns` and let the user decide.
51
+ - `numeric_summary` — `count`, `min`, `max`, `sum`, `mean`, `p25`, `median`, `p75` for a column that inferred as `number`. `count` is the values that read as numbers, which is below the column's non-missing count when the column is mixed. `count`/`min`/`max`/`sum`/`mean` are always exact; the quantiles come from a bounded sample on large columns and then `quantiles_approximate: true` says so. Use it to check the magnitude before it is locked into an AE property: a mean far below the maximum on a monotonically climbing column is the signature of a cumulative snapshot rather than a per-row measure (see [UE routing](ue-routing.md)), an amount whose values are 100× the expected size is a 分/元 unit mismatch, and a column that is entirely one number carries no signal.
52
+
53
+ Summarize both — report the shape of the distribution and the names of the values, and do not paste the whole table into the conversation. They are read out of the customer's file like `samples` are.
54
+
46
55
  ## Advanced input
47
56
 
48
57
  - **Headerless files** — inspect auto-detects a missing header row on CSV/TSV and reports `no_headers: true` with a `header_detection` verdict and `auto_headers` placeholders (`col_1..col_N`); the first row is already treated as data. `--headerless` forces the same behavior without detection. Never keep the `col_1..col_N` placeholders — they carry no business meaning. For each column, read its bounded samples and inferred type and propose a meaningful name, present every proposal to the user (column position, sample summary, suggested name), and let the user confirm or rename each one; record the confirmed names in the mapping's `headers` field. When the user already knows the names, re-run inspect with `--headers 'col1,col2,...'` so the recommended mapping carries them.
58
+ - **Title rows above the header** — an exported report often puts a caption in the first cell (`2026年3月销售明细`) and the real header row underneath. Read as-is, the caption becomes the file's only column name, every real column name is lost, and the header row is counted as a data row. Inspect reports the suspected rows under `leading_title_rows` (row ordinals and non-empty cell counts only, never the cell text) and warns — but it does **not** change what it read: the first row was still used as the header, because Excel exports legitimately carry numeric header rows (`2024`, `2025`) that look like data, and there is no flag that puts a header row back once it has been treated as data. So the report is a question for the user. When they confirm those rows are a title or banner, re-run inspect with `--skip-rows N` (N is exactly the last ordinal listed); the rerun reads the real header row, reports `skipped_rows`, and carries `skip_rows` into the recommended mapping so `convert` reads the same rows inspect profiled. On XLSX, inspect additionally reports `header_signal` when the row it used as the header looks like data — same rule: reported, not applied; resolve it with `--headers` or `--headerless`. This covers CSV/TSV and XLSX; a legacy `.xls` workbook accepts `--skip-rows` but is not scanned for title rows.
59
+ - **Summary and total rows** — an exported report ends with a 合计 row, and a grouped one repeats 小计 after every group. Those rows are not observations: uploaded, each becomes an event that never happened whose amount is the whole group's revenue, and profiled, it doubles the column's `sum` and turns its `max` into the total. Nothing in the row itself says so, so inspect flags them under `summary_rows` and warns. Each entry carries `row` (the data-row ordinal, the same numbering `invalid.rows.jsonl` and `--salvage-from` use) and `signals`: `total_label` when a cell reads as a total label (`合计` / `总计` / `小计` / `汇总` as a prefix, `Total` / `Subtotal` / `Sum` as the whole cell), naming the column in `label_column` but never the cell text; `column_total` when a number on that row equals the total of its column's other rows, listing those columns in `total_columns`. A row can raise one signal or both — a labelled group subtotal holds its group's total, not the column's, so only the label fires. Nothing is removed and nothing is changed: `convert` writes these rows as records, and every number reported for their columns counts them, which is exactly why the finding has to be read. There is no flag that drops a data row, because a row labelled `合计` is sometimes a real business record; when the user confirms a row is a total, ask them to remove it from the source file or re-export without it, then inspect again. `convert` repeats the finding in `manifest.output.summary_rows`, since by upload time a subtotal row with a plausible identity and time is indistinguishable from data. This covers every format, `.xls` included — the check reads rows, not worksheet structure.
60
+ - **Repeated business keys** — a customer re-exports a report whose range overlaps the last export, or pastes two sheets together, and the same observation arrives twice. Uploaded, each repeat is a second event: that user's revenue doubles, every funnel counts them twice, and AE appends accepted events with no way to un-send one — so the only place this is fixable is before the upload. Inspect compares each row's business key against the rows before it and reports repeats under `duplicate_keys`, plus a warning. The report carries `key_columns` (the columns actually compared — always read it, since the key is what the finding means), `checked_rows`, `duplicate_groups`, `extra_rows` (surplus records an upload would carry), and `groups`, each with `count`, the data-row `rows` (the numbering `invalid.rows.jsonl` and `--salvage-from` use), and a `key_hash` prefix that distinguishes groups without revealing values — the key's own text is never reported. `groups_truncated` / `rows_truncated` mean the list is bounded, not that the counts are; `tracking_truncated` means distinct keys outran the scan's budget and there may be more. The key comes from the mapping's identity, time, and event-name columns on `convert`, and from column-name matching on `inspect`; a single column is never a key, so a file with no recognizable time column is not scanned at all (identity alone would call every returning user's second row a repeat). Two limits to state when reporting: values are compared as written, so `2026-03-01 10:00:00` and `2026/03/01 10:00:00` are two different keys and a repeat spelled two ways is missed; and a source's own unique key (an order id) is not compared on unless the mapping names it as identity, time, or event. Nothing is removed — a repeat is sometimes a real pair of records, two order lines in the same checkout second — so ask the user whether the rows are separate observations, and if they are not, have them remove the rows from the source file and inspect again. `convert` repeats the finding in `manifest.output.duplicate_keys`, because by upload time both copies are ordinary valid records; it describes the whole source file even on a `--salvage-from` run, since every valid row of that file ends up in AE. This covers every format, `.xls` included — the check reads rows, not worksheet structure.
49
61
  - **TSV / TXT** — `.tsv` and `.tab` use a tab delimiter with no quoting convention; `.txt` and unknown extensions are content-sniffed into CSV, TSV, or NDJSON.
50
62
  - **Encoding** — text files are auto-detected (UTF-8, GBK, GB2312, Big5, and others); no flag is needed.
51
- - **Excel sheets** — `--merge-sheets` streams every worksheet in file order instead of a single selected sheet; otherwise ask which sheet/`--data-set` to use. Inspect also reports `header_consistency` (`all_same` or `different`) across a workbook's sheets, with `header_details` listing each sheet's header row when they differ; prefer `--merge-sheets` only when headers match.
63
+ - **Excel date cells** — a cell whose number format is a date or date+time is read as the wall-clock timestamp shown in Excel, not as the Excel serial number stored behind it, so the column infers as `datetime` and can serve as the time field. Inspect lists every such column in a warning. Treat that warning as a question to the user, not as a note: the same column profiled as `number` before this behavior existed, so if any part of this file was already sent to AE, the property may have been received as a number and its type is now locked — it cannot be changed to datetime, and the column has to be re-sent under a new property name. Ask whether the column was uploaded before, and only map it once the user answers. Elapsed-duration formats (`[h]:mm:ss` and the equivalent built-ins) are durations rather than points in time and stay `number`.
64
+ - **Excel formula cells** — a spreadsheet stores a formula and, next to it, the result Excel last computed. That cached result is the value: it is read normally, including a result of `0` or `""`, which are real values and not blanks. This tool never evaluates a formula and never guesses a result, so a cell holding a formula the file never computed has nothing to upload; it is read as missing and counted, as is an Excel error value (`#N/A`, `#DIV/0!`, …). Inspect reports the counts per column in a warning and `convert` repeats them in `manifest.output.unreadable_cells`. Report them: the rows are kept and the record count is unchanged, so this is the only explanation for a column that is empty in AE while the spreadsheet looks full. When a column that matters reads as missing this way, ask the user to recalculate and re-export in Excel, or to export values instead of formulas, before uploading. This covers XLSX; a legacy `.xls` workbook goes through a different parser and is not counted here.
65
+ - **Merged cells, hidden rows, and hidden columns** — a sheet maintained by hand merges a label down the rows it covers (`区域` spanning one region's block). Excel keeps that value on the block's first row only and stores every row below it as an empty cell, so a column that looks full on screen arrives mostly missing, and the AE property built from it would be empty for most events. The same worksheet may also hide a row inside a data block or hide a whole column. None of this travels with a row, so inspect scans the worksheet structure separately and reports it under `xlsx_structure`: `merged_ranges` with `merged_range_samples` (references such as `A3:A5`, never cell text), `merged_covered_cells` per column, `hidden_rows` with `hidden_row_samples` (source row numbers as Excel numbers them), and `hidden_columns` by header name. The default read is unchanged, so the report is a question for the user, and each answer is a flag: `--fill-merged-cells` copies each block's value into the cells its own range covers — bounded to the range, never overwriting a value that is there and never inventing one when the block's own cell is empty, so it is not a forward fill; `--exclude-hidden-rows` leaves hidden rows out. Neither is on by default: those cells really are empty in the file, and a row hidden inside a data block may still be real data — unlike a hidden *worksheet* (below), which is excluded by default. Hidden columns have no flag at all; when the user confirms one is not data, list it in the mapping's `exclude_columns`. Both flags are carried into the recommended mapping as `fill_merged_cells` / `exclude_hidden_rows`, which is what makes `convert` read the rows inspect profiled — `convert` has no read flags of its own — and `convert` repeats the findings in `manifest.output.xlsx_structure`, the only record of a layout the converted rows no longer show. This covers XLSX; a legacy `.xls` workbook is not scanned, so ask the user about merged labels and hidden rows there instead of trusting silence.
66
+ - **Hidden worksheets** — a worksheet hidden in the workbook is left out of the `--data-set` candidates and out of `--merge-sheets`, because a sheet the file does not show is usually scratch space or a superseded draft rather than rows anyone meant to upload, but a hidden sheet can also be a dimension / dictionary table that is meant to be loaded, and that goes through dimension routing instead of being dismissed as scratch. Inspect lists each one under `excluded_sheets` (with `reason: hidden`); report those names to the user, since they are the only explanation for a row count lower than the workbook appears to hold. Their headers are also left out of `header_consistency`, so a stale hidden draft cannot make a mergeable workbook look ragged. A hidden sheet stays readable when the user names it in `--data-set` — the command then warns on stderr that the selected sheet is hidden. Only pass a hidden sheet after the user says that is what they want. When *every* worksheet is hidden there is no candidate left, and inspect fails with `LOCAL_DATA_ALL_DATA_SETS_HIDDEN` whose hint lists the hidden sheets; treat that as a question about which sheet holds the real data, not as an unreadable file. This detection covers XLSX only: a legacy `.xls` workbook's sheet list is unfiltered, so a hidden sheet there still appears as a candidate and is still merged — for `.xls`, ask the user to confirm the sheet list instead of trusting it.
67
+ - **Excel sheets** — `--merge-sheets` streams every visible worksheet in file order instead of a single selected sheet; otherwise ask which sheet/`--data-set` to use. Inspect also reports `header_consistency` (`all_same` or `different`) across a workbook's sheets, with `header_details` listing each sheet's header row when they differ; prefer `--merge-sheets` only when headers match. Matching headers establish a shared structure, not disjoint rows: a detail sheet and a summary sheet, or `1月` and `1月修订版`, usually carry identical headers and would be merged and reported twice over. Before merging, confirm with the user that the sources are mutually exclusive partitions (one month per sheet, no overlap) rather than overlapping, revised, or derived views of the same rows, and show each sheet's row count and time coverage range in that confirmation so an overlap is visible. The same rule applies to repeated `--input-file`.
52
68
  - **Multi-file type conflicts** — when the same column has different inferred types across files, present each conflict and resolve with `--type-resolutions` on `convert` (see [transform](transform.md)).
53
69
 
54
70
  ## Nested flattening (NDJSON/JSON records and JSON-encoded CSV/TSV/Excel cells)
@@ -10,7 +10,7 @@ Generate and confirm the event/property plan **before** any transform or upload.
10
10
  ## Sub-steps
11
11
 
12
12
  1. **Event-model decision** — reuse the UE routing result: single-table single-event `track`, single-table multi-event (event-name column), `user_set`, or `mixed`. The agent may propose splitting one table into several events (for example an ad table into `ad_show`/`ad_click` by `campaign_type`); that proposal must be confirmed by the user in the confirmation gate.
13
- 2. **Column → property draft** — confirm the recommended mapping's key system fields with the user **before** drafting: `mode` (`#type`), `#account_id`/`#distinct_id` (ask together; at least one is required — a `user_id` column can be either an anonymous or a login ID and only the user knows), `#time` + source timezone + `#zone_offset`, `#event_name` (track only; the event column or a reviewed `default_event_name`), and `#ip`/`#uuid` when the data has such a column. The exact questions and the never-infer-from-a-column-name-alone rule are [transform.md](transform.md) steps 1–5; run them here. The plan is generated from this mapping, so never draft from an unconfirmed mapping. With the system fields settled, map the remaining columns to event and/or user properties (the mapping `mode` decides). Common event properties — project-level super properties attached to every event — are defined by `ae-generate-tracking-plan`, not this import path. Name events and properties in snake_case and fill **every** `display_name`, `desc`, and `event_tag` (events also carry `event_desc`). Infer all three from field names, value distribution, samples, and business-doc / prompt priors — never leave them empty: `desc`/`event_desc` state what the item means in plain language (language follows the user), and `event_tag` picks the closest category from the canonical tag list (see the `event_tag` appendix in `../../ae-generate-tracking-plan/references/business-dimension-mapping.md`). When you cannot infer a `desc` or `event_tag`, mark it pending and ask the user for it inside the confirmation gate. Infer types (`number` / `bool` / `datetime` / enum) the same way; CSV defaults to `string`. Columns that stay uncertain or conflicting are marked pending and asked only inside the confirmation gate.
13
+ 2. **Column → property draft** — confirm the recommended mapping's key system fields with the user **before** drafting: `mode` (`#type`), `#account_id`/`#distinct_id` (ask together; at least one is required — a `user_id` column can be either an anonymous or a login ID and only the user knows), `#time` + source timezone + `#zone_offset`, `#event_name` (track only; the event column or a reviewed `default_event_name`), and `#ip`/`#uuid` when the data has such a column. The exact questions and the never-infer-from-a-column-name-alone rule are [transform.md](transform.md) steps 1–5; run them here. The plan is generated from this mapping, so never draft from an unconfirmed mapping. With the system fields settled, map the remaining columns to event and/or user properties (the mapping `mode` decides). Common event properties — project-level super properties attached to every event — are defined by `ae-generate-tracking-plan`, not this import path. Name events and properties in snake_case by default (keep an uppercase event name only when the user asks to preserve it; property names stay lowercase-only) and fill **every** `display_name`, `desc`, and `event_tag` (events also carry `event_desc`). Infer all three from field names, value distribution, samples, and business-doc / prompt priors — never leave them empty: `desc`/`event_desc` state what the item means in plain language (language follows the user), and `event_tag` picks the closest category from the canonical tag list (see the `event_tag` appendix in `../../ae-generate-tracking-plan/references/business-dimension-mapping.md`). When you cannot infer a `desc` or `event_tag`, mark it pending and ask the user for it inside the confirmation gate. Infer types (`number` / `bool` / `datetime` / enum) the same way; CSV defaults to `string`. Columns that stay uncertain or conflicting are marked pending and asked only inside the confirmation gate.
14
14
  3. **Confirmation gate (single, one pass)** — present the concrete plan, never a counts-only summary: the confirmed key system-field mapping (`mode`, `#account_id`/`#distinct_id`, `#time` + source timezone, `#event_name`; `#ip`/`#uuid` when present), then a full event table (one row per event: `event_name` + `event_tag`/`event_desc` + the properties attached to it), then a full property table (one row per property: source column → target AE name → type → `display_name`/`desc`, uncertain types highlighted; a kept-whole `object`/`array_row` lists its `parent.child` sub-properties next to the parent), plus field scope (default: plan fields only, with a full-import switch) and unrecognized/dirty-data handling. For a multi-sheet workbook, group the property table by sheet so each sheet's source columns are visible. The user answers once with ok or edits (renames, types, identity/time/event columns, add/drop columns).
15
15
  4. **Merge with the existing plan** — fetch the project's current tracking plan; same-name property type conflicts are severe, same-name events are advisory; decide append vs replace. This runs for **every** file, not just the first: when the project already has a plan (an earlier file or run), diff this file's events and properties against it and put the additions — new events, new properties, new object sub-properties from flattening — in the confirmation gate. An existing plan is never a reason to skip this step; only when every addition is already present may you skip the merge, and even then state and confirm that fact with the user.
16
16
  5. **Persist the plan** — `.ae-cli/data-integration/draft.json` → `.ae-cli/data-integration/draft.xlsx` → upload with `sdk_integration_mode=none`.
@@ -17,7 +17,7 @@ Confirm the system fields with the user before touching properties. These are th
17
17
  4. **`#event_name` (track only).** Confirm the event-name column, or a reviewed `default_event_name` when no column exists.
18
18
  5. **`#ip` / `#uuid` (optional).** Ask only when the data has an IP- or UUID-like column; map it via `ip_field`/`uuid_field`, otherwise skip. `#ip` is event data only and must be a valid IPv4/IPv6 address (a private/LAN IP is kept but reported — AE cannot geolocate it); `#uuid` must be a standard 36-character UUID. A value that violates the spec is dropped from that row only (`INVALID_IP` / `INVALID_UUID`) — the row itself is kept. The program never auto-generates a `#uuid`.
19
19
 
20
- When an `#event_name`, `#account_id`, or `#distinct_id` column's values do not satisfy AE naming rules (pure Chinese, uppercase, spaces), do not stop — scan the distinct values, list them to the user, and ask for one AE-name replacement each; record the pairs in `value_mapping` (see [UE mapping](ue-mapping.md)). A value with no matching key keeps its original text and fails validation, so confirm every distinct value is covered or excluded. The same mechanism applies to a property column via that entry's own `value_mapping`.
20
+ When an `#event_name`, `#account_id`, or `#distinct_id` column's values do not satisfy AE naming rules (pure Chinese, spaces), do not stop — scan the distinct values, list them to the user, and ask for one AE-name replacement each; record the pairs in `value_mapping` (see [UE mapping](ue-mapping.md)). Event names are lowercased by default; keep an uppercase event name only when the user asks to preserve it — leave a legal uppercase value unmapped so it passes through, or map an illegal value to an uppercase target (e.g. `购买` → `Purchase`); see the casing rules in [UE mapping](ue-mapping.md). A value with no matching key keeps its original text and fails validation, so confirm every distinct value is covered or excluded. The same mechanism applies to a property column via that entry's own `value_mapping` — property names stay lowercase-only, so uppercase property values still need `value_mapping`.
21
21
 
22
22
  Then confirm the property set with the user before saving the mapping. Present every property as a readable table — one row per property with source column, target AE name, and type — never dump the raw mapping JSON at the user. Whether the rows are event or user properties follows `mode`: `track` → **event properties**, `user_set` (or another profile mode) → **user properties**, `mixed` → the same set applies to both event and user rows.
23
23
 
@@ -48,7 +48,9 @@ ae-cli data-integration convert \
48
48
  --output-dir '.ae-cli/data-integration/runs/<run-id>'
49
49
  ```
50
50
 
51
- The command never modifies the source. Inspect `manifest.json`; summarize valid and quarantined counts and the block reason. If `manifest.output.skipped_fields` is present, tell the user how many `#ip`/`#uuid` values were invalid and dropped (the rows were otherwise kept); if `manifest.output.lan_ip_records` is present, tell the user that many `#ip` values are private/LAN addresses that AE cannot geolocate. Neither blocks the manifest. If `manifest.output.flatten_misses` is present — or a stderr `Warning: flatten rule "X" did not materialize for N row(s).` fires a `flatten_rules` path missed some rows: re-check the dot path against the source shape (or confirm the column is legitimately optional); the rows are otherwise kept. A stderr `Warning: … column count different from the header row …` means ragged CSV/TSV rows were tolerated (extra fields dropped, missing fields treated as empty) — surface it as a data-quality note. A blocked manifest whose reason is `The source contained no data rows.` means the file had zero data rows; do not re-run the same command on it. Do not expose rows from `invalid.rows.jsonl` unless the user specifically asks to inspect the local quarantine. For the full failure taxonomy and how to respond, see [error handling](error-handling.md).
51
+ Merging several sources into one run assumes they are mutually exclusive partitions of the same data set. Confirm that with the user before converting matching headers do not rule out a detail/summary pair or an original/revised pairand present each source's row count and time coverage range so an overlap is visible. The same applies to `--merge-sheets` (see [source-inspect.md](source-inspect.md)).
52
+
53
+ The command never modifies the source. Inspect `manifest.json`; summarize valid and quarantined counts and the block reason, and check the conservation equation first: `manifest.output.source_rows` must equal `valid_records + invalid_records` — a mismatch means rows were dropped or duplicated between the source and the output, so report the three numbers to the user and do not upload until it is explained. A salvage run's `source_rows` is only the rows it re-processed, not the whole file. If `manifest.output.skipped_fields` is present, tell the user how many `#ip`/`#uuid` values were invalid and dropped (the rows were otherwise kept); if `manifest.output.lan_ip_records` is present, tell the user that many `#ip` values are private/LAN addresses that AE cannot geolocate. Neither blocks the manifest. If `manifest.output.flatten_misses` is present — or a stderr `Warning: flatten rule "X" did not materialize for N row(s).` fires — a `flatten_rules` path missed some rows: re-check the dot path against the source shape (or confirm the column is legitimately optional); the rows are otherwise kept. A stderr `Warning: … column count different from the header row …` means ragged CSV/TSV rows were tolerated (extra fields dropped, missing fields treated as empty) — surface it as a data-quality note. A blocked manifest whose reason is `The source contained no data rows.` means the file had zero data rows; do not re-run the same command on it. Do not expose rows from `invalid.rows.jsonl` unless the user specifically asks to inspect the local quarantine. For the full failure taxonomy and how to respond, see [error handling](error-handling.md).
52
54
 
53
55
  ## Re-report only the failed rows (salvage loop)
54
56
 
@@ -66,12 +66,14 @@ Each mapped system field carries a value spec, enforced at both inspect (warning
66
66
  | Field | Value spec | On violation (convert) |
67
67
  | --- | --- | --- |
68
68
  | `#account_id` / `#distinct_id` | Non-empty string, at most 128 characters | Row error `MISSING_USER_ID` (absent) / `USER_ID_TOO_LONG` (>128) |
69
- | `#event_name` | `^[a-z][a-z0-9_]{0,49}$` (lowercase snake_case, letter-leading, ≤50 chars) | Row error `INVALID_EVENT_NAME` |
69
+ | `#event_name` | `^[A-Za-z][A-Za-z0-9_]{0,49}$` (letter-leading letters/digits/underscore, ≤50 chars; lowercase by default, uppercase kept only on user request) | Row error `INVALID_EVENT_NAME` |
70
70
  | `#time` | One of the supported formats, within 3 years back / 3 days forward | Row error `INVALID_TIME` / `TIME_OUT_OF_RANGE` |
71
71
  | `#ip` | Valid IPv4 or IPv6. Event data only | Field skip `INVALID_IP`; a private/LAN IP is kept and reported — AE cannot geolocate it |
72
72
  | `#zone_offset` | Integer -12..14 (or an IANA name for `zone_offset_value`). Event data only | Row error `INVALID_ZONE_OFFSET` |
73
73
  | `#uuid` | Standard 36-character UUID (`xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx`). Both data kinds | Field skip `INVALID_UUID` |
74
74
 
75
+ Event-name case is decided entirely by the mapping, never by the CLI — the CLI performs no case conversion and has no flag for it. The name AE receives is the source `event_name_field` value (or `default_event_name` when there is no event column), optionally replaced by a `value_mapping.event_name` target; `data-integration plan --event-name` passes each name through unchanged. Lowercase is the default: lowercase the name when writing the mapping (a `value_mapping.event_name` entry such as `Purchase` → `purchase`, or a lowercase `default_event_name`). To keep uppercase, skip that lowercasing: leave a legal uppercase source value unmapped so it passes through, write an uppercase `value_mapping.event_name` target only when the source value is not itself a legal name (e.g. `购买` → `Purchase`), or write an uppercase `default_event_name` / `--event-name`.
76
+
75
77
  A field skip is not a row failure: the row is kept with its other fields, and the count is reported in `manifest.output.skipped_fields` so the agent tells the user at the end. `#uuid` is never auto-generated — it comes only from a mapped source column.
76
78
 
77
79
  Mapping keys: `account_id_field`, `distinct_id_field`, `time.field`, `event_name_field` (or a reviewed `default_event_name`), `ip_field`, `uuid_field`. Inspect surfaces every identity-shaped column in `identity_candidates` (name, `account`/`distinct` kind, unique and missing ratios) so the agent can present all candidates for confirmation. When the required identity column is absent, use an explicit `account_id_value`/`distinct_id_value` placeholder or `random_pool` — a user decision, never invented.
@@ -147,7 +149,8 @@ Values with an explicit offset or `Z` are parsed by the JavaScript `Date` constr
147
149
  - Identity values remain strings and are at most 128 characters.
148
150
  - Source timezone is an IANA name derived from user/project context.
149
151
  - `#zone_offset`, when set, is a whole-hour integer in -12..14 (or an IANA name/column that resolves to one) and is emitted inside `properties`, never at the top level.
150
- - Event/property names are lowercase snake_case, begin with a letter, and are at most 50 characters.
152
+ - Event names begin with a letter and are at most 50 characters; lowercase is the default, uppercase is kept only when the user asks to preserve it.
153
+ - Property names are lowercase snake_case (letter-leading, at most 50 characters).
151
154
  - Target property names are unique and do not collide with UE system fields.
152
155
  - Types are one of `string`, `number`, `boolean`, `datetime`, `list`, `object`, or `array_row`.
153
156
  - Text is at most 2 KB; numbers stay within -9E15..9E15.
@@ -20,7 +20,45 @@ Classification order:
20
20
  5. Rows that mix track and user-profile facts in one file use `mixed` with a `record_type_field`; require explicit review.
21
21
  6. Low-confidence output is a proposal, never silent approval.
22
22
 
23
- Aggregated metrics, pivot tables, cross-tabs, model outputs, free-form documents, and records without real identity/time should normally use local analysis.
23
+ Aggregated metrics, pivot tables, cross-tabs, model outputs, and free-form documents should normally use local analysis. Records without real identity/time are checked against dimension routing next and fall to local analysis only if they are not a stable-entity lookup.
24
+
25
+ ### Time coverage is not native granularity
26
+
27
+ A parseable time column establishes only when the rows are stamped, not what period each metric
28
+ covers. A daily report and a cumulative snapshot both look like one row per user per point in time,
29
+ so they satisfy every condition above and are then ingested as per-period events — inflating totals
30
+ in a way that stays invisible in ratios, because numerator and denominator scale together.
31
+
32
+ Native granularity must come from the user, a data dictionary, or a complete period structure in the
33
+ data itself. Do not infer it from the file name, the first/last date, the interval between rows, or
34
+ the row count; none of those is evidence.
35
+
36
+ When the data carries numeric columns and any of the following holds, ask the user to state whether
37
+ each row's value is the amount that occurred in that period or the total accumulated up to that
38
+ point, and do not proceed until they answer:
39
+
40
+ - Paired start/end time columns (`start_date`/`end_date`, `period_begin`/`period_end`).
41
+ - Values for one identity that never decrease over time.
42
+ - Column names carrying a to-date sense (`cumulative`, `total`, `ltv`, `累计`, `总`).
43
+
44
+ An unanswered question, a cumulative snapshot, or overlapping periods route to local analysis
45
+ instead.
46
+
47
+ ## Route to dimension table
48
+
49
+ A dimension / dictionary table describes stable entities (city, product, device): a lookup that maps an entity code to its attributes. It has no row-level identity or event time, so it fails the UE prerequisites above, but it is not local-analysis material either — it belongs in AE as a dimension table bound to a property.
50
+
51
+ Classification order (UE first, dimension second, local analysis last):
52
+
53
+ 1. Satisfy the UE must-holds above → UE ingestion wins; never route an identity/time-bearing file here.
54
+ 2. Fail the UE prerequisites **and** match most of these dimension signals → dimension routing:
55
+ - No row-level identity: no `account_id` / `distinct_id` column. A `code` / `id` / `no` key is an entity code, not a user identity.
56
+ - No row-level event time: no `#time` column. If time exists, it is an effective / expiry interval, not an event occurrence.
57
+ - Finite enumeration: few rows, each describing one entity's attributes (code → name / level), not facts accumulating over time.
58
+ - A join key: a column shared with event data (`city_code`, `sku_id`, `device_id`) whose values are descriptive attributes, not measures.
59
+ 3. Otherwise → local analysis.
60
+
61
+ See [references/dimension-routing.md](references/dimension-routing.md) for the handoff.
24
62
 
25
63
  ## Route to local analysis
26
64
 
@@ -29,6 +67,7 @@ Choose local analysis when:
29
67
  - The user wants insights, not project ingestion.
30
68
  - UE identity or time prerequisites are missing.
31
69
  - Each row is an aggregate rather than a user/event record.
70
+ - The rows are a cumulative snapshot, or their native granularity could not be established.
32
71
  - Conversion would invent semantics or discard important structure.
33
72
  - The user declines an uncertain mapping or destination.
34
73