@moda-ai/cli 1.34.1 → 1.36.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,15 +1,15 @@
1
1
  import {
2
2
  selectTenantAndCreateKey
3
- } from "./cli-4ycvzfdg.js";
3
+ } from "./cli-scjxgaba.js";
4
4
  import {
5
5
  loadAuthSession
6
- } from "./cli-eggpzt87.js";
6
+ } from "./cli-yz99t7n8.js";
7
7
  import {
8
8
  CliAuthError,
9
9
  resolveIngestUrl,
10
10
  resolveModaBaseUrl,
11
11
  stringOption
12
- } from "./cli-0h32yvwq.js";
12
+ } from "./cli-drabjmaj.js";
13
13
 
14
14
  // src/provision.ts
15
15
  async function runProvision(context, options = {}) {
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@moda-ai/cli",
3
- "version": "1.34.1",
3
+ "version": "1.36.0",
4
4
  "description": "CLI for Moda - AI agent analytics and observability",
5
5
  "type": "module",
6
6
  "bin": {
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "schema_version": "moda.skill_index.v1",
3
- "bundled_at": "2026-09-22T23:28:01.590Z",
4
- "cli_version": "1.34.1",
3
+ "bundled_at": "2026-09-27T08:58:32.181Z",
4
+ "cli_version": "1.36.0",
5
5
  "skills": [
6
6
  {
7
7
  "id": "integration-cloudflare-think",
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  name: moda-cli
3
- version: 2.5.0
4
- description: Query Moda's AI agent trace analytics and manage code-first prompt versions from the terminal — semantic/keyword/hybrid message search, overview KPIs, topic clusters, message context, user frustration detections, tool failures, and moda prompts status/sync/promote. Use when the user asks about moda, modaflows, trace or conversation analytics, prompt management, user frustrations, agent observability, tool failure debugging, wants to find traces or tool calls about a topic, or wants to investigate how their AI agent is performing.
3
+ version: 2.7.0
4
+ description: Query Moda's AI agent trace analytics and manage code-first prompt versions from the terminal — semantic/keyword/hybrid message search, overview KPIs, topic clusters, message context, user frustration detections, tool failures, custom signals, false-positive review of detections, and moda prompts status/sync/promote. Use when the user asks about moda, modaflows, trace or conversation analytics, prompt management, user frustrations, agent observability, tool failure debugging, custom signals, marking a detection as a false positive, wants to find traces or tool calls about a topic, or wants to investigate how their AI agent is performing.
5
5
  ---
6
6
 
7
7
  # Moda CLI
@@ -58,9 +58,9 @@ Optional:
58
58
 
59
59
  - `MODA_BASE_URL` defaults to `https://moda.dev`; override only for
60
60
  self-hosted or staging.
61
- - `MODA_SKILL_VERSION` — export to `2.5.0` so search-adoption telemetry can
61
+ - `MODA_SKILL_VERSION` — export to `2.6.0` so search-adoption telemetry can
62
62
  attribute usage to this skill version. Set it once per session:
63
- `export MODA_SKILL_VERSION=2.5.0`.
63
+ `export MODA_SKILL_VERSION=2.6.0`.
64
64
 
65
65
  ## Setup
66
66
 
@@ -748,6 +748,82 @@ Rules that keep the proof honest:
748
748
  `moda fix <fix_id> --wait`) drives it; a fix left `GATING` will not finish
749
749
  on its own.
750
750
 
751
+ ### 9. Custom signals and false positives
752
+
753
+ **Mark anything as a false positive with one command.** `--reason` (why the
754
+ detection is wrong) is optional but recommended; marks are logged for Moda's review.
755
+ A marked item disappears from every list and count (`moda problems`,
756
+ `hallucinations`, `laziness`, `tool-failures`, problem evidence, the dashboard);
757
+ `--undo` brings it back.
758
+
759
+ ```bash
760
+ moda false-positive problem <problem_id> --reason="two unrelated tools" # dismiss: hidden at once
761
+ moda false-positive problem-detection <problem_id> --conversation-id=<id> \
762
+ --attribution-id=<uuid> --reason="refund was approved" # one piece of evidence
763
+ moda false-positive quote <conversation_id> --family=frustration --turn=3 \
764
+ --quote="this is useless lol" --reason="user was joking" # one user quote
765
+ moda false-positive tool <conversation_id> --tool=send_invoice \
766
+ --tool-use-id=<tool_use_id> --reason="retried upstream" # one failing tool call
767
+ moda false-positive laziness <conversation_id> --reason="reload is expected"
768
+ moda false-positive hallucination <conversation_id> --reason="price came from the tool result"
769
+ moda false-positive custom-signal <signal_id> --conversation-id=<id> --reason="refund was approved"
770
+ moda false-positive laziness <conversation_id> --undo # remove a mark (any kind but problem)
771
+ moda false-positives # every current mark (--kind=tool, --all, --limit)
772
+ moda false-positives <conversation_id> # marks + problem reviews on a trace
773
+ ```
774
+
775
+ What each flag does downstream:
776
+
777
+ | Kind | Effect |
778
+ |---|---|
779
+ | `problem` | Dismisses the problem: it leaves `moda problems` and the dashboard at once, and the reason is folded into the next reconcile. Cannot be undone from the CLI |
780
+ | `problem-detection` | Drops that evidence from the problem and teaches the problem reader (same as `detection-review --verdict=false_positive`) |
781
+ | `custom-signal` | Saved as a `does_not_fit` label with the reason, so the signal learns from it and the live match is retracted (same as `signal-label`). `--undo` removes the label |
782
+ | `quote`, `tool`, `laziness`, `hallucination` | Audit only: recorded for review, never fed back to the detectors |
783
+
784
+ Find targets with the read commands: `--family`, `--turn` and `--quote` verbatim
785
+ from a `user_quotes` entry in `moda emotions`; a failing call's `tool_use_id` (or `msg_index` when it has none) from
786
+ `moda tool-failure-detail <tool_name>`;
787
+ laziness traces from `moda laziness`; hallucination traces from
788
+ `moda hallucinations`; custom-signal hits from `moda trace-signals` or
789
+ `moda signal <id> --detections`; problem evidence ids from
790
+ `moda problem <id> --evidence`.
791
+
792
+ **Custom signals** are tenant-defined behaviors (pass/fail criteria) judged on
793
+ every trace. Read them, then label hits so the signal learns:
794
+
795
+ ```bash
796
+ moda signals # every signal: criteria, status, 7-day judged/matched
797
+ moda signal <signal_id> # one signal's summary
798
+ moda signal <signal_id> --detections --days-back=30 # recent hits (limit 1–20, --offset to page)
799
+ moda trace-signals <conversation_id> # signal hits on one trace
800
+ moda signal-label <signal_id> --conversation-id=<id> --label=does_not_fit \
801
+ --note="refund was approved" --msg-index=<anchor_msg_index>
802
+ ```
803
+
804
+ `--label=does_not_fit` marks a hit a **false positive** and retracts the live
805
+ match; `--label=fits` confirms it. Pass the hit's `anchor_msg_index` (or
806
+ `--msg-dedup-token=<anchor_dedup_token>`) so the example points at the right
807
+ message, and always give a `--note` when rejecting. Labels are throttled to 10
808
+ per minute per API key and are never retried automatically.
809
+
810
+ **Problem detections** (the dashboard's flag on problem evidence) are reviewed with
811
+ `detection-review`. A `false_positive` drops the detection out of
812
+ `moda problem <id> --evidence` and is fed back to the problem reader; `clear`
813
+ undoes an earlier review.
814
+
815
+ ```bash
816
+ moda detection-reviews <conversation_id> # problem detections on a trace + review state
817
+ moda detection-review <problem_id> --conversation-id=<id> \
818
+ --attribution-id=<uuid> --verdict=false_positive --note="retry was expected"
819
+ moda detection-review <problem_id> --conversation-id=<id> \
820
+ --attribution-id=<uuid> --verdict=clear # undo
821
+ ```
822
+
823
+ `attribution_id` comes from `moda detection-reviews <conversation_id>` or
824
+ `moda problem <problem_id> --evidence`. Prefer `detection-review` over
825
+ `problem-feedback --action=flag_attribution` for a single wrong detection.
826
+
751
827
  ## Common workflow recipes
752
828
 
753
829
  ### Find where something happened (search → context)
@@ -917,7 +993,16 @@ npx: `npx -p @moda-ai/cli moda <command>`.
917
993
  | `moda tool-failure-detail <tool_name>` | Per-tool failure breakdown + examples |
918
994
  | `moda problems` | Rank cross-signal Problems by root cause (what to fix first) |
919
995
  | `moda problem <problem_id>` | One Problem: dossier, or `--evidence`/`--reports`/`--traces`/`--feedback` pages (`--conversations` is a legacy alias of `--traces`) (`--limit` 1–50, `--cursor` verbatim keyset token) |
920
- | `moda problem-feedback <problem_id>` | Write: `--action=mark_fixed\|dismiss\|flag_attribution\|rename`. `--reason` required for dismiss/flag_attribution; `--new-name` for rename; `--attribution-id` (UUID) required for flag_attribution |
996
+ | `moda problem-feedback <problem_id>` | Write: `--action=mark_fixed\|dismiss\|flag_attribution\|rename`. `--reason` required for flag_attribution, optional for dismiss; `--new-name` for rename; `--attribution-id` (UUID) required for flag_attribution |
997
+ | `moda false-positive <kind> <target>` | Write: kinds `problem\|problem-detection\|quote\|tool\|laziness\|hallucination\|custom-signal`; `--reason` optional; problem-detection needs `--conversation-id` + `--attribution-id`; quote needs `--family`, `--turn`, `--quote`; tool needs `--tool` plus `--tool-use-id` or `--msg-index`; custom-signal needs `--conversation-id`; `--undo` removes a mark (not for problem) |
998
+ | `moda false-positives [conversation_id]` | Current false-positive marks: all (`--kind`, `--all` to include removed, `--limit`) or one trace with its problem detection reviews |
999
+ | `moda laziness` | Laziness detections per trace (`--days-back`, `--limit` 1-50, `--conversation-id`, `--family`) |
1000
+ | `moda detection-reviews <conversation_id>` | Problem detections on a trace with review state (`--attribution-id=UUID` narrows to one) |
1001
+ | `moda detection-review <problem_id>` | Write: `--conversation-id`, `--attribution-id` (UUID), `--verdict=correct\|false_positive\|clear`, optional `--note` |
1002
+ | `moda signals` | Custom signals: criteria, status, 7-day judged/matched stats |
1003
+ | `moda signal <signal_id>` | One custom signal; `--detections` pages recent hits (`--days-back` 1–90, `--limit` 1–20, `--offset`) |
1004
+ | `moda trace-signals <conversation_id>` | Custom signal hits on one trace |
1005
+ | `moda signal-label <signal_id>` | Write: `--conversation-id`, `--label=fits\|does_not_fit` (false positive), optional `--note`, `--msg-index`/`--msg-dedup-token` |
921
1006
  | `moda step-scores <conversation_id>` | Graph-PRM step scores: per-segment curves, first bad step, rollup |
922
1007
  | `moda world-state <conversation_id>` | Agent memory: slots/threads/events; `--summary-only`; `--snapshot --msg-index=N` (state at a turn); `--replay --message-count=N` (state over time) |
923
1008
  | `moda failures` | Production failures worth fixing first |