ndi-cli 0.7.0__tar.gz → 0.9.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: ndi-cli
3
- Version: 0.7.0
3
+ Version: 0.9.0
4
4
  Summary: NDI platform CLI: document operations, jobs, workspaces, and the agent workspace tools
5
5
  Keywords: ndi,cli,document-intelligence,coding-agents
6
6
  Author: Nace AI
@@ -50,13 +50,14 @@ config location. Workspace-scoped verbs also take `--workspace ID`.
50
50
  | `ndi extract SOURCE -s SCHEMA` | Structured extract (`--validate` checks a schema with no job) |
51
51
  | `ndi split SOURCE --class id:label` | Logical sections |
52
52
  | `ndi classify SOURCE --class id:label` | Labels (refuses `jobid://`) |
53
- | `ndi ground SOURCE --target id=TEXT` | Locate quoted text |
53
+ | `ndi ground SOURCE --target id=TEXT` | Locate quoted text (standalone quotes only; after search, poll `grounding_job_ids`) |
54
54
  | `ndi job ID` / `ndi jobs` / `ndi cancel ID` | Inspect or cancel jobs |
55
55
  | `ndi workspace create\|list\|get\|stats\|delete\|use` | Workspace lifecycle |
56
- | `ndi files upload\|list\|get\|delete` | Workspace files (`--ingest` uploads then queues ingestion) |
56
+ | `ndi files upload\|list\|get\|delete` | Workspace files (a directory walks nested files; `--ingest` then queues ingestion) |
57
57
  | `ndi ingest` | Queue ingestion |
58
- | `ndi deep-search QUERY` / `ndi fact-search QUERY` | Agentic / single-shot search |
59
- | `ndi je-testing QUERY` | Journal-entry testing over a ledger package (`--path` to scope it, `--effort` for the level) |
58
+ | `ndi deep-search QUERY` / `ndi fact-search QUERY` | Agentic / single-shot search; then `ndi job` on each printed `grounding_job_ids` entry |
59
+ | `ndi filtered-search QUERY` | Corpus filter/rank/count over the metadata catalog (`--cursor` pages an earlier run) |
60
+ | `ndi je-testing QUERY` | Journal-entry testing over a ledger package (`--path` to scope it, `--reasoning-effort` for the thinking budget) |
60
61
  | `ndi folder-metadata` / `file-metadata` / `read-file` / `ask-file` / `run-sql` / `hybrid-search` | Read-only workspace tools (the sandbox surface) |
61
62
 
62
63
  ## Sources
@@ -73,8 +74,15 @@ config location. Workspace-scoped verbs also take `--workspace ID`.
73
74
 
74
75
  ```bash
75
76
  ndi parse a.pdf -o id | ndi extract - -s schema.json
77
+ ndi files upload ./corpus --ingest
76
78
  ```
77
79
 
80
+ `ndi files upload DIR` keeps relative paths under `--path` (default: the
81
+ directory name) and uploads in parallel (`-j`, default 8). Unsupported
82
+ files (and `.DS_Store`) are skipped; a warning on stderr lists them when
83
+ the run finishes. `--ingest` then queues ingestion in batches of 1000
84
+ file ids. A single file still requires `--path`.
85
+
78
86
  ## Output
79
87
 
80
88
  Result content goes to **stdout**; status (`job <id> queued`, `saved …`) goes
@@ -82,9 +90,20 @@ to **stderr**. `-o auto|md|json|payload|id` picks the shape. `--json` is an
82
90
  alias for `-o json`. `--save PATH` / `--out-dir DIR` write files. `--async`
83
91
  submits and prints the job id.
84
92
 
93
+ When a Parse result has external complete content, `-o auto` and `-o md` print
94
+ the inline preview. Use `-o json` or `-o payload` to inspect its disclosure and
95
+ authenticated full-content URL. A final-size preview carries
96
+ `content_truncated`; reduced-layout Office parsing carries `parse_fidelity`.
97
+ Documents use `document.content_url`; workbooks use the affected
98
+ `document.spreadsheet.sheets[].content_url`. The CLI does not download these
99
+ URLs automatically.
100
+
85
101
  API failures exit 1. Usage / config errors exit 2. Schema-validation
86
102
  failures also show up to five field constraints, each limited to 500
87
103
  characters; request input and unrelated error-detail fields are omitted.
104
+ Unsupported inputs print `error: 422 [unsupported_file_type] ...` to stderr and
105
+ exit 1. `ndi upload` prints no handle, document operations print no job ID, and
106
+ `ndi files upload --ingest` creates no workspace file or ingestion job.
88
107
 
89
108
  See the [CLI guide](https://docs.ndi.nace.ai/guides/cli) for the full command
90
109
  reference. `SKILL.md` is the agent-facing guide to the six workspace tools.
@@ -26,13 +26,14 @@ config location. Workspace-scoped verbs also take `--workspace ID`.
26
26
  | `ndi extract SOURCE -s SCHEMA` | Structured extract (`--validate` checks a schema with no job) |
27
27
  | `ndi split SOURCE --class id:label` | Logical sections |
28
28
  | `ndi classify SOURCE --class id:label` | Labels (refuses `jobid://`) |
29
- | `ndi ground SOURCE --target id=TEXT` | Locate quoted text |
29
+ | `ndi ground SOURCE --target id=TEXT` | Locate quoted text (standalone quotes only; after search, poll `grounding_job_ids`) |
30
30
  | `ndi job ID` / `ndi jobs` / `ndi cancel ID` | Inspect or cancel jobs |
31
31
  | `ndi workspace create\|list\|get\|stats\|delete\|use` | Workspace lifecycle |
32
- | `ndi files upload\|list\|get\|delete` | Workspace files (`--ingest` uploads then queues ingestion) |
32
+ | `ndi files upload\|list\|get\|delete` | Workspace files (a directory walks nested files; `--ingest` then queues ingestion) |
33
33
  | `ndi ingest` | Queue ingestion |
34
- | `ndi deep-search QUERY` / `ndi fact-search QUERY` | Agentic / single-shot search |
35
- | `ndi je-testing QUERY` | Journal-entry testing over a ledger package (`--path` to scope it, `--effort` for the level) |
34
+ | `ndi deep-search QUERY` / `ndi fact-search QUERY` | Agentic / single-shot search; then `ndi job` on each printed `grounding_job_ids` entry |
35
+ | `ndi filtered-search QUERY` | Corpus filter/rank/count over the metadata catalog (`--cursor` pages an earlier run) |
36
+ | `ndi je-testing QUERY` | Journal-entry testing over a ledger package (`--path` to scope it, `--reasoning-effort` for the thinking budget) |
36
37
  | `ndi folder-metadata` / `file-metadata` / `read-file` / `ask-file` / `run-sql` / `hybrid-search` | Read-only workspace tools (the sandbox surface) |
37
38
 
38
39
  ## Sources
@@ -49,8 +50,15 @@ config location. Workspace-scoped verbs also take `--workspace ID`.
49
50
 
50
51
  ```bash
51
52
  ndi parse a.pdf -o id | ndi extract - -s schema.json
53
+ ndi files upload ./corpus --ingest
52
54
  ```
53
55
 
56
+ `ndi files upload DIR` keeps relative paths under `--path` (default: the
57
+ directory name) and uploads in parallel (`-j`, default 8). Unsupported
58
+ files (and `.DS_Store`) are skipped; a warning on stderr lists them when
59
+ the run finishes. `--ingest` then queues ingestion in batches of 1000
60
+ file ids. A single file still requires `--path`.
61
+
54
62
  ## Output
55
63
 
56
64
  Result content goes to **stdout**; status (`job <id> queued`, `saved …`) goes
@@ -58,9 +66,20 @@ to **stderr**. `-o auto|md|json|payload|id` picks the shape. `--json` is an
58
66
  alias for `-o json`. `--save PATH` / `--out-dir DIR` write files. `--async`
59
67
  submits and prints the job id.
60
68
 
69
+ When a Parse result has external complete content, `-o auto` and `-o md` print
70
+ the inline preview. Use `-o json` or `-o payload` to inspect its disclosure and
71
+ authenticated full-content URL. A final-size preview carries
72
+ `content_truncated`; reduced-layout Office parsing carries `parse_fidelity`.
73
+ Documents use `document.content_url`; workbooks use the affected
74
+ `document.spreadsheet.sheets[].content_url`. The CLI does not download these
75
+ URLs automatically.
76
+
61
77
  API failures exit 1. Usage / config errors exit 2. Schema-validation
62
78
  failures also show up to five field constraints, each limited to 500
63
79
  characters; request input and unrelated error-detail fields are omitted.
80
+ Unsupported inputs print `error: 422 [unsupported_file_type] ...` to stderr and
81
+ exit 1. `ndi upload` prints no handle, document operations print no job ID, and
82
+ `ndi files upload --ingest` creates no workspace file or ingestion job.
64
83
 
65
84
  See the [CLI guide](https://docs.ndi.nace.ai/guides/cli) for the full command
66
85
  reference. `SKILL.md` is the agent-facing guide to the six workspace tools.
@@ -1,6 +1,6 @@
1
1
  [project]
2
2
  name = "ndi-cli"
3
- version = "0.7.0"
3
+ version = "0.9.0"
4
4
  description = "NDI platform CLI: document operations, jobs, workspaces, and the agent workspace tools"
5
5
  readme = "README.md"
6
6
  license = "MIT"
@@ -1,6 +1,6 @@
1
1
  [project]
2
2
  name = "ndi-cli"
3
- version = "0.7.0"
3
+ version = "0.9.0"
4
4
  description = "NDI platform CLI: document operations, jobs, workspaces, and the agent workspace tools"
5
5
  authors = [
6
6
  { name = "Nace AI", email = "engineering@nace.ai" }
@@ -6,4 +6,4 @@ methods. The six workspace tools (``folder-metadata``, ``file-metadata``,
6
6
  in-sandbox surface.
7
7
  """
8
8
 
9
- __version__ = "0.7.0"
9
+ __version__ = "0.9.0"
@@ -36,18 +36,17 @@ def add_parsers(sub: argparse._SubParsersAction, common: argparse.ArgumentParser
36
36
  cancel.set_defaults(run=_run_cancel, family="job", needs_workspace=False, needs_client=True)
37
37
 
38
38
 
39
- def wait_and_report(
39
+ def wait_job(
40
40
  client: NdiClient,
41
41
  job: Job,
42
42
  args: argparse.Namespace,
43
43
  *,
44
- auto: Callable[[Job], str],
45
- stem: str,
46
- ) -> int:
44
+ print_timeout_id: bool = True,
45
+ ) -> tuple[Job | None, int]:
46
+ """Wait for ``job``; status lines go to stderr. Does not render the result."""
47
47
  print(f"job {job.job_id} {job.status}", file=sys.stderr)
48
48
  if getattr(args, "async_submit", False):
49
- print(job.job_id)
50
- return 0
49
+ return job, 0
51
50
  try:
52
51
  if job.is_terminal:
53
52
  check_terminal(job, raise_on_failure=True)
@@ -56,12 +55,30 @@ def wait_and_report(
56
55
  except JobFailedError as exc:
57
56
  code = f" [{exc.job.error.code}]" if exc.job.error and exc.job.error.code else ""
58
57
  print(f"error:{code} {exc}", file=sys.stderr)
59
- return 1
58
+ return exc.job, 1
60
59
  except JobTimeoutError as exc:
61
60
  print(f"error: {exc}", file=sys.stderr)
62
- print(exc.job.job_id)
63
- return 1
61
+ if print_timeout_id:
62
+ print(exc.job.job_id)
63
+ return exc.job, 1
64
64
  print(f"job {job.job_id} {job.status}", file=sys.stderr)
65
+ return job, 0
66
+
67
+
68
+ def wait_and_report(
69
+ client: NdiClient,
70
+ job: Job,
71
+ args: argparse.Namespace,
72
+ *,
73
+ auto: Callable[[Job], str],
74
+ stem: str,
75
+ ) -> int:
76
+ job, code = wait_job(client, job, args)
77
+ if code != 0 or job is None:
78
+ return code or 1
79
+ if getattr(args, "async_submit", False):
80
+ print(job.job_id)
81
+ return 0
65
82
  return render_job(job, args, auto=auto, stem=stem)
66
83
 
67
84
 
@@ -10,6 +10,7 @@ from __future__ import annotations
10
10
 
11
11
  import json
12
12
  import sys
13
+ import uuid
13
14
 
14
15
  import html
15
16
  import re
@@ -21,6 +22,7 @@ from ndi_sdk.models.jobs import (
21
22
  ClassifyResult,
22
23
  DeepSearchV2Result,
23
24
  ExtractResult,
25
+ FilteredSearchResult,
24
26
  GroundResult,
25
27
  IngestionResult,
26
28
  IntelligentSearchResult,
@@ -657,9 +659,21 @@ def ingestion_job(job: Job) -> str:
657
659
  def search_job(job: Job) -> str:
658
660
  result = job.result
659
661
  if isinstance(result, DeepSearchV2Result):
660
- return _search_result(result.answer, result.clarification, result.evidences, receipts=len(result.sql_receipts))
662
+ return _search_result(
663
+ result.answer,
664
+ result.clarification,
665
+ result.evidences,
666
+ receipts=len(result.sql_receipts),
667
+ grounding_job_ids=result.grounding_job_ids,
668
+ )
661
669
  if isinstance(result, IntelligentSearchResult):
662
- return _search_result(result.answer, result.clarification, result.evidences, receipts=0)
670
+ return _search_result(
671
+ result.answer,
672
+ result.clarification,
673
+ result.evidences,
674
+ receipts=0,
675
+ grounding_job_ids=result.grounding_job_ids,
676
+ )
663
677
  return job_line(job)
664
678
 
665
679
 
@@ -703,10 +717,155 @@ def je_testing_job(job: Job) -> str:
703
717
  lines.append(f"quality: {result.quality.status}")
704
718
  if result.exhausted:
705
719
  lines.append("exhausted: step budget ran out before the procedure finished")
720
+ lines.extend(_grounding_job_lines(result.grounding_job_ids))
706
721
  return "\n".join(lines)
707
722
 
708
723
 
709
- def _search_result(answer: str | None, clarification: str | None, evidences: list, *, receipts: int) -> str:
724
+ #: What a cell's non-``found`` status means, glossed for a reader who has only
725
+ #: this text. A field nothing could be read from is not an absent one, and
726
+ #: neither is one that resisted normalizing — keeping them distinct is what
727
+ #: makes a thin row interpretable without opening the JSON.
728
+ _CELL_STATUS = {
729
+ "not_found": "examined; field absent",
730
+ "ambiguous": "present, not normalizable",
731
+ "unexamined": "extraction could not say",
732
+ }
733
+
734
+
735
+ def _cell_value(cell) -> str:
736
+ value = cell.value
737
+ if value is None:
738
+ return _CELL_STATUS.get(cell.status, cell.status)
739
+ match value.kind:
740
+ case "money":
741
+ # Amounts keep their own currency here because the server never
742
+ # converts one either; comparing across currencies is unsupported,
743
+ # so dropping the code would invent a comparison.
744
+ assumed = " (currency assumed)" if value.currency_source == "assumed" else ""
745
+ return f"{value.currency} {value.amount}{assumed}"
746
+ case "text":
747
+ return _clip(value.text)
748
+ case "date":
749
+ return value.date
750
+ case "number":
751
+ return value.number
752
+ case "boolean":
753
+ return "true" if value.value else "false"
754
+ return value.kind
755
+
756
+
757
+ def _page_range(row) -> str:
758
+ if row.page_start is None:
759
+ return ""
760
+ if row.page_end is not None and row.page_end != row.page_start:
761
+ return f" (pages {row.page_start}-{row.page_end})"
762
+ return f" (page {row.page_start})"
763
+
764
+
765
+ def _filtered_rows(result: FilteredSearchResult) -> list[str]:
766
+ """The rows, carrying the columns their membership hinges on.
767
+
768
+ A count-style question confirms ``total_matches`` and returns no rows at
769
+ all, so an empty listing is an answer rather than a gap — it prints nothing
770
+ instead of an empty header.
771
+ """
772
+ if not result.rows:
773
+ return []
774
+ # The server marks which columns its published query filtered or ranked on,
775
+ # and those are the values every row's membership depends on; `requested` is
776
+ # what the reader asked to see. Everything else is context, left for --json.
777
+ hinges = [column.field for column in result.columns if column.filtered]
778
+ shown = hinges + [column.field for column in result.columns if column.requested and column.field not in hinges]
779
+ lines = [f"rows ({len(result.rows)}):"]
780
+ for row in result.rows:
781
+ where = f" [{row.component}]" if row.component else ""
782
+ lines.append(f" {row.source_path}{where}{_page_range(row)}")
783
+ for field in shown or list(dict.fromkeys(cell.field for cell in row.cells)):
784
+ # A repeated column carries one cell per item, so join them rather
785
+ # than showing the first and implying the field held one value.
786
+ values = [_cell_value(cell) for cell in row.cells if cell.field == field]
787
+ if values:
788
+ lines.append(f" {field}: {'; '.join(values)}")
789
+ return lines
790
+
791
+
792
+ def _filtered_caveats(result: FilteredSearchResult) -> list[str]:
793
+ lines: list[str] = []
794
+ if not result.exhaustive:
795
+ lines.append("not exhaustive: the matches are valid, but total_matches is a floor and not a census")
796
+ if result.truncated:
797
+ lines.append("truncated: the published query matched more than this page carries")
798
+ if result.catalog_drifted:
799
+ lines.append("catalog drifted: the underlying data changed since the query first ran")
800
+ # `exhaustive` is scope-wide and almost always false, so it cannot say *what*
801
+ # was incomplete. Only a resolved value satisfies a predicate, so a field the
802
+ # extraction could not read drops its components silently — a correct query
803
+ # coming back thin for a reason unrelated to the question.
804
+ for entry in result.field_readability:
805
+ read = f"{entry.readable} of {entry.in_scope} values read"
806
+ if entry.readable == 0:
807
+ lines.append(f"unreadable: {entry.field} — {read}; a thin result here reflects that, not an absence")
808
+ else:
809
+ lines.append(f"partly readable: {entry.field} — {read}")
810
+ check = result.content_verification
811
+ if check and check.unverified_candidates:
812
+ lines.append(
813
+ f"content check: {check.unverified_candidates} of {check.candidates} candidates unverified "
814
+ f"for {_clip(check.condition, 120)!r} — unknown is not a non-match"
815
+ )
816
+ lines.extend(_coverage_lines(result.coverage))
817
+ if result.exhausted:
818
+ lines.append("exhausted: step budget ran out before the query finished")
819
+ return lines
820
+
821
+
822
+ def filtered_search_job(job: Job) -> str:
823
+ """The totals, then the rows behind them — a catalog answer read at a glance.
824
+
825
+ A filter/rank/count question is answered by a number, so this leads with the
826
+ counts and prints the published query under them: the figure and the
827
+ statement it came out of travel together, the way a JET figure names its
828
+ receipt. Rows carry the columns the matches hinge on; ``-o json`` has every
829
+ cell, citation and column the catalog returned.
830
+ """
831
+ result = job.result
832
+ if not isinstance(result, FilteredSearchResult):
833
+ return job_line(job)
834
+ lines: list[str] = []
835
+ if result.interpretation:
836
+ lines.append(f"interpretation: {_clip(result.interpretation)}")
837
+ files = f" ({result.total_files} files)" if result.total_files is not None else ""
838
+ confidence = f" | confidence: {result.confidence}" if result.confidence else ""
839
+ lines.append(f"matches: {result.total_matches}{files}{confidence}")
840
+ if result.applied_query:
841
+ lines.append(f"query: {_clip(result.applied_query, 300)}")
842
+ lines.extend(_filtered_rows(result))
843
+ lines.extend(_filtered_caveats(result))
844
+ if result.next_cursor:
845
+ lines.append(f"next: {result.next_cursor}")
846
+ return "\n".join(lines)
847
+
848
+
849
+ def _grounding_job_lines(ids: list[uuid.UUID] | None) -> list[str]:
850
+ """Tell the reader which Ground jobs to poll after this search result."""
851
+ if ids is None:
852
+ return []
853
+ if not ids:
854
+ return ["grounding_job_ids: none planned"]
855
+ lines = ["grounding_job_ids (poll each with ndi job <id>):"]
856
+ for item in ids:
857
+ lines.append(f" {item}")
858
+ return lines
859
+
860
+
861
+ def _search_result(
862
+ answer: str | None,
863
+ clarification: str | None,
864
+ evidences: list,
865
+ *,
866
+ receipts: int,
867
+ grounding_job_ids: list[uuid.UUID] | None = None,
868
+ ) -> str:
710
869
  lines: list[str] = []
711
870
  if clarification:
712
871
  lines.append(f"clarification: {clarification}")
@@ -720,4 +879,5 @@ def _search_result(answer: str | None, clarification: str | None, evidences: lis
720
879
  lines.append(f" {evidence.quote}")
721
880
  if receipts:
722
881
  lines.append(f"receipts: {receipts}")
882
+ lines.extend(_grounding_job_lines(grounding_job_ids))
723
883
  return "\n".join(lines)
@@ -9,6 +9,7 @@ from __future__ import annotations
9
9
  import re
10
10
  import sys
11
11
  import uuid
12
+ from dataclasses import dataclass
12
13
  from pathlib import Path
13
14
 
14
15
  from ndi_sdk.client import NdiClient
@@ -49,6 +50,38 @@ SUPPORTED_SUFFIXES = frozenset(
49
50
  ".mov",
50
51
  }
51
52
  )
53
+ # Workspace file upload admits more suffixes than document-ops parse: HTML,
54
+ # plaintext, markdown, email, extra media, and the JSON family that ingest
55
+ # routes even when the console picker omits them.
56
+ WORKSPACE_UPLOAD_SUFFIXES = SUPPORTED_SUFFIXES | {
57
+ ".html",
58
+ ".htm",
59
+ ".txt",
60
+ ".log",
61
+ ".md",
62
+ ".mdx",
63
+ ".eml",
64
+ ".msg",
65
+ ".oft",
66
+ ".xml",
67
+ ".docm",
68
+ ".bmp",
69
+ ".flac",
70
+ ".ogg",
71
+ ".oga",
72
+ ".mpga",
73
+ ".avi",
74
+ ".mkv",
75
+ }
76
+ _SKIP_NAMES = frozenset({".ds_store"})
77
+
78
+
79
+ @dataclass(frozen=True)
80
+ class WorkspaceUploadSelection:
81
+ """Local files a workspace folder upload will send, versus ones it skips."""
82
+
83
+ accepted: list[Path]
84
+ skipped: list[Path]
52
85
 
53
86
 
54
87
  class SourceError(ValueError):
@@ -76,6 +109,42 @@ def collect_specs(spec: str) -> list[str]:
76
109
  return [spec]
77
110
 
78
111
 
112
+ def collect_workspace_uploads(root: Path) -> WorkspaceUploadSelection:
113
+ """Split a directory into uploadable files and ones the workspace API will refuse."""
114
+ children = sorted((child for child in root.rglob("*") if child.is_file()), key=lambda path: str(path))
115
+ accepted: list[Path] = []
116
+ skipped: list[Path] = []
117
+ for child in children:
118
+ if child.name.lower() in _SKIP_NAMES or child.suffix.lower() not in WORKSPACE_UPLOAD_SUFFIXES:
119
+ skipped.append(child)
120
+ else:
121
+ accepted.append(child)
122
+ return WorkspaceUploadSelection(accepted=accepted, skipped=skipped)
123
+
124
+
125
+ def workspace_dest_path(relative: Path, *, prefix: str | None) -> str:
126
+ """Join a workspace prefix with a path relative to the uploaded directory."""
127
+ rel = relative.as_posix().lstrip("/")
128
+ if not prefix or prefix.strip("/") in ("", "."):
129
+ return rel
130
+ return f"{prefix.strip('/')}/{rel}"
131
+
132
+
133
+ def folder_upload_prefix(root: Path, path_arg: str | None) -> str:
134
+ """Workspace destination prefix for a directory upload.
135
+
136
+ ``--path`` wins. Otherwise the directory's name. ``.`` / ``..`` have no
137
+ usable ``Path.name``, so those resolve to the real folder name instead of
138
+ uploading into the workspace root.
139
+ """
140
+ if path_arg is not None:
141
+ return path_arg
142
+ name = root.name
143
+ if name in ("", ".", ".."):
144
+ name = root.resolve().name
145
+ return name
146
+
147
+
79
148
  def resolve_source(
80
149
  spec: str,
81
150
  *,
@@ -3,17 +3,27 @@
3
3
  from __future__ import annotations
4
4
 
5
5
  import argparse
6
+ import json
6
7
  import sys
8
+ from concurrent.futures import ThreadPoolExecutor, as_completed
9
+ from dataclasses import dataclass
7
10
  from pathlib import Path
8
11
 
9
12
  from ndi_sdk.client import NdiClient
10
13
  from ndi_sdk.credentials import write_config
11
- from ndi_sdk.models.jobs import Job
14
+ from ndi_sdk.errors import ErrorCode, JobFailedError, JobTimeoutError, NdiError, NdiStatusError
15
+ from ndi_sdk.models.jobs import Job, UploadFileResult
16
+ from ndi_sdk.resources.jobs import check_terminal
12
17
 
13
18
  from ndi_cli import _render
14
19
  from ndi_cli._config import ConfigError, config_path, read_config_values, resolve_settings
15
- from ndi_cli._jobs import wait_and_report
16
- from ndi_cli._output import add_job_output_args, emit_text, output_mode
20
+ from ndi_cli._jobs import wait_and_report, wait_job
21
+ from ndi_cli._output import add_job_output_args, emit_text, output_mode, _content_failure
22
+ from ndi_cli._sources import collect_workspace_uploads, folder_upload_prefix, workspace_dest_path
23
+
24
+ DEFAULT_UPLOAD_PARALLEL = 8
25
+ INGEST_FILE_ID_BATCH = 1000
26
+ _SKIP_WARNING_LIST_CAP = 20
17
27
 
18
28
 
19
29
  def add_parsers(sub: argparse._SubParsersAction, common: argparse.ArgumentParser) -> None:
@@ -50,11 +60,22 @@ def add_parsers(sub: argparse._SubParsersAction, common: argparse.ArgumentParser
50
60
  files = sub.add_parser("files", help="Upload, list, inspect, or delete workspace files.")
51
61
  fs = files.add_subparsers(dest="files_command", required=True)
52
62
 
53
- f_upload = fs.add_parser("upload", parents=[common], help="Upload a file into a workspace.")
54
- f_upload.add_argument("file")
55
- f_upload.add_argument("--path", required=True, help="Workspace destination path.")
63
+ f_upload = fs.add_parser("upload", parents=[common], help="Upload a file or folder into a workspace.")
64
+ f_upload.add_argument("file", help="Local file, or a directory to walk recursively.")
65
+ f_upload.add_argument(
66
+ "--path",
67
+ default=None,
68
+ help="Workspace destination path (required for a file). For a directory, a prefix; default is the directory name.",
69
+ )
56
70
  f_upload.add_argument("--on-conflict", choices=("reject", "new_version"), default="reject")
57
71
  f_upload.add_argument("--ingest", action="store_true", help="Queue ingestion after the upload.")
72
+ f_upload.add_argument(
73
+ "-j",
74
+ type=int,
75
+ default=DEFAULT_UPLOAD_PARALLEL,
76
+ dest="parallel",
77
+ help=f"Parallel uploads for a directory (default {DEFAULT_UPLOAD_PARALLEL}).",
78
+ )
58
79
  add_job_output_args(f_upload)
59
80
  f_upload.set_defaults(run=_run_files_upload, family="workspace", needs_workspace=True, needs_client=True)
60
81
 
@@ -95,14 +116,30 @@ def add_parsers(sub: argparse._SubParsersAction, common: argparse.ArgumentParser
95
116
  add_job_output_args(fact)
96
117
  fact.set_defaults(run=_run_fact_search, family="workspace", needs_workspace=True, needs_client=True)
97
118
 
119
+ filtered = sub.add_parser(
120
+ "filtered-search", parents=[common], help="Run a filter/rank/count query over the metadata catalog."
121
+ )
122
+ filtered.add_argument("query", nargs="?", default=None)
123
+ filtered.add_argument(
124
+ "--cursor", default=None, help="Page an earlier run: paste its result's next_cursor. A page is a new job."
125
+ )
126
+ filtered.add_argument("--prefix", default=None, dest="path_prefix")
127
+ filtered.add_argument("--context", default=None)
128
+ # No --result-unit: the server deprecates the grouping override and returns
129
+ # catalog entries. No --allow-clarification: a clarification is answered by a
130
+ # new initial query, and a shell has nowhere to keep the thread — the reason
131
+ # deep-search omits --session-id.
132
+ add_job_output_args(filtered)
133
+ filtered.set_defaults(run=_run_filtered_search, family="workspace", needs_workspace=True, needs_client=True)
134
+
98
135
  jet = sub.add_parser("je-testing", parents=[common], help="Run journal-entry testing over a ledger package.")
99
136
  jet.add_argument("query")
100
- jet.add_argument("--effort", choices=("low", "medium", "high"), default=None)
137
+ jet.add_argument("--reasoning-effort", choices=("none", "minimal", "low", "medium", "high"), default=None)
101
138
  jet.add_argument("--prefix", default=None, dest="path_prefix")
102
139
  jet.add_argument("--path", action="append", dest="paths", default=None, help="Scope to this file; repeatable.")
103
140
  jet.add_argument("--context", default=None)
104
141
  # No --session-id, for the reason deep-search omits it: a shell has nowhere
105
- # to keep a thread. Omitting --effort leaves the deployment's own floor.
142
+ # to keep a thread. Omitting --reasoning-effort uses the service default.
106
143
  add_job_output_args(jet)
107
144
  jet.set_defaults(run=_run_je_testing, family="workspace", needs_workspace=True, needs_client=True)
108
145
 
@@ -170,9 +207,14 @@ def _run_ws_use(args: argparse.Namespace) -> int:
170
207
 
171
208
  def _run_files_upload(client: NdiClient, workspace_id: str, args: argparse.Namespace) -> int:
172
209
  path = Path(args.file)
210
+ if path.is_dir():
211
+ return _run_folder_upload(client, workspace_id, args, path)
173
212
  if not path.is_file():
174
213
  print(f"error: file not found: {args.file}", file=sys.stderr)
175
214
  return 2
215
+ if not args.path:
216
+ print("error: --path is required when uploading a single file", file=sys.stderr)
217
+ return 2
176
218
  dest = args.path.lstrip("/")
177
219
  if args.ingest:
178
220
  result = client.upload_and_ingest(
@@ -187,6 +229,165 @@ def _run_files_upload(client: NdiClient, workspace_id: str, args: argparse.Names
187
229
  return _finish_job(client, job, args, _render.job_line, path.stem)
188
230
 
189
231
 
232
+ @dataclass(frozen=True)
233
+ class _UploadOutcome:
234
+ kind: str
235
+ dest: str
236
+ file_id: str | None = None
237
+ error: str | None = None
238
+
239
+
240
+ def _run_folder_upload(client: NdiClient, workspace_id: str, args: argparse.Namespace, root: Path) -> int:
241
+ selection = collect_workspace_uploads(root)
242
+ skipped = [_rel(root, path) for path in selection.skipped]
243
+ if not selection.accepted:
244
+ _print_skip_warning(skipped)
245
+ print(f"error: no supported files under {args.file}", file=sys.stderr)
246
+ return 2
247
+ prefix = folder_upload_prefix(root, args.path)
248
+ workers = max(1, args.parallel)
249
+ outcomes: list[_UploadOutcome] = []
250
+ with ThreadPoolExecutor(max_workers=workers) as pool:
251
+ futures = [
252
+ pool.submit(
253
+ _upload_one,
254
+ client,
255
+ workspace_id,
256
+ local,
257
+ workspace_dest_path(local.relative_to(root), prefix=prefix),
258
+ args,
259
+ )
260
+ for local in selection.accepted
261
+ ]
262
+ for future in as_completed(futures):
263
+ outcomes.append(future.result())
264
+ uploaded = [item for item in outcomes if item.kind == "ok"]
265
+ skipped.extend(item.dest for item in outcomes if item.kind == "skipped")
266
+ failed = [item for item in outcomes if item.kind == "failed"]
267
+ json_mode = output_mode(args) == "json"
268
+ if not json_mode:
269
+ print(f"uploaded {len(uploaded)}")
270
+ for item in sorted(uploaded, key=lambda row: row.dest):
271
+ print(f" {item.dest}")
272
+ for item in failed:
273
+ print(f"error: {item.dest}: {item.error}", file=sys.stderr)
274
+ _print_skip_warning(skipped)
275
+ ingestions: list[Job] | None = None
276
+ code = 1 if failed else 0
277
+ if not failed and args.ingest:
278
+ file_ids = [item.file_id for item in uploaded if item.file_id]
279
+ if not file_ids:
280
+ if json_mode:
281
+ _emit_folder_upload_json(uploaded, skipped, failed, ingestions=[], args=args)
282
+ print("error: no uploaded file ids to ingest", file=sys.stderr)
283
+ return 1
284
+ ingestions, code = _ingest_file_ids(client, workspace_id, args, file_ids, report=not json_mode)
285
+ if json_mode:
286
+ emit_code = _emit_folder_upload_json(uploaded, skipped, failed, ingestions, args)
287
+ return emit_code or code
288
+ return code
289
+
290
+
291
+ def _upload_one(
292
+ client: NdiClient,
293
+ workspace_id: str,
294
+ local: Path,
295
+ dest: str,
296
+ args: argparse.Namespace,
297
+ ) -> _UploadOutcome:
298
+ try:
299
+ job = client.files.upload(workspace_id, local, path=dest, on_conflict=args.on_conflict)
300
+ if not getattr(args, "async_submit", False):
301
+ if job.is_terminal:
302
+ check_terminal(job, raise_on_failure=True)
303
+ else:
304
+ job = client.jobs.wait(job.job_id, timeout=args.timeout)
305
+ return _UploadOutcome(kind="ok", dest=dest, file_id=_uploaded_file_id(job))
306
+ except NdiStatusError as exc:
307
+ if exc.status_code == 422 and exc.code == ErrorCode.UNSUPPORTED_FILE_TYPE:
308
+ return _UploadOutcome(kind="skipped", dest=dest, error=exc.message)
309
+ return _UploadOutcome(kind="failed", dest=dest, error=str(exc))
310
+ except (NdiError, JobFailedError, JobTimeoutError) as exc:
311
+ return _UploadOutcome(kind="failed", dest=dest, error=str(exc))
312
+
313
+
314
+ def _uploaded_file_id(job: Job) -> str | None:
315
+ result = job.result
316
+ if isinstance(result, UploadFileResult):
317
+ return str(result.file.file_id)
318
+ return None
319
+
320
+
321
+ def _emit_folder_upload_json(
322
+ uploaded: list[_UploadOutcome],
323
+ skipped: list[str],
324
+ failed: list[_UploadOutcome],
325
+ ingestions: list[Job] | None,
326
+ args: argparse.Namespace,
327
+ ) -> int:
328
+ payload: dict = {
329
+ "uploaded": [{"path": item.dest, "file_id": item.file_id} for item in sorted(uploaded, key=lambda row: row.dest)],
330
+ "skipped": sorted(skipped),
331
+ "failed": [{"path": item.dest, "error": item.error} for item in failed],
332
+ }
333
+ if ingestions is not None:
334
+ payload["ingestions"] = [json.loads(job.model_dump_json(exclude_none=True)) for job in ingestions]
335
+ return emit_text(json.dumps(payload, indent=2), args, default_stem="upload", suffix=".json")
336
+
337
+
338
+ def _ingest_file_ids(
339
+ client: NdiClient,
340
+ workspace_id: str,
341
+ args: argparse.Namespace,
342
+ file_ids: list[str],
343
+ *,
344
+ report: bool = True,
345
+ ) -> tuple[list[Job], int]:
346
+ jobs: list[Job] = []
347
+ code = 0
348
+ for start in range(0, len(file_ids), INGEST_FILE_ID_BATCH):
349
+ chunk = file_ids[start : start + INGEST_FILE_ID_BATCH]
350
+ try:
351
+ job = client.ingestion.ingest(workspace_id, file_ids=chunk)
352
+ except NdiError as exc:
353
+ if report:
354
+ raise
355
+ print(f"error: {exc}", file=sys.stderr)
356
+ return jobs, 1
357
+ if report:
358
+ chunk_code = _finish_job(client, job, args, _render.ingestion_job, "ingest")
359
+ if chunk_code != 0:
360
+ code = chunk_code
361
+ continue
362
+ waited, wait_code = wait_job(client, job, args, print_timeout_id=False)
363
+ if waited is not None:
364
+ jobs.append(waited)
365
+ code = code or wait_code or _content_failure(waited)
366
+ elif wait_code:
367
+ code = wait_code
368
+ return jobs, code
369
+
370
+
371
+ def _print_skip_warning(skipped: list[str]) -> None:
372
+ if not skipped:
373
+ return
374
+ ordered = sorted(skipped)
375
+ print(f"warning: skipped {len(ordered)} unsupported file(s):", file=sys.stderr)
376
+ shown = ordered[:_SKIP_WARNING_LIST_CAP]
377
+ for name in shown:
378
+ print(f" {name}", file=sys.stderr)
379
+ remaining = len(ordered) - len(shown)
380
+ if remaining:
381
+ print(f" ... and {remaining} more", file=sys.stderr)
382
+
383
+
384
+ def _rel(root: Path, path: Path) -> str:
385
+ try:
386
+ return path.relative_to(root).as_posix()
387
+ except ValueError:
388
+ return str(path)
389
+
390
+
190
391
  def _run_files_list(client: NdiClient, workspace_id: str, args: argparse.Namespace) -> int:
191
392
  page = client.files.list(workspace_id, path_prefix=args.path_prefix, ingestion_status=args.statuses)
192
393
  return _dump_or_text(page, args, text=_render.file_page(page), stem="files")
@@ -237,12 +438,37 @@ def _run_fact_search(client: NdiClient, workspace_id: str, args: argparse.Namesp
237
438
  return _finish_job(client, job, args, _render.search_job, "fact-search")
238
439
 
239
440
 
441
+ def _run_filtered_search(client: NdiClient, workspace_id: str, args: argparse.Namespace) -> int:
442
+ if args.query and args.cursor:
443
+ print("error: pass a query or --cursor, not both", file=sys.stderr)
444
+ return 2
445
+ if not args.query and not args.cursor:
446
+ print("error: pass a query, or --cursor to page an earlier run", file=sys.stderr)
447
+ return 2
448
+ if args.cursor:
449
+ # A page re-executes the parent's own stored query, so the scoping flags
450
+ # cannot apply to it. Accepting and dropping them would quietly answer a
451
+ # different question than the one typed.
452
+ if args.path_prefix or args.context:
453
+ print("error: --prefix and --context belong to the initial query, not a page", file=sys.stderr)
454
+ return 2
455
+ job = client.search.filtered_page(workspace_id, cursor=args.cursor)
456
+ else:
457
+ job = client.search.filtered(
458
+ workspace_id,
459
+ query=args.query,
460
+ context=args.context,
461
+ path_prefix=args.path_prefix,
462
+ )
463
+ return _finish_job(client, job, args, _render.filtered_search_job, "filtered-search")
464
+
465
+
240
466
  def _run_je_testing(client: NdiClient, workspace_id: str, args: argparse.Namespace) -> int:
241
467
  job = client.search.je_testing(
242
468
  workspace_id,
243
469
  query=args.query,
244
470
  context=args.context,
245
- effort=args.effort,
471
+ reasoning_effort=args.reasoning_effort,
246
472
  path_prefix=args.path_prefix,
247
473
  paths=args.paths,
248
474
  )
File without changes
File without changes
File without changes
File without changes
File without changes
File without changes
File without changes
File without changes
File without changes