periplus-python-sdk 0.11.0__py3-none-any.whl → 0.13.0__py3-none-any.whl

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: periplus-python-sdk
3
- Version: 0.11.0
3
+ Version: 0.13.0
4
4
  Summary: Typed Periplus platform client with SQL and notebook integration
5
5
  License-Expression: AGPL-3.0-only
6
6
  Project-URL: Repository, https://github.com/elei-io/periplus
@@ -21,8 +21,9 @@ Dynamic: license-file
21
21
  # Periplus Python SDK
22
22
 
23
23
  Typed organization operations and SQL access through the Periplus HTTP API.
24
- The 0.11.0 platform interface requires the matching API release. It includes
25
- the source-snapshot metadata introduced in 0.10.0 alongside customer operations.
24
+ The 0.12.0 platform interface requires the matching API release. It replaces the
25
+ discovery, retention and monitoring namespaces with captures, crawls, pins and
26
+ schedules, and includes the source-snapshot metadata introduced in 0.10.0.
26
27
  Install from PyPI:
27
28
 
28
29
  ```sh
@@ -46,7 +47,7 @@ Use HTTPS outside loopback development. Connect directly to the API origin,
46
47
  not the public marketing site.
47
48
 
48
49
  Version 0.9.0 requires personal API keys instead of database credentials. Create
49
- a key in the app's Settings → My API keys. It inherits your current access in
50
+ a Developer key in the app's Access → SDK & API. It inherits your current access in
50
51
  that organization; SQL requires your membership to have `organization:sql:exec`.
51
52
  For local development, install `./clients/periplus-python-sdk` from the repository
52
53
  root and connect to `http://localhost:8000`.
@@ -55,9 +56,12 @@ root and connect to `http://localhost:8000`.
55
56
  returns typed columns/rows for read-only queries. `schema()` returns
56
57
  visible tables, column types/descriptions and helper documentation. All SQL uses
57
58
  `POST /api/v1/sql`; schema discovery uses `GET /api/v1/schema`. ClickHouse enforces
58
- permissions. The SDK never retries automatically, including failed queries.
59
+ permissions. Buffered SQL reads retry HTTP 502, 503 and 504 up to twice with bounded
60
+ backoff starting in 0.12.1. Set `sql_retries=0` to disable this, or choose up to three retries. Each
61
+ attempt may observe a newer corpus snapshot. Mutations, transport failures and
62
+ streamed queries are not retried automatically.
59
63
 
60
- Public HTML joins use `parse_id` and `node_index`; `document_id` identifies exact
64
+ Public HTML joins use `capture_id` and `node_index`; `document_id` identifies exact
61
65
  raw bytes. Public shorthand uses the `public_v1` schema.
62
66
 
63
67
  In 0.10.0, `result.source_snapshot` is a typed `SourceSnapshot` containing
@@ -96,15 +100,16 @@ See [the public schema](../../docs/SCHEMA.md) and [query boundary](../../docs/QU
96
100
 
97
101
  ## Organization resources
98
102
 
99
- The same key selects your organization and inherits your live membership access.
103
+ A Developer key selects your organization and uses your live membership access. Agent keys authenticate only to the hosted MCP endpoint, not this SDK.
100
104
  `Client` and `AsyncClient` expose the same namespaces:
101
105
 
102
106
  | Namespace | Operations |
103
107
  | --- | --- |
104
108
  | `identity`, `availability` | `get` |
105
- | `discovery` | `create`, `get`, `list`, `iter`, `arrivals`, `cancel` |
106
- | `retention`, `monitoring` | `create`, `preview`, `activate`, `get`, `list`, `iter`, `members`, `update`, `pause`, `resume`, `delete` |
107
- | `saved_queries` | `create`, `get`, `list`, `iter`, `rename`, `delete` |
109
+ | `captures`, `crawls` | `create`, `get`, `list`, `iter`, `pages`, `cancel` |
110
+ | `pins` | `create`, `get`, `list`, `iter`, `captures`, `update`, `delete` |
111
+ | `schedules` | `create`, `get`, `list`, `iter`, `pages`, `captures`, `update`, `pause`, `resume`, `delete` |
112
+ | `saved_queries` | `create`, `get`, `list`, `iter`, `update`, `rename`, `delete` |
108
113
  | `members` | `list`, `update_role`, `remove` |
109
114
  | `invitations` | `list`, `create`, `cancel` |
110
115
  | `api_keys` | `create`, `list`, `iter`, `revoke` |
@@ -112,30 +117,61 @@ The same key selects your organization and inherits your live membership access.
112
117
  | `query_history` | `list`, `iter`, `get`, `summary` |
113
118
  | `audit` | `list`, `iter` |
114
119
 
120
+ Choose the resource by what you know and how long you need it:
121
+
122
+ | Resource | Use it when |
123
+ | --- | --- |
124
+ | `captures` | You know the URLs (1–100) and need them in SQL now |
125
+ | `crawls` | You know a starting point, not the URLs |
126
+ | `pins` | You need exact captures for longer than seven days |
127
+ | `schedules` | You need the same pages again later |
128
+
115
129
  ```python
116
130
  from periplus_sdk import Client
117
131
 
118
132
  with Client() as client:
119
- preview = client.retention.preview(
133
+ capture = client.captures.create(urls=["https://example.com/pricing"])
134
+ capture = client.captures.get(capture, wait_seconds=30)
135
+ if capture.finished and capture.queryable:
136
+ for page in client.captures.pages(capture).items:
137
+ print(page.url, page.status, page.capture_id)
138
+
139
+ crawl = client.crawls.create(seeds=["https://example.com/docs/"], page_limit=200,
140
+ allowed_paths=["/docs/*"])
141
+
142
+ pin = client.pins.create(
120
143
  name="Research sources", days=90,
121
144
  sql="SELECT capture_id FROM captures WHERE domain(url) = ?",
122
145
  parameters=["example.com"],
123
146
  )
124
- print(preview.id, preview.capture_count, preview.sample)
125
- # Review the selection before calling:
126
- # active = client.retention.activate(preview)
147
+ schedule = client.schedules.create(name="Pricing", every_days=7,
148
+ urls=["https://example.com/pricing"])
149
+ print(pin.captures, schedule.projected_pages_per_month)
127
150
  ```
128
151
 
129
- SQL previews persist inactive, frozen selections; activation never reruns SQL.
130
- Explicit IDs/URLs passed to `create` activate immediately. Monitoring URLs must
131
- already be known in the corpus. Use `members` and its `next_cursor` for complete
132
- membership: policy detail contains a sample. Updates/deletion accept a fetched
133
- `Policy` or an ID with `expected_version`; conflicts are never retried.
152
+ Every capture is kept for seven days after collection unless a pin keeps it. A
153
+ capture always fetches again and never returns an existing capture, so check
154
+ coverage with SQL first. `request_id` (a capture or crawl `id`) is what you
155
+ ordered; each page's `capture_id` is what the crawler produced and is the SQL key.
156
+ `get(..., wait_seconds=N)` returns as soon as the request is finished and queryable,
157
+ or after at most 30 seconds; the SDK never polls on its own.
158
+
159
+ Pins and schedules accept explicit IDs/URLs or one read-only SQL query with
160
+ parameters, and are active immediately. SQL runs once at creation and its exact
161
+ result is locked in; running the same query beforehand is only an estimate.
162
+ Schedule URLs need not be in the corpus. `pin_days` on a capture, crawl or schedule
163
+ pins each successful capture it produces and requires `organization:pins:write`.
164
+ `pins.update` changes `days` (moving every capture's expiry by the difference) or
165
+ the name; `schedules.update` changes `every_days` or the name. Both accept a fetched
166
+ object or an ID with `expected_version`; conflicts are never retried. Deletion needs
167
+ no version: deleting a pin removes future protection, not captures.
134
168
 
135
169
  Paginated lists return `Page[T]` (`items`, `next_cursor`). Pass cursors unchanged
136
170
  or use `iter()`; membership/invitation lists are bounded snapshots instead.
137
- Offset-based lists can shift during concurrent changes. Query history and arrivals
138
- use keyset cursors. UTC usage ranges have an exclusive end date, at most 93 days.
171
+ Offset-based lists can shift during concurrent changes. Query history, request
172
+ page feeds and scheduled capture feeds use opaque keyset cursors. UTC usage ranges have an exclusive end date, at most 93 days. Usage reports pages
173
+ captured (broken down by capture, crawl or schedule), challenge-resolution pages and
174
+ pinned capture-days.
139
175
  Query history is best-effort and expires after 30 days; it is not a billing ledger.
140
176
 
141
177
  `ApiError` includes HTTP status, code, optional request ID, validation fields and
@@ -0,0 +1,22 @@
1
+ periplus_python_sdk-0.13.0.dist-info/licenses/licensing/LICENSE,sha256=DZak_2itbUtvHzD3E7GNUYSRK6jdOJ-GqncQ2weavLA,34523
2
+ periplus_python_sdk-0.13.0.dist-info/licenses/licensing/NOTICE,sha256=W4W3pzZe-a0ebw4MxgWVmFGj7NSGL1qj8hFx4YG9iF0,105
3
+ periplus_python_sdk-0.13.0.dist-info/licenses/licensing/README.md,sha256=NABTC2wm2Y9iRNMnlMeyVMA2dsAw7XRmCrsV1B2TQdk,1996
4
+ periplus_python_sdk-0.13.0.dist-info/licenses/licensing/THIRD_PARTY_NOTICES.md,sha256=-egdv-dVsKhyRE2pPaHZgg4XYdpQfpH0VwM-sHYjUTg,9643
5
+ periplus_sdk/__init__.py,sha256=QzgQOvgN2YBTr6Vf65uCwKzAPjjnxsq4kdXy2XWxRAs,1020
6
+ periplus_sdk/async_resources.py,sha256=APJ2S-5OsjLwkOzdl3N76yZXtCq8F49EiNlW4kZrRaw,20583
7
+ periplus_sdk/async_stream.py,sha256=O5hO_w_-ooQuDV1fOZRX4Qei1-w7Wri8eIlsn3yCpGQ,4592
8
+ periplus_sdk/client.py,sha256=zVt1uso8LXQuwjVjOz3Nm9PH8bCAv1ZwLbuAvSj5BAI,15592
9
+ periplus_sdk/dbapi.py,sha256=dkJA-DkrMtcgMGjLQ5Vv8-pb6riRrAaG4Yqrn22uRZY,12064
10
+ periplus_sdk/errors.py,sha256=nnbmRzEu1GEFGeKNsLqEo6pHCFu0oFy__5t-yBLO2hw,965
11
+ periplus_sdk/py.typed,sha256=AbpHGcgLb-kRsJGnwFEktk7uzpZOCcBY74-YBdrKVGs,1
12
+ periplus_sdk/resources.py,sha256=Ui08saEs6cUgDttGU7R_VBLDQmFUmk5xJ07uNS3Db2E,20178
13
+ periplus_sdk/resources_types.py,sha256=_UBaTWaut8xq9jDcrJKZdswTDFjSi179P7V6LFjOhQU,8949
14
+ periplus_sdk/sql_api.py,sha256=uUWdwMRSZFMNXXy3-yLozcuT6WluvT6BTjYDB0lNrAU,1917
15
+ periplus_sdk/sqlalchemy.py,sha256=XE_OzdDh9OZ8goWWRoKlBzp7Q0FM-eS9-hjH7Ofzy-E,6292
16
+ periplus_sdk/stream.py,sha256=6Ytb6FkP3Y5C4sDA_SDdU0MJHBQdgHde1lNM6YxHdIQ,5560
17
+ periplus_sdk/types.py,sha256=TSaQRjN5vng_yV4hjZV03s5sU_RHL874qJe40-Hb1KQ,3644
18
+ periplus_python_sdk-0.13.0.dist-info/METADATA,sha256=j0tHG6yuYZ5l-Wnn1ZZIzH3x7CMFT7r4Lb40pB4EQgs,9208
19
+ periplus_python_sdk-0.13.0.dist-info/WHEEL,sha256=YVMoNqKzERt-wjUZwJ33xBGAwnFl-4cqbYkTtWa4itE,91
20
+ periplus_python_sdk-0.13.0.dist-info/entry_points.txt,sha256=Pr14L_7AhLinq-4qDxB4awFVubrvR1BEfkaFVRpdCbU,73
21
+ periplus_python_sdk-0.13.0.dist-info/top_level.txt,sha256=o41t5TzwgoxzSmbKoP6olWW1FyEAGWVYTjeoKadBK40,13
22
+ periplus_python_sdk-0.13.0.dist-info/RECORD,,
@@ -7,7 +7,7 @@ The Periplus license does not replace these terms.
7
7
 
8
8
  ## Selectolax and Lexbor
9
9
 
10
- The backend pins Selectolax 0.4.11 (MIT) and its bundled Lexbor 3.1.0
10
+ The backend pins Selectolax 0.4.12 (MIT) and its bundled Lexbor 3.1.0
11
11
  (Apache-2.0). The DOM adapter's read-only native layouts follow Lexbor's DOM and
12
12
  HTML interface headers. Selectolax retains its packaged MIT license; Periplus
13
13
  includes Lexbor's license and notice under `licensing/third-party/lexbor/`
periplus_sdk/__init__.py CHANGED
@@ -3,9 +3,11 @@ from .dbapi import connect
3
3
  from .client import AsyncClient, Client
4
4
  from .errors import ApiError, ConfigurationError, PeriplusError, ResponseError, TransportError
5
5
  from .types import Diagnostic, PreparedQuery, QueryHelper, QueryHelpers, QueryResult, SourceSnapshot
6
- from .resources_types import Page, Policy, Discovery, Member, Invitation, ApiKey, CreatedApiKey, Identity, Usage
6
+ from .resources_types import (FollowRule, QueryCondition, Page, Request, RequestPage, Pin, PinCapture, Schedule, SchedulePage, Member,
7
+ Invitation, ApiKey, CreatedApiKey, Identity, Usage)
7
8
 
8
9
  __all__ = ["connect", "AsyncClient", "Client", "ApiError", "ConfigurationError", "PeriplusError",
9
10
  "ResponseError", "TransportError", "Diagnostic", "PreparedQuery", "QueryHelper",
10
- "QueryHelpers", "QueryResult", "Page", "Policy", "Discovery", "Member", "Invitation",
11
+ "QueryHelpers", "QueryResult", "FollowRule", "QueryCondition", "Page", "Request", "RequestPage", "Pin", "PinCapture",
12
+ "Schedule", "SchedulePage", "Member", "Invitation",
11
13
  "ApiKey", "CreatedApiKey", "Identity", "Usage", "SourceSnapshot"]
@@ -1,7 +1,7 @@
1
1
  """Async customer operations; the same contract as resources.py."""
2
2
  from __future__ import annotations
3
3
  from collections.abc import AsyncIterator, Sequence
4
- from datetime import date
4
+ from datetime import date, datetime
5
5
  from typing import Generic, Literal, TypeVar
6
6
  from urllib.parse import quote
7
7
  from uuid import UUID, uuid4
@@ -9,9 +9,9 @@ from uuid import UUID, uuid4
9
9
  from pydantic import JsonValue
10
10
 
11
11
  from .resources_types import (
12
- ApiKey, ArrivalsPage, AuditEvent, Availability, CreatedApiKey, Discovery,
13
- Identity, Invitation, Member, Members, Page, Policy, PolicyMember,
14
- QueryExecution, QueryUsage, Role, SavedQuery, Usage,
12
+ ApiKey, AuditEvent, Availability, CreatedApiKey, Identity, Invitation, Member, Members,
13
+ FollowRule, Page, Pin, PinCapture, QueryExecution, QueryUsage, Request, RequestPagesPage, Role,
14
+ SavedQuery, Schedule, SchedulePage, Usage,
15
15
  )
16
16
 
17
17
 
@@ -33,14 +33,28 @@ def page_size(limit: int) -> int:
33
33
  return limit
34
34
 
35
35
 
36
- def policy_ref(policy: Policy | str | UUID, expected_version: int | None, kind: str):
37
- if isinstance(policy, Policy):
38
- if policy.kind != kind:
39
- raise ValueError("Policy belongs to a different resource")
40
- return policy.id, policy.version if expected_version is None else expected_version
36
+ def versioned_ref(value: Pin | Schedule | str | UUID, expected_version: int | None, model: type):
37
+ if isinstance(value, (Pin, Schedule)):
38
+ if not isinstance(value, model):
39
+ raise ValueError(f"Expected a {model.__name__}")
40
+ return value.id, value.version if expected_version is None else expected_version
41
41
  if expected_version is None:
42
- raise ValueError("An ID requires expected_version; fetch the policy before changing it")
43
- return policy, expected_version
42
+ raise ValueError("An ID requires expected_version; fetch the resource before changing it")
43
+ return value, expected_version
44
+
45
+
46
+ def object_id(value) -> str:
47
+ return path_id(value.id if isinstance(value, (Request, Pin, Schedule)) else value)
48
+
49
+
50
+ def selection(ids_name: str, ids: Sequence | None, sql: str | None, parameters: Sequence[JsonValue]) -> dict:
51
+ if (ids is None) == (sql is None):
52
+ raise ValueError(f"Choose either {ids_name} or sql")
53
+ if sql is None:
54
+ if parameters:
55
+ raise ValueError("parameters require sql")
56
+ return {ids_name: [str(value) for value in ids]}
57
+ return {"sql": sql, "parameters": list(parameters)}
44
58
 
45
59
 
46
60
  def member_ref(member: Member | str, expected_role: Role | None):
@@ -83,133 +97,186 @@ class AvailabilityResource(Resource):
83
97
  return await self._client._resource_request("GET", "access", Availability)
84
98
 
85
99
 
86
- class DiscoveryResource(PagedResource[Discovery]):
87
- async def select_urls(self, sql: str, parameters: Sequence[JsonValue] = ()) -> list[str]:
88
- """Resolve a complete starting-page selection without creating discovery."""
89
- from .selection import starting_urls
90
- return starting_urls(await self._client.execute(sql, parameters))
91
-
92
- async def create(self, *, urls: Sequence[str], link_depth: int = 2, page_limit: int = 10_000,
93
- follow_external_links: bool = False, allowed_hosts: Sequence[str] = (),
94
- excluded_hosts: Sequence[str] = (), allowed_paths: Sequence[str] = (),
95
- excluded_paths: Sequence[str] = (), retain_days: int | None = None,
96
- id: UUID | None = None) -> Discovery:
97
- return await self._client._resource_request("POST", "discoveries", Discovery, json={
98
- "id": str(id or uuid4()), "specification": {
99
- "seed_urls": list(urls), "max_depth": link_depth, "page_limit": page_limit,
100
- "follow_scope": "linked_sites" if follow_external_links else "starting_sites",
101
- "allowed_hosts": list(allowed_hosts), "excluded_hosts": list(excluded_hosts),
102
- "allowed_paths": list(allowed_paths), "excluded_paths": list(excluded_paths),
103
- "retain_days": retain_days}})
104
-
105
- async def list(self, *, cursor: str | None = None, limit: int = 20,
106
- status: Literal["active", "paused", "settled"] | None = None) -> Page[Discovery]:
107
- return await self._client._resource_request("GET", "discoveries", Page[Discovery],
108
- params={"offset": page_offset(cursor), "limit": page_size(limit), "status": status})
109
-
110
- async def get(self, id: str | UUID) -> Discovery:
111
- return await self._client._resource_request("GET", f"discoveries/{path_id(id)}", Discovery)
112
-
113
- async def cancel(self, id: str | UUID) -> Discovery:
114
- return await self._client._resource_request("POST", f"discoveries/{path_id(id)}/actions",
115
- Discovery, json={"action": "cancel"})
116
-
117
- async def arrivals(self, id: str | UUID, *, cursor: str | None = None, limit: int = 20) -> ArrivalsPage:
118
- return await self._client._resource_request("GET", f"discoveries/{path_id(id)}/arrivals",
119
- ArrivalsPage, params={"cursor": cursor, "limit": page_size(limit)})
120
-
121
-
122
- class PolicyResource(PagedResource[Policy]):
123
- kind: Literal["retention", "refresh"]
124
-
125
- async def list(self, *, cursor: str | None = None, limit: int = 20,
126
- include_drafts: bool = True) -> Page[Policy]:
127
- return await self._client._resource_request("GET", "page-policies", Page[Policy],
128
- params={"kind": self.kind, "offset": page_offset(cursor), "limit": page_size(limit),
129
- "include_drafts": include_drafts})
130
-
131
- async def get(self, id: str | UUID) -> Policy:
132
- return await self._client._resource_request("GET", f"page-policies/{path_id(id)}", Policy,
133
- params={"kind": self.kind, "members_limit": 20})
134
-
135
- async def members(self, id: str | UUID, *, cursor: str | None = None, limit: int = 100) -> Page[PolicyMember]:
136
- return await self._client._resource_request("GET", f"page-policies/{path_id(id)}/members",
137
- Page[PolicyMember], params={"kind": self.kind, "offset": page_offset(cursor), "limit": page_size(limit)})
138
-
139
- async def _create(self, name, days, selection, retain_days, id) -> Policy:
140
- return await self._client._resource_request("POST", "page-policies", Policy, json={
141
- "id": str(id or uuid4()), "kind": self.kind, "name": name, "days": days,
142
- "selection": selection, "retain_days": retain_days})
143
-
144
- async def _action(self, policy, action, expected_version) -> Policy:
145
- id, version = policy_ref(policy, expected_version, self.kind)
146
- return await self._client._resource_request("POST", f"page-policies/{path_id(id)}/actions", Policy,
147
- params={"kind": self.kind}, json={"action": action, "expected_version": version})
148
-
149
- async def activate(self, policy: Policy | str | UUID, *, expected_version: int | None = None) -> Policy:
150
- """Activate the persisted SQL selection without executing SQL again."""
151
- return await self._action(policy, "activate", expected_version)
152
-
153
- async def pause(self, policy: Policy | str | UUID, *, expected_version: int | None = None) -> Policy:
154
- return await self._action(policy, "pause", expected_version)
155
-
156
- async def resume(self, policy: Policy | str | UUID, *, expected_version: int | None = None) -> Policy:
157
- """Resume without restarting the original retention period."""
158
- return await self._action(policy, "resume", expected_version)
159
-
160
- async def delete(self, policy: Policy | str | UUID, *, expected_version: int | None = None) -> None:
161
- id, version = policy_ref(policy, expected_version, self.kind)
162
- return await self._client._resource_request("DELETE", f"page-policies/{path_id(id)}", None,
163
- params={"kind": self.kind, "expected_version": version})
164
-
165
- async def _update(self, policy, expected_version, fields) -> Policy:
166
- id, version = policy_ref(policy, expected_version, self.kind)
167
- return await self._client._resource_request("PATCH", f"page-policies/{path_id(id)}", Policy,
168
- params={"kind": self.kind}, json={"expected_version": version, **fields})
169
-
170
-
171
- class RetentionResource(PolicyResource):
172
- kind = "retention"
173
-
174
- async def create(self, *, name: str, capture_ids: Sequence[str | UUID], days: int,
175
- id: UUID | None = None) -> Policy:
176
- """Protect explicit capture IDs immediately."""
177
- return await self._create(name, days, {"capture_ids": [str(id) for id in capture_ids]}, None, id)
178
-
179
- async def preview(self, *, name: str, sql: str, days: int,
180
- parameters: Sequence[JsonValue] = (), id: UUID | None = None) -> Policy:
181
- """Save an inactive, frozen SQL selection. Call activate separately."""
182
- return await self._create(name, days, {"sql": sql, "parameters": list(parameters)}, None, id)
183
-
184
- async def update(self, policy: Policy | str | UUID, *, name: str | None = None,
185
- days: int | None = None, expected_version: int | None = None) -> Policy:
186
- return await self._update(policy, expected_version,
187
- {key: value for key, value in {"name": name, "days": days}.items() if value is not None})
188
-
189
-
190
- class MonitoringResource(PolicyResource):
191
- kind = "refresh"
192
-
193
- async def create(self, *, name: str, urls: Sequence[str], interval_days: int,
194
- retain_days: int | None = None, id: UUID | None = None) -> Policy:
195
- """Monitor exact known corpus URLs immediately; does not follow links."""
196
- return await self._create(name, interval_days, {"urls": list(urls)}, retain_days, id)
197
-
198
- async def preview(self, *, name: str, sql: str, interval_days: int,
199
- parameters: Sequence[JsonValue] = (), retain_days: int | None = None,
200
- id: UUID | None = None) -> Policy:
201
- return await self._create(name, interval_days, {"sql": sql, "parameters": list(parameters)}, retain_days, id)
100
+ class RequestResource(PagedResource[Request]):
101
+ """Shared reads for captures and crawls: one request view and one page feed."""
102
+ noun: Literal["captures", "crawls"]
103
+
104
+ async def list(self, *, finished: bool | None = None, cursor: str | None = None,
105
+ limit: int = 20) -> Page[Request]:
106
+ return await self._client._resource_request("GET", self.noun, Page[Request],
107
+ params={"finished": finished, "offset": page_offset(cursor), "limit": page_size(limit)})
108
+
109
+ async def get(self, id: Request | str | UUID, *, wait_seconds: int = 0) -> Request:
110
+ """Return as soon as the request is finished and queryable, or after `wait_seconds` (0-30).
111
+
112
+ Waiting never delays capture. The SDK does not poll; call again to keep waiting."""
113
+ if not 0 <= wait_seconds <= 30:
114
+ raise ValueError("wait_seconds must be between 0 and 30")
115
+ return await self._client._resource_request("GET", f"{self.noun}/{object_id(id)}", Request,
116
+ params={"wait_seconds": wait_seconds or None})
117
+
118
+ async def pages(self, id: Request | str | UUID, *,
119
+ status: Literal["captured", "failed", "pending"] | None = None,
120
+ search: str | None = None, host: str | None = None, depth: int | None = None,
121
+ failure_code: str | None = None, render_incomplete: bool | None = None,
122
+ cursor: str | None = None, limit: int = 20) -> RequestPagesPage:
123
+ """Requested pages, newest first; `capture_id` is the SQL key. Pass cursors unchanged.
124
+
125
+ Failure-code filtering scans a bounded window, so an empty page can still have a cursor.
126
+ `render_incomplete=True` lists captured pages whose HTML was only a loading placeholder."""
127
+ return await self._client._resource_request("GET", f"{self.noun}/{object_id(id)}/pages", RequestPagesPage,
128
+ params={"status": status, "search": search, "host": host, "depth": depth,
129
+ "failure_code": failure_code, "render_incomplete": render_incomplete,
130
+ "cursor": cursor, "limit": page_size(limit)})
131
+
132
+ async def cancel(self, id: Request | str | UUID) -> Request:
133
+ """Stop remaining work. Captures already made are kept."""
134
+ return await self._client._resource_request("POST", f"{self.noun}/{object_id(id)}/cancel", Request)
135
+
136
+
137
+ class CapturesResource(RequestResource):
138
+ noun = "captures"
139
+
140
+ async def create(self, *, urls: Sequence[str], pin_days: int | None = None,
141
+ id: UUID | None = None) -> Request:
142
+ """Fetch 1-100 known URLs once, now.
143
+
144
+ Always fetches again and never returns an existing capture: check coverage with SQL
145
+ first. `pin_days` pins each successful capture (requires organization:pins:write).
146
+ Keep `id` to reconcile an uncertain write; the same ID and input return the same request."""
147
+ return await self._client._resource_request("POST", "captures", Request, json={
148
+ "id": str(id or uuid4()), "urls": list(urls),
149
+ **({"pin_days": pin_days} if pin_days is not None else {})})
150
+
151
+
152
+ class CrawlsResource(RequestResource):
153
+ noun = "crawls"
154
+
155
+ async def admission_preview(self, *, seeds: Sequence[str], accepted_queue_wait_seconds: int | None = None) -> Availability:
156
+ return await self._client._resource_request("POST", "access/crawl-preview", Availability, json={
157
+ "seed_urls": list(seeds), "accepted_queue_wait_seconds": accepted_queue_wait_seconds})
158
+
159
+ async def create(self, *, seeds: Sequence[str], page_limit: int, max_depth: int = 2,
160
+ follow_scope: Literal["starting_sites", "linked_sites"] = "starting_sites",
161
+ allowed_hosts: Sequence[str] = (), excluded_hosts: Sequence[str] = (),
162
+ allowed_paths: Sequence[str] = (), excluded_paths: Sequence[str] = (),
163
+ follow_rules: Sequence[FollowRule | dict] = (),
164
+ accepted_queue_wait_seconds: int | None = None,
165
+ pin_days: int | None = None, id: UUID | None = None) -> Request:
166
+ """Follow links from seeds when the URLs are not known yet.
167
+
168
+ `page_limit` means up to N pages including seeds, not guaranteed coverage. `max_depth`
169
+ is at least 1; use `captures.create` for known URLs. Hosts include subdomains; path
170
+ patterns match the whole URL-encoded path and only `*` is special."""
171
+ return await self._client._resource_request("POST", "crawls", Request, json={
172
+ "id": str(id or uuid4()), "seeds": list(seeds), "page_limit": page_limit,
173
+ "max_depth": max_depth, "follow_scope": follow_scope,
174
+ **({"accepted_queue_wait_seconds": accepted_queue_wait_seconds} if accepted_queue_wait_seconds is not None else {}),
175
+ "allowed_hosts": list(allowed_hosts), "excluded_hosts": list(excluded_hosts),
176
+ "allowed_paths": list(allowed_paths), "excluded_paths": list(excluded_paths),
177
+ "follow_rules": [rule.model_dump(exclude_none=True) if isinstance(rule, FollowRule) else rule
178
+ for rule in follow_rules],
179
+ **({"pin_days": pin_days} if pin_days is not None else {})})
180
+
181
+
182
+ class PinsResource(PagedResource[Pin]):
183
+ async def create(self, *, name: str, days: int, capture_ids: Sequence[str | UUID] | None = None,
184
+ sql: str | None = None, parameters: Sequence[JsonValue] = (),
185
+ saved_query_id: str | UUID | None = None, id: UUID | None = None) -> Pin:
186
+ """Keep exact captures for `days`, active immediately.
187
+
188
+ Pass `capture_ids`, one read-only `sql` query returning `capture_id`, or the
189
+ `saved_query_id` of such a query in your organization. SQL runs once at creation and its
190
+ exact result is locked in; running it beforehand is only an estimate. Membership never
191
+ changes. SQL selection also requires organization:sql:exec; a saved query also
192
+ requires organization:sql:read."""
193
+ if saved_query_id is not None:
194
+ if capture_ids is not None or sql is not None:
195
+ raise ValueError("Choose one of capture_ids, sql or saved_query_id")
196
+ source = {"saved_query_id": str(saved_query_id), "parameters": list(parameters)}
197
+ else:
198
+ source = selection("capture_ids", capture_ids, sql, parameters)
199
+ return await self._client._resource_request("POST", "pins", Pin, json={
200
+ "id": str(id or uuid4()), "name": name, "days": days, **source})
201
+
202
+ async def list(self, *, cursor: str | None = None, limit: int = 20) -> Page[Pin]:
203
+ return await self._client._resource_request("GET", "pins", Page[Pin],
204
+ params={"offset": page_offset(cursor), "limit": page_size(limit)})
202
205
 
203
- async def update(self, policy: Policy | str | UUID, *, name: str | None = None,
204
- interval_days: int | None = None, retain_days: int | None = None,
205
- clear_retention: bool = False, expected_version: int | None = None) -> Policy:
206
- if clear_retention and retain_days is not None:
207
- raise ValueError("Choose retain_days or clear_retention, not both")
208
- fields = {key: value for key, value in {"name": name, "days": interval_days,
209
- "retain_days": retain_days}.items() if value is not None}
210
- if clear_retention:
211
- fields["retain_days"] = None
212
- return await self._update(policy, expected_version, fields)
206
+ async def get(self, id: Pin | str | UUID) -> Pin:
207
+ return await self._client._resource_request("GET", f"pins/{object_id(id)}", Pin)
208
+
209
+ async def captures(self, id: Pin | str | UUID, *, cursor: str | None = None,
210
+ limit: int = 100) -> Page[PinCapture]:
211
+ """Pinned captures with each one's expiry."""
212
+ return await self._client._resource_request("GET", f"pins/{object_id(id)}/captures", Page[PinCapture],
213
+ params={"offset": page_offset(cursor), "limit": page_size(limit)})
214
+
215
+ async def update(self, pin: Pin | str | UUID, *, name: str | None = None, days: int | None = None,
216
+ expected_version: int | None = None) -> Pin:
217
+ """Rename, or change `days`: every capture's expiry moves by the difference."""
218
+ id, version = versioned_ref(pin, expected_version, Pin)
219
+ return await self._client._resource_request("PATCH", f"pins/{path_id(id)}", Pin, json={
220
+ "expected_version": version,
221
+ **{key: value for key, value in {"name": name, "days": days}.items() if value is not None}})
222
+
223
+ async def delete(self, id: Pin | str | UUID) -> None:
224
+ """Remove future protection. Captures and past usage are not deleted."""
225
+ return await self._client._resource_request("DELETE", f"pins/{object_id(id)}", None)
226
+
227
+
228
+ class SchedulesResource(PagedResource[Schedule]):
229
+ async def create(self, *, name: str, every_days: int, urls: Sequence[str] | None = None,
230
+ sql: str | None = None, parameters: Sequence[JsonValue] = (),
231
+ pin_days: int | None = None, id: UUID | None = None) -> Schedule:
232
+ """Capture a fixed URL list every `every_days`, active immediately.
233
+
234
+ Pass `urls` (they need not be in the corpus) or one read-only `sql` query returning a
235
+ `url` column; SQL runs once and its URLs are locked in. Scheduled runs follow no links.
236
+ `pin_days` pins every successful scheduled capture (requires organization:pins:write)."""
237
+ return await self._client._resource_request("POST", "schedules", Schedule, json={
238
+ "id": str(id or uuid4()), "name": name, "every_days": every_days,
239
+ **selection("urls", urls, sql, parameters),
240
+ **({"pin_days": pin_days} if pin_days is not None else {})})
241
+
242
+ async def list(self, *, cursor: str | None = None, limit: int = 20) -> Page[Schedule]:
243
+ return await self._client._resource_request("GET", "schedules", Page[Schedule],
244
+ params={"offset": page_offset(cursor), "limit": page_size(limit)})
245
+
246
+ async def get(self, id: Schedule | str | UUID) -> Schedule:
247
+ return await self._client._resource_request("GET", f"schedules/{object_id(id)}", Schedule)
248
+
249
+ async def pages(self, id: Schedule | str | UUID, *, search: str | None = None,
250
+ cursor: str | None = None, limit: int = 100) -> Page[SchedulePage]:
251
+ """Scheduled URLs with their latest success, failure and next due time."""
252
+ return await self._client._resource_request("GET", f"schedules/{object_id(id)}/pages", Page[SchedulePage],
253
+ params={"search": search, "offset": page_offset(cursor), "limit": page_size(limit)})
254
+
255
+ async def captures(self, id: Schedule | str | UUID, *, status: Literal["captured", "failed"] | None = None,
256
+ search: str | None = None, render_incomplete: bool | None = None,
257
+ cursor: str | None = None, limit: int = 20) -> RequestPagesPage:
258
+ """Results of scheduled runs, newest first, in the request page-feed shape."""
259
+ return await self._client._resource_request("GET", f"schedules/{object_id(id)}/captures", RequestPagesPage,
260
+ params={"status": status, "search": search, "render_incomplete": render_incomplete,
261
+ "cursor": cursor, "limit": page_size(limit)})
262
+
263
+ async def update(self, schedule: Schedule | str | UUID, *, name: str | None = None,
264
+ every_days: int | None = None, expected_version: int | None = None) -> Schedule:
265
+ id, version = versioned_ref(schedule, expected_version, Schedule)
266
+ return await self._client._resource_request("PATCH", f"schedules/{path_id(id)}", Schedule, json={
267
+ "expected_version": version, **{key: value for key, value in
268
+ {"name": name, "every_days": every_days}.items() if value is not None}})
269
+
270
+ async def pause(self, id: Schedule | str | UUID) -> Schedule:
271
+ """Stop new runs; an unstarted run is cancelled, started captures finish."""
272
+ return await self._client._resource_request("POST", f"schedules/{object_id(id)}/pause", Schedule)
273
+
274
+ async def resume(self, id: Schedule | str | UUID) -> Schedule:
275
+ return await self._client._resource_request("POST", f"schedules/{object_id(id)}/resume", Schedule)
276
+
277
+ async def delete(self, id: Schedule | str | UUID) -> None:
278
+ """Remove the schedule. Captures it produced are kept."""
279
+ return await self._client._resource_request("DELETE", f"schedules/{object_id(id)}", None)
213
280
 
214
281
 
215
282
  class SavedQueriesResource(PagedResource[SavedQuery]):
@@ -224,6 +291,13 @@ class SavedQueriesResource(PagedResource[SavedQuery]):
224
291
  return await self._client._resource_request("POST", "saved-queries", SavedQuery,
225
292
  json={"id": str(id or uuid4()), "name": name, "sql": sql})
226
293
 
294
+ async def update(self, id: str | UUID, *, name: str | None = None, sql: str | None = None) -> SavedQuery:
295
+ """Revise the shared definition without executing SQL or changing pinned results."""
296
+ payload = {key: value for key, value in {"name": name, "sql": sql}.items() if value is not None}
297
+ if not payload:
298
+ raise ValueError("Supply name, sql, or both")
299
+ return await self._client._resource_request("PATCH", f"saved-queries/{path_id(id)}", SavedQuery, json=payload)
300
+
227
301
  async def rename(self, id: str | UUID, *, name: str) -> SavedQuery:
228
302
  return await self._client._resource_request("PATCH", f"saved-queries/{path_id(id)}", SavedQuery, json={"name": name})
229
303
 
@@ -251,7 +325,7 @@ class InvitationsResource(Resource):
251
325
  async def list(self) -> list[Invitation]:
252
326
  return (await self._client._resource_request("GET", "members", Members)).invitations
253
327
 
254
- async def create(self, *, email: str, role: Role = "member") -> None:
328
+ async def create(self, *, email: str, role: Role = "developer") -> None:
255
329
  return await self._client._resource_request("POST", "members/invitations", None,
256
330
  json={"email": email, "role": role})
257
331
 
@@ -264,10 +338,11 @@ class ApiKeysResource(PagedResource[ApiKey]):
264
338
  return await self._client._resource_request("GET", "api-keys", Page[ApiKey],
265
339
  params={"offset": page_offset(cursor), "limit": page_size(limit)})
266
340
 
267
- async def create(self, *, name: str, id: UUID | None = None) -> CreatedApiKey:
341
+ async def create(self, *, name: str, id: UUID | None = None, profile: Literal["developer", "agent"] = "developer", expires_at: datetime | None = None) -> CreatedApiKey:
268
342
  """Secret is returned once; access via secret.get_secret_value()."""
269
343
  return await self._client._resource_request("POST", "api-keys", CreatedApiKey,
270
- json={"id": str(id or uuid4()), "name": name})
344
+ json={"id": str(id or uuid4()), "name": name, **({"profile": profile} if profile != "developer" else {}),
345
+ **({"expires_at": expires_at.isoformat()} if expires_at is not None else {})})
271
346
 
272
347
  async def revoke(self, id: str | UUID) -> None:
273
348
  return await self._client._resource_request("DELETE", f"api-keys/{path_id(id)}", None)