relaygpu-client 0.1.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (58) hide show
  1. relaygpu_client-0.1.0/.gitignore +12 -0
  2. relaygpu_client-0.1.0/CHANGELOG.md +42 -0
  3. relaygpu_client-0.1.0/LICENSE +21 -0
  4. relaygpu_client-0.1.0/PKG-INFO +567 -0
  5. relaygpu_client-0.1.0/README.md +535 -0
  6. relaygpu_client-0.1.0/pyproject.toml +62 -0
  7. relaygpu_client-0.1.0/relaygpu/__init__.py +21 -0
  8. relaygpu_client-0.1.0/relaygpu/_async/__init__.py +1 -0
  9. relaygpu_client-0.1.0/relaygpu/_async/_http.py +146 -0
  10. relaygpu_client-0.1.0/relaygpu/_async/account.py +83 -0
  11. relaygpu_client-0.1.0/relaygpu/_async/audio.py +86 -0
  12. relaygpu_client-0.1.0/relaygpu/_async/client.py +218 -0
  13. relaygpu_client-0.1.0/relaygpu/_async/estimate.py +18 -0
  14. relaygpu_client-0.1.0/relaygpu/_async/files.py +293 -0
  15. relaygpu_client-0.1.0/relaygpu/_async/image.py +157 -0
  16. relaygpu_client-0.1.0/relaygpu/_async/keys.py +103 -0
  17. relaygpu_client-0.1.0/relaygpu/_async/models.py +75 -0
  18. relaygpu_client-0.1.0/relaygpu/_async/run.py +99 -0
  19. relaygpu_client-0.1.0/relaygpu/_async/tasks.py +66 -0
  20. relaygpu_client-0.1.0/relaygpu/_async/video.py +114 -0
  21. relaygpu_client-0.1.0/relaygpu/_async/webhooks.py +67 -0
  22. relaygpu_client-0.1.0/relaygpu/_async/workflows.py +149 -0
  23. relaygpu_client-0.1.0/relaygpu/_clock.py +31 -0
  24. relaygpu_client-0.1.0/relaygpu/_core.py +208 -0
  25. relaygpu_client-0.1.0/relaygpu/_errors.py +119 -0
  26. relaygpu_client-0.1.0/relaygpu/_estimate.py +224 -0
  27. relaygpu_client-0.1.0/relaygpu/_exceptions.py +166 -0
  28. relaygpu_client-0.1.0/relaygpu/_generated/__init__.py +1 -0
  29. relaygpu_client-0.1.0/relaygpu/_generated/error_codes.py +400 -0
  30. relaygpu_client-0.1.0/relaygpu/_generated/models_map.py +229 -0
  31. relaygpu_client-0.1.0/relaygpu/_generated/types.py +2676 -0
  32. relaygpu_client-0.1.0/relaygpu/_images.py +70 -0
  33. relaygpu_client-0.1.0/relaygpu/_run_common.py +124 -0
  34. relaygpu_client-0.1.0/relaygpu/_shapes.py +85 -0
  35. relaygpu_client-0.1.0/relaygpu/_sync/__init__.py +2 -0
  36. relaygpu_client-0.1.0/relaygpu/_sync/_http.py +147 -0
  37. relaygpu_client-0.1.0/relaygpu/_sync/account.py +84 -0
  38. relaygpu_client-0.1.0/relaygpu/_sync/audio.py +87 -0
  39. relaygpu_client-0.1.0/relaygpu/_sync/client.py +219 -0
  40. relaygpu_client-0.1.0/relaygpu/_sync/estimate.py +19 -0
  41. relaygpu_client-0.1.0/relaygpu/_sync/files.py +294 -0
  42. relaygpu_client-0.1.0/relaygpu/_sync/image.py +158 -0
  43. relaygpu_client-0.1.0/relaygpu/_sync/keys.py +104 -0
  44. relaygpu_client-0.1.0/relaygpu/_sync/models.py +76 -0
  45. relaygpu_client-0.1.0/relaygpu/_sync/run.py +100 -0
  46. relaygpu_client-0.1.0/relaygpu/_sync/tasks.py +67 -0
  47. relaygpu_client-0.1.0/relaygpu/_sync/video.py +115 -0
  48. relaygpu_client-0.1.0/relaygpu/_sync/webhooks.py +68 -0
  49. relaygpu_client-0.1.0/relaygpu/_sync/workflows.py +150 -0
  50. relaygpu_client-0.1.0/relaygpu/_util.py +73 -0
  51. relaygpu_client-0.1.0/relaygpu/_version.py +1 -0
  52. relaygpu_client-0.1.0/relaygpu/async_client.py +5 -0
  53. relaygpu_client-0.1.0/relaygpu/client.py +5 -0
  54. relaygpu_client-0.1.0/relaygpu/errors.py +10 -0
  55. relaygpu_client-0.1.0/relaygpu/inputs.py +306 -0
  56. relaygpu_client-0.1.0/relaygpu/py.typed +0 -0
  57. relaygpu_client-0.1.0/relaygpu/types.py +166 -0
  58. relaygpu_client-0.1.0/relaygpu/webhooks.py +212 -0
@@ -0,0 +1,12 @@
1
+ .venv/
2
+ __pycache__/
3
+ *.pyc
4
+ dist/
5
+ build/
6
+ *.egg-info/
7
+ .env
8
+ .env.*
9
+ .mypy_cache/
10
+ .ruff_cache/
11
+ .pytest_cache/
12
+ .DS_Store
@@ -0,0 +1,42 @@
1
+ # Changelog
2
+
3
+ `relaygpu-client` versions track the TypeScript client [`@relaygpu/client`](https://www.npmjs.com/package/@relaygpu/client)
4
+ minor for minor: 0.1.x here is the port of 0.1.x there. The API may change before 1.0.
5
+
6
+ ## 0.1.0 (unreleased)
7
+
8
+ First release: a port of `@relaygpu/client` 0.1.0.
9
+
10
+ - **Clients**: `Relay` (sync, `httpx.Client`) and `AsyncRelay` (`asyncio`, `httpx.AsyncClient`) over one core;
11
+ `Relay` is generated from `AsyncRelay` (`scripts/unasync.py`), so both expose the same namespaces, methods and
12
+ types. `api_key` (`X-API-Key`) or `jwt` (Bearer); the key is passed explicitly, never read from the environment,
13
+ and never appears in `repr()` or an error. Context managers, `timeout` in seconds (default 600 per attempt),
14
+ `retry` policy (`max_retries` 2, `max_rate_limit_retries` 3, `max_retry_after` 60 s, or `False`),
15
+ `default_headers`, bring-your-own `http_client`.
16
+ - **Any model**: `run()` resolves every model through `models.get` (`GET /v2/models/{name}`, cached 5 min): route,
17
+ body shape, async default and schemas. Unknown → `ModelNotFoundError`, retired → `ModelRetiredError`, both before
18
+ anything is billed. `is_accepted()` narrows the `202` envelope.
19
+ - **Families**: `image.generate` / `image.edit` (→ `ImageResult` of `RelayImage`: `url` / `b64`, `to_bytes()`,
20
+ `save()`), `video.generate` (the `202`, or with `wait=True` the completed task), `audio.speech`,
21
+ `audio.transcribe`.
22
+ - **Tasks**: `tasks.get`, `tasks.wait` (long-polls `?wait=30`; `on_progress` on status transitions;
23
+ `TaskFailedError` carrying the task's `error_code`).
24
+ - **Idempotency and retries**: every async submit and workflow run carries an `Idempotency-Key` (yours or a UUID)
25
+ and is retried safely; replays come back with `replayed: True`. Sync submits are never retried.
26
+ - **Files**: `files.upload` (bytes, path, binary file object, iterator of bytes; streamed), `copy`, `get`, `list`,
27
+ `list_all`, `delete`; implicit upload of file values in any `*_url` / `*_urls` field (`upload={"retention": ...}`,
28
+ default `relay1h`), `inline_images` for images ≤ 4 MB.
29
+ - **Webhooks**: `webhooks.verify` / `verify_webhook` (Standard Webhooks, ±5 min, multi-secret for the 24 h rotation
30
+ window; a plain call on both clients), `secret`, `rotate_secret`, `deliveries.list` / `deliveries.get`; typed
31
+ `task.*`, `workflow.*` and `instance.*` events.
32
+ - **Workflows**: `list`, `get`, `run` (optionally `wait=True`), `get_run`, `wait_run`, `cancel_run`.
33
+ - **Account and keys**: credits, credit history, usage, usage by key, usage timeseries, metrics, pricing, profile,
34
+ model allowlist; key `list` / `list_all` / `get` / `create` / `update` / `rename` / `revoke` / `unrevoke` /
35
+ `delete` / `topup` / `promote`.
36
+ - **Catalog**: `models.list` / `models.get`, `pricing.get`, `tiers.list`, `health()`; `estimate_cost()` from the
37
+ public pricing rows (an estimate, never an invoice).
38
+ - **Errors**: a class per HTTP status (13) and per catalog error code, with `status`, `code`, `request_id`,
39
+ `detail`, `retry_after`; `APIConnectionError`, `APITimeoutError`, `TaskFailedError`, `WebhookVerificationError`.
40
+ - **Types**: `TypedDict`s generated from Relay's public OpenAPI spec (`scripts/gen.py`) in `relaygpu.types`;
41
+ `py.typed`, `mypy --strict` on the package. Python 3.10–3.13; one runtime dependency, `httpx`.
42
+ - `request()`: any route with the same errors and retry policy.
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 OpenGPU Network
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,567 @@
1
+ Metadata-Version: 2.5
2
+ Name: relaygpu-client
3
+ Version: 0.1.0
4
+ Summary: Python client for the Relay API: image, video and audio generation, tasks, files, webhooks, workflows and typed errors.
5
+ Project-URL: Homepage, https://relaygpu.com
6
+ Project-URL: Repository, https://github.com/OpenGPU-Network/relaygpu-python
7
+ Author: OpenGPU Network
8
+ License-Expression: MIT
9
+ License-File: LICENSE
10
+ Keywords: ai,image-generation,inference,relay,relaygpu,sdk,video-generation
11
+ Classifier: Development Status :: 4 - Beta
12
+ Classifier: Intended Audience :: Developers
13
+ Classifier: Operating System :: OS Independent
14
+ Classifier: Programming Language :: Python :: 3
15
+ Classifier: Programming Language :: Python :: 3 :: Only
16
+ Classifier: Programming Language :: Python :: 3.10
17
+ Classifier: Programming Language :: Python :: 3.11
18
+ Classifier: Programming Language :: Python :: 3.12
19
+ Classifier: Programming Language :: Python :: 3.13
20
+ Classifier: Topic :: Software Development :: Libraries :: Python Modules
21
+ Classifier: Typing :: Typed
22
+ Requires-Python: >=3.10
23
+ Requires-Dist: httpx<1,>=0.27
24
+ Provides-Extra: dev
25
+ Requires-Dist: build>=1.2; extra == 'dev'
26
+ Requires-Dist: mypy>=1.11; extra == 'dev'
27
+ Requires-Dist: pytest-asyncio>=0.24; extra == 'dev'
28
+ Requires-Dist: pytest>=8; extra == 'dev'
29
+ Requires-Dist: ruff>=0.6; extra == 'dev'
30
+ Requires-Dist: twine>=5; extra == 'dev'
31
+ Description-Content-Type: text/markdown
32
+
33
+ # relaygpu-client
34
+
35
+ > **Status: pre-release (0.x).** The API may change before 1.0.
36
+
37
+ The Python client for [Relay](https://relaygpu.com): image, video and audio generation, async tasks, file
38
+ uploads, webhook verification, workflows, account and keys, with typed errors. Request and response types are
39
+ generated from Relay's public OpenAPI spec. Chat is not wrapped: point the OpenAI or Anthropic SDK at Relay
40
+ ([quickstart 6](#6-chat-use-the-sdk-you-already-have)). A port of the TypeScript client
41
+ [`@relaygpu/client`](https://www.npmjs.com/package/@relaygpu/client), version for version (0.1.x ↔ 0.1.x).
42
+
43
+ ```bash
44
+ pip install relaygpu-client
45
+ # or
46
+ uv add relaygpu-client
47
+ ```
48
+
49
+ ```python
50
+ import relaygpu
51
+ ```
52
+
53
+ The distribution is `relaygpu-client`; the import is `relaygpu`. Python 3.10–3.13, one runtime dependency
54
+ (`httpx`).
55
+
56
+ ## The client
57
+
58
+ ```python
59
+ import os
60
+
61
+ from relaygpu import Relay
62
+
63
+ relay = Relay(
64
+ os.environ["RELAY_API_KEY"], # relay_sk_…, sent as X-API-Key. Or jwt=... for a dashboard login token.
65
+ base_url="https://relaygpu.com", # the default
66
+ timeout=600, # seconds per HTTP attempt (default 10 min)
67
+ retry={"max_retries": 2, "max_rate_limit_retries": 3, "max_retry_after": 60}, # the defaults; retry=False disables retries
68
+ )
69
+ ```
70
+
71
+ The key is not read from the environment: pass it. It never appears in an error, a log line or `repr(relay)`.
72
+ Catalog reads (`models`, `pricing`, `tiers`, `health`) and task polls need no credential.
73
+
74
+ The client holds a connection pool. Use it as a context manager, or call `relay.close()` when you are done:
75
+
76
+ ```python
77
+ with Relay(os.environ["RELAY_API_KEY"]) as relay:
78
+ print(relay.health())
79
+ ```
80
+
81
+ `AsyncRelay` is the same client for `asyncio`: every method is awaited, everything else is identical
82
+ ([Sync and async](#sync-and-async)).
83
+
84
+ ```python
85
+ from relaygpu import AsyncRelay
86
+
87
+ async with AsyncRelay(os.environ["RELAY_API_KEY"]) as relay:
88
+ result = await relay.image.generate("Qwen/qwen-image", {"prompt": "a lighthouse at dusk"})
89
+ ```
90
+
91
+ ## Quickstarts
92
+
93
+ Each one is a runnable file in [`examples/`](examples/) (how to run them:
94
+ [examples/README.md](examples/README.md)). They read `RELAY_API_KEY`, and `RELAY_BASE_URL` when set.
95
+
96
+ ### 1. Image
97
+
98
+ Generate an image with `Qwen/qwen-image`, print its link and save it.
99
+ ([`examples/01_image.py`](examples/01_image.py))
100
+
101
+ ```python
102
+ import os
103
+ import sys
104
+ import tempfile
105
+ from pathlib import Path
106
+
107
+ from relaygpu import ContentPolicyDeclinedError, InsufficientCreditsError, Relay
108
+
109
+ with Relay(os.environ["RELAY_API_KEY"], base_url=os.environ.get("RELAY_BASE_URL")) as relay:
110
+ try:
111
+ result = relay.image.generate(
112
+ "Qwen/qwen-image",
113
+ {"prompt": "A red fox in an autumn forest", "size": "1024x1024"},
114
+ )
115
+ image = result.images[0]
116
+ print("url:", image.url) # a link that lives 1 h; pass store_output="relay7d" to keep it longer
117
+
118
+ ext = (image.mime_type or "image/bin").split("/")[1]
119
+ path = Path(tempfile.gettempdir()) / f"relay-image.{ext}"
120
+ image.save(path) # downloads the link (without your key) or decodes base64
121
+ print("saved:", path)
122
+ except InsufficientCreditsError as e:
123
+ # Status picks the class, e.code refines it: KeyBudgetExhaustedError is an InsufficientCreditsError (402).
124
+ sys.exit(f"out of credit ({e.code}), request {e.request_id}")
125
+ except ContentPolicyDeclinedError as e:
126
+ sys.exit(f"the provider declined this prompt: {e}")
127
+ ```
128
+
129
+ `result.images` is a list of `RelayImage` (`url` or `b64`, `mime_type`, `to_bytes()`, `save(path)`);
130
+ `result.raw` is the response body exactly as the API sent it.
131
+
132
+ ### 2. Video, waiting for the result
133
+
134
+ Kling v3 text-to-video, 3 seconds, standard quality. `wait=True` long-polls the task until it ends.
135
+ ([`examples/02_video_wait.py`](examples/02_video_wait.py); the `asyncio` version is
136
+ [`examples/02_video_wait_async.py`](examples/02_video_wait_async.py))
137
+
138
+ ```python
139
+ import os
140
+ import sys
141
+
142
+ from relaygpu import Relay, TaskFailedError
143
+
144
+ with Relay(os.environ["RELAY_API_KEY"], base_url=os.environ.get("RELAY_BASE_URL")) as relay:
145
+ try:
146
+ task = relay.video.generate(
147
+ "KlingTeam/v3-T2V",
148
+ {"prompt": "A paper boat drifting down a rain gutter", "duration": 3, "quality_mode": "std"},
149
+ wait=True,
150
+ # Status transitions and elapsed time only: there is no queue position, log stream or cancel.
151
+ on_progress=lambda p: print(f"{p['status']} after {p['elapsed_seconds']}s"),
152
+ )
153
+ urls = (task.get("result") or {}).get("urls") or []
154
+ print("video:", urls[0] if urls else None) # expires 1 h after completion unless you pass store_output
155
+ except TaskFailedError as e:
156
+ # The task ran and failed: e.code is the task's error_code (e.g. CONTENT_POLICY_DECLINED, UPSTREAM_TIMEOUT).
157
+ sys.exit(f"task {e.task_id} failed: {e.code or 'unclassified'}: {e}")
158
+ ```
159
+
160
+ ### 3. Upload a file, then motion control
161
+
162
+ Both ways to send a local file: upload it explicitly (`files.upload`, here with 1-day retention), or put the
163
+ file straight into a `*_url` field and let the SDK upload it. The submit returns the `202` without waiting.
164
+ Arguments: a local video, then a character image URL.
165
+ ([`examples/03_upload_motion_control.py`](examples/03_upload_motion_control.py))
166
+
167
+ ```python
168
+ import os
169
+ import sys
170
+ from pathlib import Path
171
+
172
+ from relaygpu import FileTooLargeError, Relay, ValidationError
173
+
174
+ video_path, image_url = Path(sys.argv[1]), sys.argv[2]
175
+
176
+ with Relay(os.environ["RELAY_API_KEY"], base_url=os.environ.get("RELAY_BASE_URL")) as relay:
177
+ try:
178
+ # Form 1, explicit: host the file yourself and get a link any *_url field takes.
179
+ # relay1h is free; relay1d / relay7d / relay30d are billed per file (relay.pricing.get()["media_storage"]).
180
+ file = relay.files.upload(video_path, retention="relay1d") # streamed from disk; media type sniffed
181
+ print("uploaded:", file["file_id"], file["url"], "expires", file["expires_at"])
182
+
183
+ # Form 2, implicit: put the file straight into a *_url field; the SDK uploads it first (relay1h unless
184
+ # upload={"retention": ...} says otherwise) and sends the link. "video_url": file["url"] works the same.
185
+ accepted = relay.video.generate(
186
+ "KlingTeam/v3-Motion-Control",
187
+ {
188
+ "video_url": video_path.read_bytes(),
189
+ "image_url": image_url,
190
+ "character_orientation": "video", # "image" | "video": which input decides the facing direction
191
+ "duration": 5,
192
+ "quality_mode": "std",
193
+ },
194
+ )
195
+ # No wait: the 202 envelope. Poll later with relay.tasks.wait(accepted["task_id"]), or pass webhook_url.
196
+ print("task:", accepted["task_id"], "(replayed)" if accepted.get("replayed") else "")
197
+
198
+ # Take the explicit upload down now (idempotent; the fee is not refunded).
199
+ relay.files.delete(file["file_id"])
200
+ except FileTooLargeError:
201
+ sys.exit("files are capped at 100 MB")
202
+ except ValidationError as e:
203
+ sys.exit(f"rejected ({e.code}): {e}")
204
+ ```
205
+
206
+ ### 4. Webhook receiver
207
+
208
+ Verify each delivery on its raw body and branch on `event`. The example file adds `--self-test`, which signs
209
+ deliveries with a throwaway secret and posts them to itself, so it runs offline and without a tunnel.
210
+ ([`examples/04_webhook_receiver.py`](examples/04_webhook_receiver.py))
211
+
212
+ ```python
213
+ import os
214
+ from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
215
+
216
+ from relaygpu import Relay, WebhookVerificationError
217
+
218
+ relay = Relay(os.environ.get("RELAY_API_KEY")) # verify() itself needs no credential
219
+ # The account's signing secret: relay.webhooks.secret() (JWT or superkey). During the 24 h after a
220
+ # rotation, pass [current, previous] instead.
221
+ SECRET = os.environ["RELAY_WEBHOOK_SECRET"]
222
+
223
+
224
+ class Receiver(BaseHTTPRequestHandler):
225
+ def do_POST(self) -> None:
226
+ # Verify the bytes as sent, never a re-serialised json.loads result.
227
+ raw = self.rfile.read(int(self.headers.get("Content-Length") or 0))
228
+ try:
229
+ event = relay.webhooks.verify(raw, self.headers, SECRET)
230
+ except WebhookVerificationError as e:
231
+ print("rejected delivery:", e.code)
232
+ self.send_response(400)
233
+ self.end_headers()
234
+ return
235
+ # Delivery is at-least-once: dedupe on self.headers["webhook-id"] in production.
236
+ name = event["event"]
237
+ if name == "task.completed":
238
+ print("task done:", event["task_id"], event.get("result"))
239
+ elif name == "task.failed":
240
+ print("task failed:", event["task_id"], event.get("error_code"), event.get("error"))
241
+ elif name in ("workflow.completed", "workflow.failed"):
242
+ # Run events: the top-level status is always "completed" (the envelope). The run's own
243
+ # outcome is result["status"]: completed | failed | cancelled.
244
+ print(f"{name}:", event["task_id"], (event.get("result") or {}).get("status"))
245
+ else:
246
+ print("other event:", name) # instance.*
247
+ self.send_response(204)
248
+ self.end_headers()
249
+
250
+
251
+ ThreadingHTTPServer(("0.0.0.0", 8787), Receiver).serve_forever()
252
+ ```
253
+
254
+ `verify` takes `self.headers` as is (any object with `.get` or `.items`, case-insensitive), a `dict`, or
255
+ `httpx.Headers`. In a framework, pass the raw body: `await request.body()` (Starlette/FastAPI),
256
+ `request.get_data()` (Flask), `request.body` (Django).
257
+
258
+ ### 5. Workflow run
259
+
260
+ `script-voiceover`: an LLM writes a short line from your brief, a TTS model speaks it. Each step is billed as
261
+ an ordinary request. ([`examples/05_workflow_run.py`](examples/05_workflow_run.py))
262
+
263
+ ```python
264
+ import json
265
+ import os
266
+ import sys
267
+
268
+ from relaygpu import Relay, TaskFailedError
269
+
270
+ with Relay(os.environ["RELAY_API_KEY"], base_url=os.environ.get("RELAY_BASE_URL")) as relay:
271
+ try:
272
+ # Inputs follow the template's input_schema: relay.workflows.get("script-voiceover").
273
+ # Pass webhook_url for one signed workflow.* delivery instead of waiting.
274
+ run = relay.workflows.run(
275
+ "script-voiceover",
276
+ {"messages": [{"role": "user", "content": "a lighthouse at dusk"}], "voice": "Cherry"},
277
+ wait=True,
278
+ on_progress=lambda r: print(r["status"]),
279
+ )
280
+ print("run:", run["run_id"], run["status"])
281
+ print("output:", json.dumps(run.get("output"), indent=2)) # the last step's output (the speech)
282
+ except TaskFailedError as e:
283
+ # A failed or cancelled run raises; e.task is the run as polled.
284
+ sys.exit(f"run {e.task_id} did not complete: {e}")
285
+ ```
286
+
287
+ ### 6. Chat: use the SDK you already have
288
+
289
+ Chat and completions run on the official OpenAI and Anthropic SDKs: point them at Relay. This SDK does not wrap
290
+ chat. `openai` is an example dependency only (`pip install openai`); `relaygpu-client` does not need it.
291
+ ([`examples/06_chat_base_url.py`](examples/06_chat_base_url.py))
292
+
293
+ ```python
294
+ import os
295
+
296
+ from openai import OpenAI
297
+
298
+ base = os.environ.get("RELAY_BASE_URL") or "https://relaygpu.com"
299
+ client = OpenAI(api_key=os.environ["RELAY_API_KEY"], base_url=f"{base}/v2/openai/v1")
300
+
301
+ completion = client.chat.completions.create(
302
+ model="openai/gpt-4o-mini",
303
+ messages=[{"role": "user", "content": "Say hello in five words."}],
304
+ max_tokens=20,
305
+ )
306
+ print(completion.choices[0].message.content)
307
+ ```
308
+
309
+ ## Any model: `run()` and model resolution
310
+
311
+ ```python
312
+ from relaygpu import is_accepted
313
+
314
+ out = relay.run("Qwen/qwen-image", {"prompt": "a lighthouse at dusk"}) # the output body
315
+ sub = relay.run("KlingTeam/v3-T2V", {"prompt": "waves", "duration": 3}, wait=False)
316
+ if is_accepted(sub):
317
+ print(sub["task_id"]) # the 202 envelope
318
+ ```
319
+
320
+ The SDK ships no model table. Every call resolves the model through `relay.models.get(name)`
321
+ (`GET /v2/models/{name}`, cached 5 min per client): its route, whether `model` goes in the body, whether the
322
+ route is async by default, and its request/response schemas. A model added to Relay after this release works
323
+ through `run()` and the family helpers (`image`, `video`, `audio`) without an upgrade; the models this release
324
+ knows also get `Literal` hints for autocomplete.
325
+
326
+ - An unknown name raises `ModelNotFoundError` (404); a retired one raises `ModelRetiredError` (403). Both
327
+ before anything is submitted or billed.
328
+ - `relay.models.list(tag="text-to-video")` lists the catalog; `relay.models.get(name)` returns one model with
329
+ its `request_schema`, `request_example` and `pricing`.
330
+ - `run()` returns a sync route's body, or waits for an async route and returns the task's `result`.
331
+ `wait=False` returns the `202` envelope instead (narrow it with `is_accepted`).
332
+ - Request options are keyword arguments on every call: `timeout`, `on_progress`, `mode`, `store_output`,
333
+ `webhook_url`, `idempotency_key`, `async_` (trailing underscore: `async` is a keyword), `upload`,
334
+ `inline_images`.
335
+
336
+ ## Async tasks and `tasks.wait`
337
+
338
+ Video (and any call with `async_=True`) answers `202` with a `task_id`. `relay.video.generate` returns that
339
+ envelope unless you pass `wait=True`; `run()` and the image/audio helpers wait by default.
340
+
341
+ ```python
342
+ task = relay.tasks.wait(
343
+ task_id,
344
+ timeout=20 * 60, # seconds; the default. The task keeps running past it (APITimeoutError)
345
+ on_progress=lambda p: print(p["status"], p["elapsed_seconds"]),
346
+ )
347
+ ```
348
+
349
+ `tasks.wait` long-polls `GET /v2/tasks/{id}?wait=30`: a 15 s task costs one or two requests, not a poll every
350
+ second. `on_progress` fires on status transitions (`queued` → `running` → `completed`) with `elapsed_seconds`,
351
+ and that is all there is: no queue position, no logs and no cancel, by design. A failed task raises
352
+ `TaskFailedError` with `code` = the task's `error_code`. Task polls need no key: the task id is the access
353
+ token. A task's result expires 1 hour after it finishes. `relay.tasks.get(task_id)` reads the status once.
354
+
355
+ Workflow runs have no long-poll: `workflows.wait_run` polls the run every 1 s, backing off to 5 s (default
356
+ budget 30 min).
357
+
358
+ ## Idempotency and retries
359
+
360
+ Every async submit (`202` routes, `video.generate`, `run()` on an async route) and every workflow run is sent
361
+ with an `Idempotency-Key`: yours (`idempotency_key=...`) or a generated UUID. That makes the submit safe to
362
+ retry, so the SDK retries it on a network error, a timeout or a 5xx (up to 2 times, same key). When the server
363
+ recognises a key it has already accepted, it answers with the original `202` and the SDK returns it with
364
+ `"replayed": True`: no second task, no second charge. The same key with a different body raises
365
+ `IdempotencyKeyReusedError` (422). Keys live 24 hours.
366
+
367
+ ```python
368
+ a = relay.video.generate("KlingTeam/v3-T2V", {"prompt": "waves", "duration": 3}, idempotency_key="order-1234")
369
+ b = relay.video.generate("KlingTeam/v3-T2V", {"prompt": "waves", "duration": 3}, idempotency_key="order-1234")
370
+ # b["task_id"] == a["task_id"], b["replayed"] is True
371
+ ```
372
+
373
+ Sync calls (an image route answering `200`, chat, TTS) carry no key and are **never retried**: a retry would
374
+ run, and bill, the request again. GETs retry on 429/503 (honouring `Retry-After`, up to 3 times) and on other
375
+ 5xx (up to 2). A `Retry-After` longer than `max_retry_after` (60 s) is not waited out: the error is raised.
376
+ `retry=False` turns every retry off.
377
+
378
+ ## Files and implicit uploads
379
+
380
+ ```python
381
+ file = relay.files.upload(Path("dance.mp4"), retention="relay1d") # {file_id, url, expires_at, …}
382
+ relay.files.copy("https://example.com/clip.mp4", retention="relay7d") # Relay fetches it
383
+ relay.files.get(file["file_id"])
384
+ relay.files.list(source="upload") # one page; files.list_all(...) iterates every page
385
+ relay.files.delete(file["file_id"]) # the link goes down now; no refund
386
+ ```
387
+
388
+ - `files.upload` takes `bytes`, a path (`pathlib.Path` or `str`), a binary file object (`open(p, "rb")`) or an
389
+ iterator of bytes (`AsyncRelay` also takes an async iterator). Paths and file objects are streamed in
390
+ chunks, never read whole into memory.
391
+ - Any `*_url` / `*_urls` input also takes `bytes`, a `Path`, a binary file object or an iterator of bytes: the
392
+ SDK uploads it first and sends the link (quickstart 3). There a `str` is always a URL, never a path.
393
+ Implicit uploads use `relay1h` unless you pass `upload={"retention": "relay1d"}`.
394
+ - `inline_images=True` sends images of 4 MB or less as base64 instead, on routes whose schema takes it; larger
395
+ ones are still uploaded.
396
+ - One file is at most 100 MB (the server answers `413`, `FileTooLargeError`). The media type comes from
397
+ `content_type=`, else from the file's first bytes.
398
+ - `relay1h` is free within a daily quota; `relay1d`, `relay7d` and `relay30d` are billed per file at upload.
399
+
400
+ ## How long result links live
401
+
402
+ Result links (`urls`, `audio_url`, …) and uploaded files expire. By default a result link lives **1 hour**
403
+ after the task finishes. To keep outputs longer, buy storage per request with `store_output`:
404
+
405
+ ```python
406
+ relay.image.generate("Qwen/qwen-image", {"prompt": "a lighthouse"}, store_output="relay7d")
407
+ ```
408
+
409
+ | `store_output` / `retention` | Lifetime | Fee |
410
+ |---|---|---|
411
+ | `provider` (default) / `relay1h` | 1 hour | free |
412
+ | `relay1d` | 1 day | per file |
413
+ | `relay7d` | 7 days | per file |
414
+ | `relay30d` | 30 days | per file |
415
+
416
+ The fees are in `relay.pricing.get()["media_storage"]`; read them there rather than hard-coding them.
417
+
418
+ ## Webhooks
419
+
420
+ Pass `webhook_url` on an async submit (one `task.completed` or `task.failed`) or on a workflow run (one
421
+ `workflow.completed` or `workflow.failed`). Deliveries are signed with
422
+ [Standard Webhooks](https://www.standardwebhooks.com/); `relay.webhooks.verify(raw_body, headers, secret)`
423
+ checks the signature and the timestamp (±5 min, `tolerance=300`) and returns the typed event, discriminated on
424
+ `event` (quickstart 4). It makes no request, needs no credential, and is a plain (not awaited) call on both
425
+ `Relay` and `AsyncRelay`; `relaygpu.verify_webhook(...)` is the same function without a client.
426
+
427
+ - Verify the **raw** body bytes, not a re-serialised `json.loads` result.
428
+ - Rotation: `relay.webhooks.rotate_secret()` issues a new secret and the previous one keeps working for 24
429
+ hours. During that window pass both: `verify(raw, headers, [current, previous])`.
430
+ - On `workflow.*` events the top-level `status` is always `"completed"` (it describes the delivery). Branch on
431
+ `event`, or on `result["status"]` (`completed`, `failed`, `cancelled`; a cancelled run arrives as
432
+ `workflow.failed`).
433
+ - Delivery is at-least-once: dedupe on the `webhook-id` header.
434
+ - `relay.webhooks.secret()`, `.rotate_secret()` and `.deliveries.list()` / `.deliveries.get(task_id_or_run_id)`
435
+ need a JWT or a partner superkey.
436
+
437
+ ## Errors
438
+
439
+ Every error extends `RelayError`. HTTP errors extend `RelayAPIError`; the HTTP status picks the class:
440
+
441
+ | Status | Class |
442
+ |---|---|
443
+ | 400 | `InvalidRequestError` |
444
+ | 401 | `AuthenticationError` |
445
+ | 402 | `InsufficientCreditsError` |
446
+ | 403 | `PermissionDeniedError` |
447
+ | 404 | `NotFoundError` |
448
+ | 409 | `ConflictError` |
449
+ | 410 | `GoneError` |
450
+ | 422 | `ValidationError` |
451
+ | 429 | `RateLimitError` |
452
+ | 500 | `RelayInternalError` |
453
+ | 502 | `ProviderError` |
454
+ | 503 | `CapacityError` |
455
+ | 504 | `UpstreamTimeoutError` |
456
+
457
+ The `error.code` then picks a subclass: `KEY_BUDGET_EXHAUSTED` raises `KeyBudgetExhaustedError`, which is an
458
+ `InsufficientCreditsError`; `CONTENT_POLICY_DECLINED` raises `ContentPolicyDeclinedError` (a
459
+ `ValidationError`). Every code in Relay's error catalog has a class (all in `relaygpu.errors`, re-exported
460
+ from `relaygpu`). A code this release does not know falls back to the status class and keeps `e.code`;
461
+ `e.code` may also be `None`.
462
+
463
+ ```python
464
+ from relaygpu import KeyBudgetExhaustedError, RateLimitError, RelayAPIError
465
+
466
+ try:
467
+ relay.image.generate("Qwen/qwen-image", {"prompt": "a lighthouse"})
468
+ except KeyBudgetExhaustedError:
469
+ print("this key's budget is spent")
470
+ except RateLimitError as e:
471
+ print(f"retry in {e.retry_after}s")
472
+ except RelayAPIError as e:
473
+ print(e.status, e.code, e.request_id, e.detail)
474
+ ```
475
+
476
+ - `e.request_id` is the id to quote to support; `e.retry_after` (seconds) is set from `Retry-After`;
477
+ `e.detail` is the server's detail.
478
+ - `TaskFailedError`: an async task (or workflow run) ended `failed`; `e.code` is the task's `error_code`,
479
+ `e.task_id` and `e.task` carry the rest.
480
+ - `APIConnectionError` (no HTTP answer), `APITimeoutError` (a `timeout` ran out),
481
+ `WebhookVerificationError` (a delivery failed verification).
482
+ - Branch on the class or `e.code`, never on the message, and never on `error.source` in the response body.
483
+ - `relaygpu.FileNotFoundError` (`FILE_NOT_FOUND`, a `NotFoundError`) is Relay's error, not the builtin.
484
+ `from relaygpu import *` shadows the builtin `FileNotFoundError` in your module; import the names you use,
485
+ or write `relaygpu.FileNotFoundError` / `relaygpu.errors.FileNotFoundError`.
486
+
487
+ ## Account and keys
488
+
489
+ ```python
490
+ import time
491
+
492
+ relay = Relay(jwt=dashboard_jwt) # or Relay(superkey) for a partner (custom) tier
493
+
494
+ relay.account.credits()
495
+ relay.account.usage()
496
+ relay.account.usage_timeseries(start_time=int(time.time()) - 86_400)
497
+
498
+ created = relay.keys.create(name="end-user-42", key_budget=5)
499
+ print(created["key"]) # the secret, shown ONCE
500
+ for k in relay.keys.list_all():
501
+ print(k["key_id"], k["name"])
502
+ relay.keys.topup(key_id, 2)
503
+ relay.keys.revoke(key_id)
504
+ ```
505
+
506
+ These routes take a dashboard JWT or a custom-tier superkey. A plain inference key gets the server's 403
507
+ (`PermissionDeniedError`); the SDK does not guess the key class. `keys.create` returns the full secret in
508
+ `key` once; every later response masks it, so store it then. Address keys by `key_id`. `key_budget` is for
509
+ custom tiers only. Query parameters are keyword arguments with the API's names (`from` is spelled `from_`).
510
+
511
+ ## Cost estimates
512
+
513
+ ```python
514
+ est = relay.estimate_cost("KlingTeam/v3-T2V", {"duration_seconds": 3, "quality_mode": "std"})
515
+ print(est["usd"], est["basis"])
516
+ ```
517
+
518
+ `estimate_cost` computes from the public `/v2/pricing` rows. It is an **estimate, not an invoice**: you are
519
+ billed from provider-reported usage, and custom-tier prices are not in the public list.
520
+
521
+ ## Sync and async
522
+
523
+ One core, two transports. `AsyncRelay` (`httpx.AsyncClient`) is the source; `Relay` (`httpx.Client`) is
524
+ generated from it, so the two have the same namespaces, methods, arguments and return types. The only
525
+ differences:
526
+
527
+ | | `Relay` | `AsyncRelay` |
528
+ |---|---|---|
529
+ | Calls | `relay.image.generate(...)` | `await relay.image.generate(...)` |
530
+ | Lifetime | `with Relay(...) as relay` / `relay.close()` | `async with AsyncRelay(...) as relay` / `await relay.aclose()` |
531
+ | Paging | `for k in relay.keys.list_all()` | `async for k in relay.keys.list_all()` |
532
+ | Image bytes | `image.to_bytes()`, `image.save(p)` | `await image.to_bytes()`, `await image.save(p)` |
533
+ | Upload streams | iterator of bytes | iterator or async iterator of bytes |
534
+ | `webhooks.verify` | plain call | plain call (no `await`: it does no I/O) |
535
+
536
+ There is no `AbortSignal`. Cancellation belongs to the event loop: `asyncio.timeout(...)`, `asyncio.wait_for`
537
+ or `task.cancel()` stop an `AsyncRelay` call where it is (on a wait, the task itself keeps running
538
+ server-side). In sync code, `timeout=` bounds each call. Bring your own `httpx` client with
539
+ `http_client=httpx.Client(...)` / `httpx.AsyncClient(...)` (proxies, transports, mounts); the SDK does not
540
+ close a client you passed in. `default_headers={...}` adds headers to every request.
541
+
542
+ ## Types
543
+
544
+ Request and response bodies are `TypedDict`s generated from Relay's public OpenAPI spec, in `relaygpu.types`
545
+ (for example `TaskStatus`, `FileObject`, `WorkflowRunState`, `WebhookEvent`, `UsageInput`), so mypy and pyright
546
+ check the keys you read. The package ships `py.typed` and is `mypy --strict` clean. Responses are plain
547
+ `dict`s at runtime: no wrapper objects, `json.dumps` works on them. The image helpers are the exception:
548
+ `ImageResult` / `RelayImage` are small frozen dataclasses.
549
+
550
+ ## Runtimes
551
+
552
+ CPython 3.10, 3.11, 3.12 and 3.13, tested in CI on each. One runtime dependency, `httpx`. Sync and async
553
+ (`AsyncRelay` runs on `asyncio`). `relay.request(method, path, json=...)` reaches any route
554
+ with the same errors and retry policy.
555
+
556
+ ## Chat
557
+
558
+ Chat is a base-URL swap on the official SDKs, in any language:
559
+
560
+ | SDK | Base URL |
561
+ |---|---|
562
+ | OpenAI | `https://relaygpu.com/v2/openai/v1` |
563
+ | Anthropic | `https://relaygpu.com/v2/anthropic` |
564
+
565
+ ## License
566
+
567
+ MIT