tai-sdk 0.1.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (37) hide show
  1. tai_sdk-0.1.0/.gitignore +7 -0
  2. tai_sdk-0.1.0/API_CONTRACT.md +366 -0
  3. tai_sdk-0.1.0/CHANGELOG.md +45 -0
  4. tai_sdk-0.1.0/LICENSE +21 -0
  5. tai_sdk-0.1.0/PKG-INFO +503 -0
  6. tai_sdk-0.1.0/README.md +476 -0
  7. tai_sdk-0.1.0/examples/assistant_thread.py +111 -0
  8. tai_sdk-0.1.0/examples/async_chat.py +76 -0
  9. tai_sdk-0.1.0/examples/quickstart.py +85 -0
  10. tai_sdk-0.1.0/examples/streaming.py +94 -0
  11. tai_sdk-0.1.0/pyproject.toml +59 -0
  12. tai_sdk-0.1.0/src/tai_sdk/__init__.py +137 -0
  13. tai_sdk-0.1.0/src/tai_sdk/_client.py +220 -0
  14. tai_sdk-0.1.0/src/tai_sdk/_config.py +332 -0
  15. tai_sdk-0.1.0/src/tai_sdk/_http.py +538 -0
  16. tai_sdk-0.1.0/src/tai_sdk/_streaming.py +275 -0
  17. tai_sdk-0.1.0/src/tai_sdk/_version.py +1 -0
  18. tai_sdk-0.1.0/src/tai_sdk/errors.py +437 -0
  19. tai_sdk-0.1.0/src/tai_sdk/py.typed +0 -0
  20. tai_sdk-0.1.0/src/tai_sdk/resources/__init__.py +11 -0
  21. tai_sdk-0.1.0/src/tai_sdk/resources/_base.py +138 -0
  22. tai_sdk-0.1.0/src/tai_sdk/resources/assistants.py +99 -0
  23. tai_sdk-0.1.0/src/tai_sdk/resources/chat.py +101 -0
  24. tai_sdk-0.1.0/src/tai_sdk/resources/models.py +39 -0
  25. tai_sdk-0.1.0/src/tai_sdk/resources/threads.py +159 -0
  26. tai_sdk-0.1.0/src/tai_sdk/resources/usage.py +19 -0
  27. tai_sdk-0.1.0/src/tai_sdk/types.py +544 -0
  28. tai_sdk-0.1.0/tests/conftest.py +192 -0
  29. tai_sdk-0.1.0/tests/test_assistants.py +196 -0
  30. tai_sdk-0.1.0/tests/test_chat.py +186 -0
  31. tai_sdk-0.1.0/tests/test_cleanup_and_models.py +195 -0
  32. tai_sdk-0.1.0/tests/test_client.py +192 -0
  33. tai_sdk-0.1.0/tests/test_config.py +246 -0
  34. tai_sdk-0.1.0/tests/test_errors.py +215 -0
  35. tai_sdk-0.1.0/tests/test_retries.py +361 -0
  36. tai_sdk-0.1.0/tests/test_streaming.py +440 -0
  37. tai_sdk-0.1.0/tests/test_threads.py +266 -0
@@ -0,0 +1,7 @@
1
+ __pycache__/
2
+ *.py[cod]
3
+ .pytest_cache/
4
+ build/
5
+ dist/
6
+ *.egg-info/
7
+ .venv/
@@ -0,0 +1,366 @@
1
+ # TAI Assistant API — contract v1
2
+
3
+ This is the interface that **both** the server implementation and the Python SDK
4
+ must follow exactly. Nothing here is OpenAI-compatible on purpose; it is our own
5
+ shape.
6
+
7
+ Base path: `/api/v1`
8
+ Auth: `Authorization: Bearer sk-tai-...` on every endpoint.
9
+ Content type: `application/json` (except SSE streams).
10
+
11
+ ---
12
+
13
+ ## 1. Conventions
14
+
15
+ ### 1.1 Error shape
16
+
17
+ Every non-2xx response is:
18
+
19
+ ```json
20
+ { "code": "invalid_request", "message": "Human readable sentence", "param": "model" }
21
+ ```
22
+
23
+ `param` is optional. Stable codes:
24
+
25
+ | code | HTTP | meaning |
26
+ |---|---|---|
27
+ | `missing_api_key` | 401 | no `Authorization: Bearer` header |
28
+ | `invalid_api_key` | 401 | unknown or revoked key |
29
+ | `account_disabled` | 403 | the owning account is disabled |
30
+ | `insufficient_balance` | 402 | balance exhausted (reserved; not enforced yet) |
31
+ | `not_found` | 404 | no such object, or it belongs to another account |
32
+ | `invalid_request` | 422 | validation failed |
33
+ | `model_not_available` | 409 | the model exists but is not live (e.g. still training) |
34
+ | `model_not_found` | 404 | unknown model id |
35
+ | `rate_limited` | 429 | too many requests |
36
+ | `backend_unavailable` | 502 | the model backend failed |
37
+ | `stream_aborted` | 499 | client disconnected mid-stream (server-side only) |
38
+
39
+ ### 1.2 Object ids
40
+
41
+ Opaque strings with a type prefix, generated with `secrets.token_urlsafe(18)`:
42
+
43
+ ```
44
+ asst_<token> assistant
45
+ thrd_<token> thread
46
+ msg_<token> message
47
+ chat_<token> stateless chat completion
48
+ ```
49
+
50
+ ### 1.3 Timestamps
51
+
52
+ UTC ISO-8601 with a `Z`-less `+00:00` offset, second precision:
53
+ `2026-09-30T15:04:05+00:00`. Field names are `created_at` / `updated_at`.
54
+
55
+ ### 1.4 Pagination
56
+
57
+ List endpoints accept `limit` (1–100, default 20) and `after` (an object id).
58
+ They return:
59
+
60
+ ```json
61
+ { "object": "list", "data": [ ... ], "has_more": false, "first_id": "asst_…", "last_id": "asst_…" }
62
+ ```
63
+
64
+ `data` is newest-first.
65
+
66
+ ---
67
+
68
+ ## 2. Assistants
69
+
70
+ A reusable configuration: model + system instructions + a name.
71
+
72
+ ```
73
+ POST /api/v1/assistants
74
+ GET /api/v1/assistants?limit=&after=
75
+ GET /api/v1/assistants/{assistant_id}
76
+ PATCH /api/v1/assistants/{assistant_id}
77
+ DELETE /api/v1/assistants/{assistant_id} -> 204
78
+ ```
79
+
80
+ **Assistant object**
81
+
82
+ ```json
83
+ {
84
+ "id": "asst_…",
85
+ "object": "assistant",
86
+ "name": "Study Buddy",
87
+ "model": "tfmf",
88
+ "instructions": "You are a patient tutor.",
89
+ "metadata": {},
90
+ "created_at": "2026-09-30T15:04:05+00:00",
91
+ "updated_at": "2026-09-30T15:04:05+00:00"
92
+ }
93
+ ```
94
+
95
+ **Create body**
96
+
97
+ | field | type | required | notes |
98
+ |---|---|---|---|
99
+ | `model` | string | yes | must be a **live** model id |
100
+ | `name` | string | no | 1–80 chars, default `"Assistant"` |
101
+ | `instructions` | string | no | ≤ 8000 chars, default `""` |
102
+ | `metadata` | object | no | ≤ 16 keys, flat string values |
103
+
104
+ **PATCH body** — any subset of `name`, `model`, `instructions`, `metadata`.
105
+
106
+ ---
107
+
108
+ ## 3. Threads
109
+
110
+ A persistent conversation.
111
+
112
+ ```
113
+ POST /api/v1/threads
114
+ GET /api/v1/threads?limit=&after=
115
+ GET /api/v1/threads/{thread_id}
116
+ PATCH /api/v1/threads/{thread_id}
117
+ DELETE /api/v1/threads/{thread_id} -> 204
118
+ ```
119
+
120
+ **Thread object**
121
+
122
+ ```json
123
+ {
124
+ "id": "thrd_…",
125
+ "object": "thread",
126
+ "title": "Factorial help",
127
+ "assistant_id": "asst_…",
128
+ "metadata": {},
129
+ "message_count": 4,
130
+ "created_at": "…",
131
+ "updated_at": "…"
132
+ }
133
+ ```
134
+
135
+ **Create body**
136
+
137
+ | field | type | required | notes |
138
+ |---|---|---|---|
139
+ | `assistant_id` | string | no | if omitted, `model` is required at message time |
140
+ | `model` | string | no | default model for this thread when no assistant is set |
141
+ | `title` | string | no | default `"New thread"` |
142
+ | `metadata` | object | no | |
143
+
144
+ **PATCH body** — any subset of `title`, `assistant_id`, `model`, `metadata`.
145
+
146
+ Deleting a thread deletes its messages (cascade).
147
+
148
+ ---
149
+
150
+ ## 4. Messages
151
+
152
+ ```
153
+ POST /api/v1/threads/{thread_id}/messages append + generate
154
+ GET /api/v1/threads/{thread_id}/messages?limit=&after=&order=asc|desc
155
+ ```
156
+
157
+ ### 4.1 POST body
158
+
159
+ | field | type | default | notes |
160
+ |---|---|---|---|
161
+ | `content` | string | — | required, 1–32000 chars |
162
+ | `stream` | bool | `false` | see §6 |
163
+ | `temperature` | float | `0.7` | 0.0–2.0 |
164
+ | `max_output_tokens` | int | `512` | 1–8192 |
165
+
166
+ Also accepts an optional `role` (only `"user"` is allowed from clients).
167
+
168
+ ### 4.2 Non-streaming response (200)
169
+
170
+ Both messages are returned so the caller does not need a second round trip:
171
+
172
+ ```json
173
+ {
174
+ "object": "message.exchange",
175
+ "thread_id": "thrd_…",
176
+ "user_message": { …message… },
177
+ "assistant_message": { …message… },
178
+ "usage": {
179
+ "model": "tfmf",
180
+ "input_tokens": 128,
181
+ "output_tokens": 42,
182
+ "cost_cny": 0.00019,
183
+ "estimated": true
184
+ }
185
+ }
186
+ ```
187
+
188
+ ### 4.3 Message object
189
+
190
+ ```json
191
+ {
192
+ "id": "msg_…",
193
+ "object": "message",
194
+ "thread_id": "thrd_…",
195
+ "role": "user",
196
+ "content": "How do I reverse a list?",
197
+ "model": null,
198
+ "usage": null,
199
+ "created_at": "…"
200
+ }
201
+ ```
202
+
203
+ `model` and `usage` are populated on assistant messages only.
204
+
205
+ ### 4.4 GET
206
+
207
+ Returns a list envelope (§1.4) of message objects. `order` defaults to `desc`.
208
+
209
+ ---
210
+
211
+ ## 5. Stateless chat
212
+
213
+ ```
214
+ POST /api/v1/chat
215
+ ```
216
+
217
+ | field | type | default | notes |
218
+ |---|---|---|---|
219
+ | `model` | string | — | required unless `assistant_id` given |
220
+ | `assistant_id` | string | — | supplies model + instructions |
221
+ | `messages` | array | — | 1–100 items, each `{role, content}`, role in `user`/`assistant`/`system` |
222
+ | `stream` | bool | `false` | |
223
+ | `temperature` | float | `0.7` | |
224
+ | `max_output_tokens` | int | `512` | |
225
+
226
+ **Response (200)**
227
+
228
+ ```json
229
+ {
230
+ "id": "chat_…",
231
+ "object": "chat.completion",
232
+ "model": "tfmf",
233
+ "output": { "role": "assistant", "content": "…" },
234
+ "usage": { "input_tokens": 12, "output_tokens": 40, "cost_cny": 0.000126, "estimated": true },
235
+ "created_at": "…"
236
+ }
237
+ ```
238
+
239
+ ---
240
+
241
+ ## 6. Streaming (SSE)
242
+
243
+ When `stream` is `true` the response is `text/event-stream` with these events.
244
+ Every `data:` line is a JSON object on a single line.
245
+
246
+ ```
247
+ event: message.start
248
+ data: {"id":"msg_…","model":"tfmf","thread_id":"thrd_…"}
249
+
250
+ event: message.delta
251
+ data: {"delta":"Hello"}
252
+
253
+ event: message.done
254
+ data: {"id":"msg_…","content":"Hello there","usage":{"input_tokens":12,"output_tokens":40,"cost_cny":0.000126,"estimated":true}}
255
+
256
+ event: error
257
+ data: {"code":"backend_unavailable","message":"…"}
258
+ ```
259
+
260
+ Rules:
261
+
262
+ * `message.start` is always first; `message.done` or `error` is always last.
263
+ * `thread_id` is `null` for `/api/v1/chat` streams.
264
+ * On completion the assistant message is persisted (thread streams only) and
265
+ `usage` is recorded exactly once, then the dashboard reflects it.
266
+ * A client disconnect must still persist whatever was generated and record usage.
267
+
268
+ ---
269
+
270
+ ## 7. Models and usage
271
+
272
+ ```
273
+ GET /api/v1/models -> list of live models
274
+ GET /api/v1/usage -> same body as the cookie-authenticated /api/usage
275
+ ```
276
+
277
+ `GET /api/v1/models` response:
278
+
279
+ ```json
280
+ {
281
+ "object": "list",
282
+ "data": [
283
+ {
284
+ "id": "tfmf",
285
+ "object": "model",
286
+ "name": "TFMF",
287
+ "live": true,
288
+ "context_window": 8192,
289
+ "input_cny_per_1m": 0.5,
290
+ "output_cny_per_1m": 3.0
291
+ }
292
+ ]
293
+ }
294
+ ```
295
+
296
+ `GET /api/v1/usage` returns the `db.usage_summary()` payload plus
297
+ `"object": "usage"`.
298
+
299
+ ---
300
+
301
+ ## 8. Server implementation notes
302
+
303
+ * New tables: `assistants`, `threads`, `messages`. All scoped by `user_id` and
304
+ cascading on user delete. New columns `threads.model`, `messages.model`,
305
+ `messages.usage_json`.
306
+ * Every generation path must call `db.record_usage(...)` **exactly once**, so the
307
+ existing dashboard totals stay correct. `GET /api/usage` must not break.
308
+ * The model backend is pluggable, selected by `TAI_MODEL_BACKEND`:
309
+ * `stub` (default) — deterministic, offline, no model. Responses carry
310
+ `"backend": "stub"` so nobody mistakes them for real output.
311
+ * `openai_compatible` — POSTs `TAI_MODEL_BASE_URL` + `/chat/completions` with
312
+ `TAI_MODEL_API_KEY`. Works with vLLM, llama.cpp server, Ollama, LM Studio.
313
+ * Token counts: use the backend's reported usage when present, otherwise the
314
+ documented estimate `max(1, ceil(len(text)/4))`, and set `"estimated": true`.
315
+ * `stub` must never claim to be a real model in the response.
316
+
317
+ ## 9. SDK notes
318
+
319
+ * Package `tai-sdk`, import name `tai_sdk`, `src/` layout, `pyproject.toml`
320
+ with hatchling.
321
+ * Runtime dependency: `httpx` only.
322
+ * Public surface: `TAI` (sync) and `AsyncTAI` (async), plus `errors` types.
323
+ * Resource namespaces: `client.assistants`, `client.threads`,
324
+ `client.threads.messages`, `client.chat`, `client.models`, `client.usage`.
325
+ * Base URL resolution order:
326
+ 1. `base_url=` argument
327
+ 2. `TAI_BASE_URL` environment variable
328
+ 3. `~/.tai/config.json` (`{"base_url": "…"}`)
329
+ 4. the discovery document (§10), cached under `~/.tai/endpoint-cache.json`
330
+ 5. `https://api.tai-research.dev` (last-resort default)
331
+ * API key resolution: `api_key=` argument, then `TAI_API_KEY`.
332
+ * Send `ngrok-skip-browser-warning: true` on every request so free ngrok tunnels
333
+ do not inject their interstitial HTML into API responses.
334
+ * Retry idempotent requests (GET/DELETE) and connection errors up to
335
+ `max_retries` times with exponential backoff + jitter. Never retry a
336
+ non-idempotent POST unless the failure happened before the request was sent.
337
+ * Streaming must be incremental, and the sync iterator must close the response
338
+ when the caller breaks out of the loop.
339
+
340
+ ## 10. Endpoint discovery
341
+
342
+ GitHub Pages cannot proxy requests, so it is used as a **discovery document**
343
+ instead. `TAI()` with no `base_url` fetches:
344
+
345
+ ```
346
+ GET https://ltyleo.github.io/platform/api-endpoint.json
347
+ ```
348
+
349
+ ```json
350
+ {
351
+ "object": "endpoint",
352
+ "base_url": "https://xxxx.ngrok-free.app",
353
+ "updated_at": "2026-09-30T15:04:05+00:00",
354
+ "note": "TAI Developer Platform API. Update this file when the tunnel changes."
355
+ }
356
+ ```
357
+
358
+ Rules:
359
+
360
+ * Short timeout (3 s). Failure is non-fatal — fall through to the next source.
361
+ * Cache the result in `~/.tai/endpoint-cache.json` and reuse it for 12 hours, so
362
+ the SDK works offline and does not hit Pages on every client construction.
363
+ * `TAI_DISCOVERY_URL` overrides the document location; `discover=False`
364
+ disables it entirely.
365
+ * Because the ngrok URL can change, every SDK error message about connectivity
366
+ should mention the resolved base URL so users can see what it tried.
@@ -0,0 +1,45 @@
1
+ # Changelog
2
+
3
+ All notable changes to `tai-sdk` are documented here. This project follows
4
+ [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
5
+
6
+ ## 0.1.0 — 2026-09-30
7
+
8
+ First public release, implementing the TAI Assistant API contract v1.
9
+
10
+ ### Added
11
+
12
+ - `TAI` (sync) and `AsyncTAI` (async) clients, both usable as context managers.
13
+ - Resource namespaces: `client.assistants`, `client.threads`,
14
+ `client.threads.messages` (also `client.messages`), `client.chat`,
15
+ `client.models`, `client.usage`.
16
+ - Assistants: create, list (with `limit`/`after` pagination), get, update
17
+ (`PATCH`), delete.
18
+ - Threads: create, list, get, update, delete (messages cascade server-side).
19
+ - Messages: append + generate (returns both the user and assistant message plus
20
+ usage in one round trip), and list with `limit`/`after`/`order`.
21
+ - Stateless chat: `client.chat.create(...)`.
22
+ - Models list and account usage summary.
23
+ - Incremental SSE streaming via `client.chat.stream(...)` and
24
+ `client.threads.messages.create(..., stream=True)`, yielding typed
25
+ `MessageStart`, `MessageDelta` and `MessageDone` events and raising on
26
+ `error` frames. Works with `for` (sync) and `async for` (async); the HTTP
27
+ response is always closed, including when the caller breaks out early.
28
+ - Base URL resolution: `base_url=` argument, `TAI_BASE_URL`,
29
+ `~/.tai/config.json`, the GitHub Pages discovery document (3 s timeout,
30
+ cached 12 h in `~/.tai/endpoint-cache.json`), then
31
+ `https://api.tai-research.dev`. `discover=False` disables discovery and
32
+ `TAI_DISCOVERY_URL` relocates the document.
33
+ - API key resolution: `api_key=` argument, then `TAI_API_KEY`.
34
+ - `ngrok-skip-browser-warning: true` on every request.
35
+ - Full error hierarchy: every non-2xx becomes the matching exception carrying
36
+ `code`, `message`, `param`, `status_code` and `request_id`. Connectivity
37
+ failures include the resolved base URL in the message.
38
+ - Retries with exponential backoff and jitter on connection errors and
39
+ HTTP 429/5xx, honouring `Retry-After`. Only idempotent methods (plus POSTs
40
+ that provably never left the client) are retried; `max_retries` defaults to 2.
41
+ - Configurable timeouts (30 s connect/read by default; the read timeout is
42
+ disabled for streams).
43
+ - Response objects are small dataclasses with attribute access, a `.raw` dict
44
+ and a `__repr__`. No pydantic.
45
+ - Typed package (`py.typed`), Python 3.9+, runtime dependency: `httpx` only.
tai_sdk-0.1.0/LICENSE ADDED
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 TAI Research
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.