ag_ui_antigravity 0.1.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2025
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,658 @@
1
+ Metadata-Version: 2.4
2
+ Name: ag_ui_antigravity
3
+ Version: 0.1.0
4
+ Summary: Google Antigravity integration for the AG-UI Protocol
5
+ License-Expression: MIT
6
+ License-File: LICENSE
7
+ Requires-Dist: ag-ui-protocol>=1.0.0,<2.0
8
+ Requires-Dist: google-antigravity>=0.1.8,<0.2.0
9
+ Requires-Dist: fastapi>=0.115.2
10
+ Requires-Dist: pydantic>=2.11.7
11
+ Requires-Dist: sse-starlette>=2.1.0
12
+ Requires-Dist: uvicorn>=0.35.0
13
+ Requires-Python: >=3.10, <3.15
14
+ Project-URL: Homepage, https://github.com/ag-ui-protocol/ag-ui/tree/main/integrations/antigravity/python
15
+ Project-URL: Issues, https://github.com/ag-ui-protocol/ag-ui/issues
16
+ Description-Content-Type: text/markdown
17
+
18
+ # AG-UI ⚡ Google Antigravity
19
+
20
+ Implementation of the [AG-UI protocol](https://github.com/ag-ui-protocol/ag-ui)
21
+ for [Google Antigravity](https://github.com/google-antigravity/antigravity-sdk-python).
22
+
23
+ Antigravity is not a Python agent loop. The SDK drives a bundled Go
24
+ `localharness` subprocess over a WebSocket, and that subprocess does real file
25
+ and shell work on the host. So this is an **ADK-class integration** — stateful
26
+ and session-based — not a stateless LangGraph-class one. The stream translation
27
+ is small; the session lifecycle is the substance.
28
+
29
+ Sessions share those subprocesses rather than owning one each — see
30
+ [Harness pooling](#harness-pooling).
31
+
32
+ ## Experimental features
33
+
34
+ These are off unless you use them, and their names and behaviour may change in
35
+ a later release:
36
+
37
+ | Feature | How to use it |
38
+ |---|---|
39
+ | Built-in `get_app_context` tool | `AntigravityAgent(experimental_app_context=True)` |
40
+ | Built-in `get_shared_state` tool | `AntigravityAgent(experimental_app_state=True)` |
41
+ | Shared state from server tools | `experimental_get_state()` and `experimental_set_state()` |
42
+ | App context from server tools | `experimental_get_context()` |
43
+ | Interrupts from your own tools | `experimental_interrupt()` |
44
+
45
+ See [What the model sees](#what-the-model-sees-app-context-and-shared-state),
46
+ [Shared state from server tools](#shared-state-from-server-tools) and
47
+ [Interrupts from your own tools](#interrupts-from-your-own-tools).
48
+
49
+ ## Why the HITL story is unusually clean
50
+
51
+ Antigravity's hooks and custom tools are **async, awaited, and carry no
52
+ timeout**. When the model needs a human, the Go harness calls into Python and
53
+ blocks on the returned coroutine. That awaited coroutine is a *native suspension
54
+ primitive*: the integration emits AG-UI events, parks an `asyncio.Future`, lets
55
+ the SSE response for run *N* close, and resolves the future from run *N+1*. The
56
+ model's tool result is the tool's actual return value — no proxy tool, no
57
+ fire-and-forget long-running-tool workaround.
58
+
59
+ This depends on the Go side not abandoning a pending hook while the stream is
60
+ closed, which the SDK source could not answer. It was verified empirically
61
+ (`tests/test_parking_gate.py`, and the live gate described below): the harness
62
+ survived **45 s and 180 s** parks with the consumer detached and resumed
63
+ correctly in both cases.
64
+
65
+ ## Install
66
+
67
+ ```bash
68
+ pip install ag-ui-antigravity
69
+ ```
70
+
71
+ `google-antigravity` ships platform-specific wheels (~32 MB) that bundle the
72
+ `localharness` binary; no separate download step is needed.
73
+
74
+ ## Usage
75
+
76
+ ```python
77
+ from ag_ui_antigravity import AntigravityAgent, create_antigravity_app
78
+
79
+ agent = AntigravityAgent(
80
+ model="gemini-3.6-flash",
81
+ api_key="...", # or GEMINI_API_KEY in the environment
82
+ system_instructions="You are a helpful assistant.",
83
+ workspaces=["/path/to/a/sandbox"],
84
+ )
85
+
86
+ app = create_antigravity_app(agent, path="/")
87
+ ```
88
+
89
+ Serve several demo agents from one process by passing a mapping:
90
+
91
+ ```python
92
+ app = create_antigravity_app({
93
+ "agentic_chat": chat_agent,
94
+ "human_in_the_loop": hitl_agent,
95
+ })
96
+ ```
97
+
98
+ ### Custom Gemini endpoints
99
+
100
+ `endpoint` sends the native path's Gemini requests somewhere other than
101
+ Google's API, such as a gateway or a mock server like
102
+ [aimock](https://github.com/CopilotKit/aimock). It takes the SDK's own endpoint
103
+ types, so headers ride along on every model call the harness makes:
104
+
105
+ ```python
106
+ from google.antigravity.types import GeminiAPIEndpoint
107
+
108
+ agent = AntigravityAgent(
109
+ model="gemini-3.6-flash",
110
+ endpoint=GeminiAPIEndpoint(
111
+ base_url="http://localhost:4010",
112
+ http_headers={"X-AIMock-Context": "my-app"},
113
+ ),
114
+ )
115
+ ```
116
+
117
+ The adapter pins both the text model and the image model to the endpoint, so no
118
+ call falls back to Google's API. The harness still requires a Gemini API key on
119
+ this path, wherever `base_url` points: pass `api_key` (or set `GEMINI_API_KEY`),
120
+ and against a mock any non-empty value works. `VertexEndpoint` works the same
121
+ way for Vertex AI.
122
+
123
+ ### Local OpenAI-compatible servers
124
+
125
+ ```python
126
+ agent = AntigravityAgent(model="gemma3", base_url="http://localhost:11434")
127
+ ```
128
+
129
+ This is the harness's path for unauthenticated local servers such as Ollama or
130
+ LM Studio. It cannot reach hosted OpenAI: see
131
+ [Known gaps](#known-gaps-in-google-antigravity-018019-openai-compatible-path).
132
+ Pass the **root** URL, not `.../v1` — the harness appends
133
+ `/v1/chat/completions` itself. (The SDK's own docstring example is misleading
134
+ on this point.) `base_url` and `endpoint` are mutually exclusive.
135
+
136
+ ## Event mapping
137
+
138
+ | Antigravity `Step` / signal | AG-UI event(s) |
139
+ |---|---|
140
+ | `run()` entry | `RUN_STARTED`, declaring `protocolVersion: "1.0"` |
141
+ | first `content_delta` on a step | `TEXT_MESSAGE_START` |
142
+ | subsequent `content_delta` | `TEXT_MESSAGE_CONTENT` (the delta, not `content`) |
143
+ | same step reaches `DONE` | `TEXT_MESSAGE_END` |
144
+ | `thinking_delta` | `REASONING_START`, `REASONING_MESSAGE_*`, `REASONING_END` (one id per span) |
145
+ | `TOOL_CALL` (built-in / MCP) | `TOOL_CALL_START` / `ARGS` / `END` / `RESULT` |
146
+ | `start_subagent` | `STEP_STARTED` / `STEP_FINISHED` around the delegated work |
147
+ | `FINISH.structured_output` | `STATE_SNAPSHOT` (or `CUSTOM`) |
148
+ | `experimental_set_state()` inside a server tool | `STATE_SNAPSHOT` |
149
+ | iterator exhaustion | `RUN_FINISHED` |
150
+ | raised `Antigravity*Error` | `RUN_ERROR` (except a mid-stream cancel, which ends `RUN_FINISHED`) |
151
+ | parked hook (question / approval) | `RUN_FINISHED` with an interrupt outcome |
152
+ | parked frontend tool | `RUN_FINISHED`, no outcome — the client replies with a `ToolMessage` |
153
+ | step with `status=ERROR` | `RUN_ERROR` |
154
+
155
+ Two details that only show up against a live harness:
156
+
157
+ * Steps whose `source` is `USER` are the harness echoing the prompt back. They
158
+ are never translated — doing so would replay the user's own message as
159
+ assistant output.
160
+ * **Failures usually arrive as a step, not an exception.** `receive_steps()`
161
+ raises only for `source=SYSTEM` errors carrying HTTP 400/401/403. Rate limits,
162
+ 5xx and model-side failures are *yielded* with `status=ERROR`, so a loop that
163
+ only catches exceptions reports them to the client as an empty success. The
164
+ run loop inspects `status` for exactly this reason.
165
+ * **Built-in tool results have no fixed key.** The harness reports a tool's
166
+ outcome by *growing* the `args` dict at DONE, under a tool-specific name
167
+ (`list_directory` adds `results`). Some tools — `view_file` — add nothing at
168
+ all, because their output goes to the model out of band. The translator
169
+ therefore takes whatever keys appeared after the call was first seen. A
170
+ failure is reported in words (`TOOL_CALL_RESULT` has no error channel), and
171
+ so is "completed with no output" — an empty string would make the two
172
+ indistinguishable and render a failed call as a successful one.
173
+ * **`TOOL_CALL_ARGS` deltas are concatenated by the client**
174
+ (`function.arguments += delta`), so everything sent for one call must join
175
+ into a single JSON document. Antigravity hands over the whole args dict each
176
+ time rather than streaming fragments, and *grows* it with the result at DONE
177
+ — and a grown JSON object is not a string extension of the smaller one. The
178
+ args are therefore sent once; the result travels on `TOOL_CALL_RESULT`.
179
+
180
+ ## Human-in-the-loop
181
+
182
+ Three cases, one primitive (emit → park a Future → resolve from the next run):
183
+
184
+ * **Frontend tools** — every `RunAgentInput.tools` entry becomes a custom async
185
+ Antigravity tool built from its JSON Schema. The client answers with a
186
+ `ToolMessage` carrying the `tool_call_id`.
187
+ * **Model questions** — `OnInteractionHook` maps `AskQuestionInteractionSpec`
188
+ onto a `RunFinishedInterruptOutcome`; the client answers via
189
+ `RunAgentInput.resume`.
190
+ * **Tool approval** (`tool_approval=True`) — `PreToolCallDecideHook` round-trips
191
+ an approval interrupt. Registering it also satisfies the SDK's mandatory
192
+ safety guard, so write and MCP tools stop raising without a separate policy.
193
+ It needs a client that implements the interrupt protocol — the dojo answers
194
+ `ToolMessage`s, not `resume` entries, so the demos leave it off.
195
+
196
+ ### Answering an interrupt
197
+
198
+ Two wire shapes are accepted, because clients disagree:
199
+
200
+ * **AG-UI `resume`** — `RunAgentInput.resume` entries carrying an
201
+ `interrupt_id`. What the dojo and the protocol itself use.
202
+ * **`forwardedProps.command`** — what CopilotKit Channels sends. Its run loop
203
+ re-enters an interrupted run with
204
+ `runAgent({ forwardedProps: { command: resume } })` and attaches no interrupt
205
+ id, because it tracks a single outstanding interrupt per thread.
206
+
207
+ For the second shape the answer is matched to an explicit id in the payload if
208
+ there is one, otherwise to the single parked *interrupt*. With several parked
209
+ and no id it is refused and logged: resolving the wrong request is
210
+ unrecoverable, a warning is not. A parked frontend tool never counts, and an
211
+ interrupt answer can never resolve one: its result is the `ToolMessage` that
212
+ carries its `tool_call_id`, and letting a bare command stand in for it would
213
+ hand the model the user's reply as the tool's return value. Supporting only
214
+ `resume` left a channel-driven approval parked forever, which looks like a bot
215
+ that has silently gone quiet.
216
+
217
+ ### Interrupts from your own tools
218
+
219
+ *Experimental.* A server tool can pause itself on an interrupt of its own design with
220
+ `experimental_interrupt()`, then carry on with the user's answer:
221
+
222
+ ```python
223
+ from ag_ui_antigravity import experimental_interrupt
224
+
225
+ async def schedule_meeting(topic: str, attendee: str) -> str:
226
+ """Books a meeting once the user picks a time."""
227
+ answer = await experimental_interrupt(
228
+ "schedule_meeting",
229
+ message=f"Pick a time for {topic}",
230
+ metadata={"topic": topic, "attendee": attendee},
231
+ )
232
+ if not answer.resolved:
233
+ return f"Not scheduled ({answer.status})."
234
+ return f"Scheduled for {answer.payload['chosen_label']}."
235
+ ```
236
+
237
+ The run ends with `RUN_FINISHED` carrying an interrupt outcome whose interrupt
238
+ has that `reason`, `message` and `metadata`, the calling tool's
239
+ `tool_call_id`, and any extra keyword fields at the top level (AG-UI's
240
+ `Interrupt` allows them). Put what the UI needs in `metadata`, though:
241
+ CopilotKit's runtime relays only the protocol's own interrupt fields, so extra
242
+ top-level ones never reach the browser. The resume payload comes back unchanged in
243
+ `answer.payload`. `answer.status` is `"resolved"`, `"cancelled"` (the user
244
+ declined) or `"abandoned"` (the user sent a new message instead); tell the model
245
+ which, so it does not report a decline that never happened. As with
246
+ `experimental_get_state()`, it only works inside a server tool.
247
+
248
+ ### One turn, several runs
249
+
250
+ An Antigravity *turn* that parks on a human spans several AG-UI *runs*, and on
251
+ each later run the harness re-delivers the steps of that turn it has already
252
+ sent. Three pieces of state are therefore scoped to the turn, not the run, and
253
+ are retired together by `AntigravitySession.reset_stream()`:
254
+
255
+ * the `receive_steps()` iterator (and any in-flight `__anext__()`, which is
256
+ parked rather than cancelled — cancelling discards the step being delivered),
257
+ * the `EventTranslator`, which records which steps and tool calls it has
258
+ already finished,
259
+ * the bridge's per-turn frontend-tool results.
260
+
261
+ Without this, every run re-translates the same tool call, the client re-executes
262
+ it, and the conversation never converges.
263
+
264
+ ### Repeated tool calls
265
+
266
+ The harness escalates a slow custom tool to a **background task** and lets the
267
+ model continue without waiting for it. The model then commonly re-issues the
268
+ call — sometimes with slightly different arguments — which would make the client
269
+ run a side-effecting action a second time.
270
+
271
+ So a frontend tool is dispatched to the client **at most once per turn**. An
272
+ identical repeat gets the cached result; a repeat with different arguments gets
273
+ a plain statement of what already ran, so the model reports the result instead
274
+ of retrying. Set `deduplicate_tool_calls=False` if a tool is genuinely meant to
275
+ run repeatedly within one turn.
276
+
277
+ ### Server-side tools
278
+
279
+ Pass your own Python callables as `tools=[...]` and they run in this process,
280
+ with the call and its result streamed to the client — that is what the dojo's
281
+ `backend_tool_rendering` demo draws its weather card from.
282
+
283
+ The adapter emits those events itself rather than reading them off the step
284
+ stream, because the harness reports a custom tool as a **single**
285
+ `TOOL_CALL`/`ACTIVE` step: there is no DONE step, and `Step` carries no result
286
+ field at all, since the return value goes back over the WebSocket straight to
287
+ the model. A client waiting for a `TOOL_CALL_RESULT` from the step stream would
288
+ wait forever. Built-in tools are different — the harness re-reports those at
289
+ DONE with their output folded into the call arguments.
290
+
291
+ The wrapper preserves each function's signature and docstring, so the SDK still
292
+ derives the same tool schema. Return a JSON-serializable value (or a string);
293
+ a raised exception is reported to the client as
294
+ `There was an error executing <tool>: ...` and re-raised.
295
+
296
+ ### Shared state from server tools
297
+
298
+ *Experimental.* A server tool can read and write the AG-UI shared state of the session it runs
299
+ in:
300
+
301
+ ```python
302
+ from ag_ui_antigravity import experimental_get_state, experimental_set_state
303
+
304
+ async def research_agent(task: str) -> str:
305
+ """Delegates a research task."""
306
+ facts = await run_research(task)
307
+ state = experimental_get_state()
308
+ delegations = state.get("delegations", []) + [facts]
309
+ experimental_set_state({**state, "delegations": delegations})
310
+ return facts
311
+ ```
312
+
313
+ `experimental_get_state()` returns a copy of the session's state: what the
314
+ client sent with the run (`RunAgentInput.state`), or what the last
315
+ `experimental_set_state()` stored if the client has not received that yet.
316
+ `experimental_set_state()` replaces the whole state and emits a `STATE_SNAPSHOT` at once, before the tool's `TOOL_CALL_RESULT`, so a UI
317
+ bound to agent state updates while the turn is still running. The state must be
318
+ a JSON-serializable dict; anything else raises `TypeError` in the tool.
319
+
320
+ Both raise `RuntimeError` outside a server tool. The session is found through a
321
+ context variable set for the duration of the call, not through a tool
322
+ parameter, because the SDK would put such a parameter into the tool's schema.
323
+
324
+ A write made while the turn is parked (no run attached) is queued and delivered
325
+ on the next run. Until then the adapter keeps its own copy rather than the
326
+ client's, since the client's copy predates the write.
327
+
328
+ `experimental_get_context()` works the same way and returns the run's
329
+ `RunAgentInput.context` (what CopilotKit's `useAgentContext` shares) as a list of
330
+ `{"description", "value"}` entries.
331
+
332
+ ### What the model sees: app context and shared state
333
+
334
+ *Experimental, off by default.* Antigravity fixes an agent's instructions when
335
+ the harness session starts, so per-run input cannot be folded into the prompt
336
+ the way other integrations do. Without help, the model never sees the run's
337
+ context or shared state. The adapter can give an agent two read-only tools the
338
+ model calls when it needs them:
339
+
340
+ | Tool | Returns | Turn it on with |
341
+ |---|---|---|
342
+ | `get_app_context` | the run's `RunAgentInput.context` | `experimental_app_context=True` |
343
+ | `get_shared_state` | the session's shared state, including the user's edits in the UI | `experimental_app_state=True` |
344
+
345
+ ```python
346
+ agent = AntigravityAgent(
347
+ model="gemini-2.5-flash",
348
+ experimental_app_context=True,
349
+ experimental_app_state=True,
350
+ system_instructions=(
351
+ "Call get_app_context and get_shared_state before answering anything "
352
+ "about the user or the app."
353
+ ),
354
+ )
355
+ ```
356
+
357
+ Both run silently: the client gets no `TOOL_CALL_*` events for them, so no tool
358
+ card appears for what is, elsewhere, an invisible prompt update. Their
359
+ docstrings tell the model to call them whenever the answer may depend on the
360
+ user, the page or the app state; say so in `system_instructions` too when it
361
+ matters. A server or client tool with the same name takes precedence.
362
+
363
+ The difference from prompt injection is that the model has to ask. A preference
364
+ the user changed in the UI reaches the model on its next call to
365
+ `get_shared_state`, not before.
366
+
367
+ ### Attachments
368
+
369
+ Image, document, audio and video parts of a user message reach the model when
370
+ their bytes travel inline: a `data` source, or a `data:` URL. They become the
371
+ SDK's `Image`/`Document`/`Audio`/`Video` objects, sent in order with the text.
372
+ The SDK accepts PNG, JPEG, WebP and BMP images and PDF, plain-text, CSV, JSON,
373
+ HTML and XML documents, among others.
374
+
375
+ The harness cannot fetch anything itself, so an `https://` URL, a provider file
376
+ reference or an unsupported type such as GIF is replaced by a one-line note in
377
+ the prompt (`[Attached image 'x.png' was not forwarded: ...]`). The model then
378
+ says it could not see the file rather than answering as if nothing was attached.
379
+ `/capabilities` advertises `multimodal.input` accordingly.
380
+
381
+ ### Built-in tools worth disabling
382
+
383
+ The harness exposes its whole built-in toolset by default. `search_web` returns
384
+ an *empty* summary unless the harness has Google credentials — the model then
385
+ retries it indefinitely and the conversation never settles. Pass a
386
+ `CapabilitiesConfig(enabled_tools=[...])` naming only what the agent needs;
387
+ `BuiltinTools.FINISH` must stay, since the harness uses it to end a turn. The
388
+ chat demos in `examples/` enable nothing else.
389
+
390
+ ## Sessions
391
+
392
+ `SessionManager` keys a live `Conversation` by `thread_id`.
393
+
394
+ * **Hot resume** — a session with a parked coroutine stays in memory, because a
395
+ suspended coroutine cannot be serialized. It gets a longer grace period than
396
+ an idle session rather than an exemption: `parked_timeout_seconds` (2 h)
397
+ instead of `session_timeout_seconds` (30 min). A run in flight holds the
398
+ session lock and is never reclaimed.
399
+
400
+ A session occupies memory for its whole life, not only while parked — from
401
+ the first message on a `thread_id` until it times out, is evicted at
402
+ `max_sessions`, or is rebuilt because the client's tool contract changed —
403
+ a tool's name, description *or* parameter schema, since Antigravity fixes
404
+ the tool configuration when it connects. Parking does not allocate anything;
405
+ it extends how long the allocation is held.
406
+ * **Cold resume** — a recycled session with nothing parked is rebuilt from
407
+ `conversation_id` + `session_continuation_mode` + `save_dir`. This covers a
408
+ thread that comes back **after** its session was swept, not only one rebuilt
409
+ in place: the manager keeps `thread_id -> (conversation_id, forwarded
410
+ prompts)` for closed sessions, so a user returning past the idle timeout
411
+ continues where they left off rather than meeting an agent with amnesia while
412
+ their history sits unreachable in `save_dir`.
413
+
414
+ That map is currently unbounded — one short string and a small set per thread
415
+ the process has ever seen. Cap it (LRU or TTL) before running at a scale where
416
+ that matters.
417
+
418
+ Because Antigravity fixes the tool list in the harness config at connect time,
419
+ a client that changes its `tools` between runs forces a cold-resume rebuild
420
+ rather than running against a stale list.
421
+
422
+ ## Harness pooling
423
+
424
+ Sessions do **not** get a subprocess each. A `HarnessPool` shares one
425
+ `localharness` process between up to `max_conversations_per_process` (default 8)
426
+ conversations, because Antigravity configures a harness twice:
427
+
428
+ | sent | when | contains |
429
+ |---|---|---|
430
+ | `InputConfig` | process stdin at startup | `save_dir`, `env` |
431
+ | `HarnessConfig` | WebSocket, **per conversation** | tools, model, system instructions, capabilities, MCP servers, hooks, subagents, `response_schema`, `conversation_id`, **`workspaces`** |
432
+
433
+ Almost everything varying per thread — including `workspaces`, so per-thread
434
+ filesystem isolation is unaffected — is per-conversation. Only `save_dir` and
435
+ `env` are process-wide, and they form the pool's partition key.
436
+
437
+ Measured with 8 concurrent conversations, one turn each:
438
+
439
+ | | pooled (1 process) | one process each |
440
+ |---|---|---|
441
+ | idle | **101 MB** | 752 MB |
442
+ | mid-turn | **154 MB** | 1040 MB |
443
+ | wall clock | 12.7 s | 12.3 s |
444
+
445
+ So an extra idle conversation costs ~1 MB rather than ~95 MB. Throughput is
446
+ unchanged; per-turn p50 rises ~1.3× only when all 8 turn simultaneously. A
447
+ 20-second tool call in one conversation was measured **not** to delay its
448
+ neighbours (median inflation 0.95× against a control).
449
+
450
+ Pooling is configured on the agent:
451
+
452
+ | Option | Default | Effect |
453
+ |---|---|---|
454
+ | `max_conversations_per_process` | `8` | conversations sharing one harness process; `1` gives every conversation its own process |
455
+ | `harness_idle_grace_seconds` | `30.0` | how long an empty process stays up before it is stopped |
456
+ | `harness_pool` | a pool per agent | pass one `HarnessPool` to several agents to share processes between them (see below) |
457
+
458
+ Pooled conversations stay isolated at the harness level (each has its own
459
+ workspaces, tools and instructions), but they share one OS process. If your
460
+ deployment needs process-level isolation between users, set
461
+ `max_conversations_per_process=1`.
462
+
463
+ ### Several agents in one server
464
+
465
+ Each `AntigravityAgent` owns a pool, so a server hosting four agents gets four
466
+ harness processes — a ~95 MB floor per agent, independent of traffic. It does
467
+ not affect scaling (threads of one agent still share), and it affects nothing
468
+ about correctness, but it is silent: you find out by counting processes.
469
+
470
+ Sharing needs **both** a pool and a `save_dir`. Passing only `harness_pool=`
471
+ changes nothing, because each agent otherwise mints its own `tempfile.mkdtemp()`
472
+ save directory and `save_dir` is half the pool's partition key:
473
+
474
+ ```python
475
+ from ag_ui_antigravity.harness_pool import HarnessPool
476
+
477
+ pool = HarnessPool()
478
+ save_dir = "/var/lib/myapp/antigravity" # both, or you still get a process each
479
+
480
+ chat = AntigravityAgent(model=..., harness_pool=pool, save_dir=save_dir)
481
+ research = AntigravityAgent(model=..., harness_pool=pool, save_dir=save_dir)
482
+ ```
483
+
484
+ Agents that must not share process-level storage should keep separate
485
+ `save_dir` values — they will land on separate processes by design. See
486
+ `examples/server/api/_common.py`.
487
+
488
+ * **Blast radius.** One dead process fails every conversation on it. They raise
489
+ promptly rather than hanging (`test_process_death_raises_rather_than_hangs`),
490
+ but this is the tradeoff pooling buys — hence the modest default.
491
+ * **Parked sessions.** Parking and pooling compose: a conversation parked on a
492
+ human does not block its co-tenants
493
+ (`test_a_parked_conversation_does_not_block_its_siblings`). But a parked
494
+ session pins its process, and the pool cannot know in advance which
495
+ conversations will park, so under park-heavy load the saving is bounded by
496
+ fragmentation rather than by the ~1 MB marginal figure.
497
+ * **Not a tenancy boundary.** `save_dir` is shared by every conversation on a
498
+ process. Use `workspaces` for isolation, and give tenants separate `save_dir`
499
+ values (which partitions them onto separate processes) if storage must be
500
+ isolated too.
501
+
502
+ ## Operational notes
503
+
504
+ * **Sandboxing.** Real filesystem and shell access, scoped per conversation by
505
+ `workspaces`. Multi-tenant hosting needs per-thread workspace isolation,
506
+ resource caps, and reliable cleanup. Always set `workspaces=[...]`.
507
+ * **SDK churn.** `google-antigravity` is young; the dependency is pinned to
508
+ `<0.2.0` deliberately.
509
+
510
+ ### A crashed harness loses its history
511
+
512
+ A conversation is pinned to one harness process for its whole life -- there is
513
+ no migration -- so if that process dies the session is gone with it. The run
514
+ that meets the corpse reports `RUN_ERROR`, and the next run on that thread
515
+ rebuilds the session automatically and carries on.
516
+
517
+ What does not survive is the conversation history. The harness writes
518
+ trajectories through SQLite's WAL and reopens them with `immutable=1`, which
519
+ ignores WAL files, so a killed process leaves everything uncheckpointed:
520
+ measured after a `SIGKILL`, the main database showed **0 steps while its
521
+ 465 KB WAL held all 4**. The resume therefore sees an empty conversation and
522
+ `CREATE_OR_RESUME` quietly starts a new one. This is upstream and not
523
+ something the integration can work around; a clean shutdown checkpoints
524
+ normally and resumes fine.
525
+
526
+ The rebuild logs a warning naming the thread, so silent context loss is at
527
+ least visible in the logs.
528
+
529
+ ### Keep workspace paths short
530
+
531
+ A long workspace path makes runs fail intermittently, and the failure looks
532
+ nothing like its cause.
533
+
534
+ If a prompt leads the model to write an absolute path into a tool call — "read
535
+ `/var/folders/0t/pq2_7rn97834lcsvc_qy4t8r0000gn/T/ag-ui-antigravity-wz_4xjbd/notes.txt`"
536
+ — it sometimes reproduces that path wrongly, truncating it or repeating a chunk
537
+ of it. macOS temp directories are ~75 characters of high-entropy text, which is
538
+ about the worst case. Measured: **0/14 runs failed with a
539
+ 9-character workspace path, 2/14 with a 75-character one.**
540
+
541
+ The harness treats the resulting bad path as a fatal
542
+ `AntigravityExecutionError` rather than returning the error for the model to
543
+ retry, so the whole run dies mid-tool-call with
544
+ `RUN_ERROR: The model produced an invalid tool call`. Nothing in that message
545
+ points at path length.
546
+
547
+ Set `ANTIGRAVITY_WORKSPACE` (or `workspaces=[...]`) to something short and
548
+ stable — `/srv/agents/w1`, not a generated temp directory. This is a model
549
+ limitation rather than an integration bug: it reproduces identically on older
550
+ commits.
551
+
552
+ ### Known gaps in `google-antigravity` 0.1.8–0.1.9 (OpenAI-compatible path)
553
+
554
+ These are upstream, not integration bugs. They affect only `base_url` usage:
555
+
556
+ 1. **No API-key field.** `GemmaEndpoint` carries only `base_url`, and the Go
557
+ harness reads no `OPENAI_API_KEY`. The path targets unauthenticated local
558
+ servers (Ollama, LM Studio), so hosted OpenAI rejects every request.
559
+ 2. **Gemini-shaped tool schemas.** Custom-tool schemas are generated with
560
+ `api_option="GEMINI_API"`, emitting proto-style uppercase types (`"STRING"`)
561
+ that OpenAI rejects. Tools registered via `ToolWithSchema` — which is how
562
+ this integration builds *frontend* tools — pass their schema through
563
+ untouched and are unaffected.
564
+ 3. **`session_continuation_mode` dropped.** `LocalOpenAIAgentConfig.create_strategy`
565
+ does not forward it, disabling cold resume. Worked around by
566
+ `_ResumableOpenAIConfig` in `agent.py`.
567
+
568
+ ### Tool calls do not stream their arguments
569
+
570
+ Not specific to the OpenAI path, and not workable around: the harness hands over
571
+ a tool call **fully formed, in a single step**. Measured with a custom tool
572
+ whose argument was 1850 characters — still one `TOOL_CALL` step carrying the
573
+ complete arguments. Presumably the Go side parses the model's tool-call JSON
574
+ before dispatching, since it needs valid JSON to invoke anything, and surfaces
575
+ the parsed call rather than the token stream.
576
+
577
+ So `TOOL_CALL_ARGS` arrives as one delta rather than filling in progressively.
578
+ There are no fragments to forward; this would need the harness to expose partial
579
+ tool-call deltas. What *does* stream is the call lifecycle: the call is emitted
580
+ as soon as the harness dispatches it, so a client renders its pending state
581
+ while the tool runs and swaps in the result when it returns.
582
+
583
+ ## Not implemented yet
584
+
585
+ Deliberate gaps, so the surface above is not mistaken for more than it is:
586
+
587
+ * **Triggers** (async inbound messages) — the SDK supports them; nothing here
588
+ maps them to AG-UI yet.
589
+ * **`STATE_DELTA`** — structured output and `experimental_set_state()` are emitted as whole
590
+ snapshots only.
591
+ * **`forwardedProps`** — apart from interrupt answers, not passed to the model
592
+ or to tools.
593
+ * **MCP servers** — passed through to the SDK config and covered by the
594
+ approval hook, but not exercised by a live test.
595
+ * **`predictive_state_updates`** dojo feature — it streams a tool's
596
+ arguments into state while the model writes them, and the harness hands over
597
+ tool calls whole. (`shared_state` is in the menu, backed by
598
+ `examples/server/api/shared_state.py`.)
599
+ * Subagent bracketing is unit-tested against recorded step shapes, not against
600
+ a live multi-agent run.
601
+
602
+ ## Development
603
+
604
+ ```bash
605
+ uv sync
606
+ uv run pytest # 315 unit tests; live tests are deselected by default
607
+ ```
608
+
609
+ The live checks start a real harness subprocess and call Gemini:
610
+
611
+ ```bash
612
+ export GEMINI_API_KEY=...
613
+ uv run pytest tests/ -m live
614
+ ```
615
+
616
+ `ANTIGRAVITY_TEST_MODEL` picks the model (default: the SDK's default model), and
617
+ `GOOGLE_GEMINI_BASE_URL` points the same tests at a Gemini-compatible gateway.
618
+
619
+ `tests/test_parking_gate.py` is the important one — it re-verifies the Go-side
620
+ no-timeout property the whole HITL design depends on. Raise the park duration
621
+ to reproduce the long soak:
622
+
623
+ ```bash
624
+ PARK_SECONDS=180 uv run pytest tests/test_parking_gate.py -m live
625
+ ```
626
+
627
+ ### Dojo
628
+
629
+ ```bash
630
+ # terminal 1
631
+ cd examples
632
+ GEMINI_API_KEY=... uv run dev # serves on :8027
633
+
634
+ # terminal 2
635
+ cd apps/dojo && pnpm dev
636
+ ```
637
+
638
+ `pnpm run-dojo-everything` starts it alongside the other integrations.
639
+
640
+ To replay aimock fixtures instead of calling Gemini, point the server at aimock.
641
+ The harness insists on a key, but aimock ignores its value:
642
+ `GOOGLE_GEMINI_BASE_URL=http://localhost:4010 AIMOCK_CONTEXT=<fixture-set> GEMINI_API_KEY=unused uv run dev`.
643
+
644
+ Then open `/antigravity/feature/agentic_chat`.
645
+
646
+ ## Verification status
647
+
648
+ Verified on 2026-09-30 against `google-antigravity` 0.1.9 on the native Gemini
649
+ path, with no proxy in between:
650
+
651
+ | Check | Result |
652
+ |---|---|
653
+ | 315 unit tests (translator, bridge, sessions, endpoint, config) | pass |
654
+ | 19 live tests against Gemini (streaming, multi-turn, frontend-tool park/resume, built-in tools, SSE, server tools, cold resume, pooling) | 15 pass; 4 were cut short by the free-tier key's quota (429) or Gemini overload (503) errors |
655
+ | Harness parked with the stream closed, then resumed | pass in the CopilotKit showcase's human-in-the-loop cells; the 30 s `test_parking_gate.py` soak ran into the quota limit |
656
+ | `endpoint=GeminiAPIEndpoint(...)` against aimock: text, server-tool round trip, reasoning | pass (reasoning streams as `REASONING_*` events) |
657
+ | Example server against aimock via `GOOGLE_GEMINI_BASE_URL` | pass |
658
+ | `/capabilities` payload against `AgentCapabilitiesSchema` (zod) | valid |