ag_ui_antigravity 0.1.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- ag_ui_antigravity-0.1.0/LICENSE +21 -0
- ag_ui_antigravity-0.1.0/PKG-INFO +658 -0
- ag_ui_antigravity-0.1.0/README.md +641 -0
- ag_ui_antigravity-0.1.0/pyproject.toml +40 -0
- ag_ui_antigravity-0.1.0/pyproject.toml.orig +45 -0
- ag_ui_antigravity-0.1.0/src/ag_ui_antigravity/__init__.py +36 -0
- ag_ui_antigravity-0.1.0/src/ag_ui_antigravity/agent.py +1072 -0
- ag_ui_antigravity-0.1.0/src/ag_ui_antigravity/builtin_tools.py +38 -0
- ag_ui_antigravity-0.1.0/src/ag_ui_antigravity/endpoint.py +190 -0
- ag_ui_antigravity-0.1.0/src/ag_ui_antigravity/event_translator.py +584 -0
- ag_ui_antigravity-0.1.0/src/ag_ui_antigravity/harness_pool.py +688 -0
- ag_ui_antigravity-0.1.0/src/ag_ui_antigravity/session_manager.py +497 -0
- ag_ui_antigravity-0.1.0/src/ag_ui_antigravity/ui_bridge.py +881 -0
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2025
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
|
@@ -0,0 +1,658 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: ag_ui_antigravity
|
|
3
|
+
Version: 0.1.0
|
|
4
|
+
Summary: Google Antigravity integration for the AG-UI Protocol
|
|
5
|
+
License-Expression: MIT
|
|
6
|
+
License-File: LICENSE
|
|
7
|
+
Requires-Dist: ag-ui-protocol>=1.0.0,<2.0
|
|
8
|
+
Requires-Dist: google-antigravity>=0.1.8,<0.2.0
|
|
9
|
+
Requires-Dist: fastapi>=0.115.2
|
|
10
|
+
Requires-Dist: pydantic>=2.11.7
|
|
11
|
+
Requires-Dist: sse-starlette>=2.1.0
|
|
12
|
+
Requires-Dist: uvicorn>=0.35.0
|
|
13
|
+
Requires-Python: >=3.10, <3.15
|
|
14
|
+
Project-URL: Homepage, https://github.com/ag-ui-protocol/ag-ui/tree/main/integrations/antigravity/python
|
|
15
|
+
Project-URL: Issues, https://github.com/ag-ui-protocol/ag-ui/issues
|
|
16
|
+
Description-Content-Type: text/markdown
|
|
17
|
+
|
|
18
|
+
# AG-UI ⚡ Google Antigravity
|
|
19
|
+
|
|
20
|
+
Implementation of the [AG-UI protocol](https://github.com/ag-ui-protocol/ag-ui)
|
|
21
|
+
for [Google Antigravity](https://github.com/google-antigravity/antigravity-sdk-python).
|
|
22
|
+
|
|
23
|
+
Antigravity is not a Python agent loop. The SDK drives a bundled Go
|
|
24
|
+
`localharness` subprocess over a WebSocket, and that subprocess does real file
|
|
25
|
+
and shell work on the host. So this is an **ADK-class integration** — stateful
|
|
26
|
+
and session-based — not a stateless LangGraph-class one. The stream translation
|
|
27
|
+
is small; the session lifecycle is the substance.
|
|
28
|
+
|
|
29
|
+
Sessions share those subprocesses rather than owning one each — see
|
|
30
|
+
[Harness pooling](#harness-pooling).
|
|
31
|
+
|
|
32
|
+
## Experimental features
|
|
33
|
+
|
|
34
|
+
These are off unless you use them, and their names and behaviour may change in
|
|
35
|
+
a later release:
|
|
36
|
+
|
|
37
|
+
| Feature | How to use it |
|
|
38
|
+
|---|---|
|
|
39
|
+
| Built-in `get_app_context` tool | `AntigravityAgent(experimental_app_context=True)` |
|
|
40
|
+
| Built-in `get_shared_state` tool | `AntigravityAgent(experimental_app_state=True)` |
|
|
41
|
+
| Shared state from server tools | `experimental_get_state()` and `experimental_set_state()` |
|
|
42
|
+
| App context from server tools | `experimental_get_context()` |
|
|
43
|
+
| Interrupts from your own tools | `experimental_interrupt()` |
|
|
44
|
+
|
|
45
|
+
See [What the model sees](#what-the-model-sees-app-context-and-shared-state),
|
|
46
|
+
[Shared state from server tools](#shared-state-from-server-tools) and
|
|
47
|
+
[Interrupts from your own tools](#interrupts-from-your-own-tools).
|
|
48
|
+
|
|
49
|
+
## Why the HITL story is unusually clean
|
|
50
|
+
|
|
51
|
+
Antigravity's hooks and custom tools are **async, awaited, and carry no
|
|
52
|
+
timeout**. When the model needs a human, the Go harness calls into Python and
|
|
53
|
+
blocks on the returned coroutine. That awaited coroutine is a *native suspension
|
|
54
|
+
primitive*: the integration emits AG-UI events, parks an `asyncio.Future`, lets
|
|
55
|
+
the SSE response for run *N* close, and resolves the future from run *N+1*. The
|
|
56
|
+
model's tool result is the tool's actual return value — no proxy tool, no
|
|
57
|
+
fire-and-forget long-running-tool workaround.
|
|
58
|
+
|
|
59
|
+
This depends on the Go side not abandoning a pending hook while the stream is
|
|
60
|
+
closed, which the SDK source could not answer. It was verified empirically
|
|
61
|
+
(`tests/test_parking_gate.py`, and the live gate described below): the harness
|
|
62
|
+
survived **45 s and 180 s** parks with the consumer detached and resumed
|
|
63
|
+
correctly in both cases.
|
|
64
|
+
|
|
65
|
+
## Install
|
|
66
|
+
|
|
67
|
+
```bash
|
|
68
|
+
pip install ag-ui-antigravity
|
|
69
|
+
```
|
|
70
|
+
|
|
71
|
+
`google-antigravity` ships platform-specific wheels (~32 MB) that bundle the
|
|
72
|
+
`localharness` binary; no separate download step is needed.
|
|
73
|
+
|
|
74
|
+
## Usage
|
|
75
|
+
|
|
76
|
+
```python
|
|
77
|
+
from ag_ui_antigravity import AntigravityAgent, create_antigravity_app
|
|
78
|
+
|
|
79
|
+
agent = AntigravityAgent(
|
|
80
|
+
model="gemini-3.6-flash",
|
|
81
|
+
api_key="...", # or GEMINI_API_KEY in the environment
|
|
82
|
+
system_instructions="You are a helpful assistant.",
|
|
83
|
+
workspaces=["/path/to/a/sandbox"],
|
|
84
|
+
)
|
|
85
|
+
|
|
86
|
+
app = create_antigravity_app(agent, path="/")
|
|
87
|
+
```
|
|
88
|
+
|
|
89
|
+
Serve several demo agents from one process by passing a mapping:
|
|
90
|
+
|
|
91
|
+
```python
|
|
92
|
+
app = create_antigravity_app({
|
|
93
|
+
"agentic_chat": chat_agent,
|
|
94
|
+
"human_in_the_loop": hitl_agent,
|
|
95
|
+
})
|
|
96
|
+
```
|
|
97
|
+
|
|
98
|
+
### Custom Gemini endpoints
|
|
99
|
+
|
|
100
|
+
`endpoint` sends the native path's Gemini requests somewhere other than
|
|
101
|
+
Google's API, such as a gateway or a mock server like
|
|
102
|
+
[aimock](https://github.com/CopilotKit/aimock). It takes the SDK's own endpoint
|
|
103
|
+
types, so headers ride along on every model call the harness makes:
|
|
104
|
+
|
|
105
|
+
```python
|
|
106
|
+
from google.antigravity.types import GeminiAPIEndpoint
|
|
107
|
+
|
|
108
|
+
agent = AntigravityAgent(
|
|
109
|
+
model="gemini-3.6-flash",
|
|
110
|
+
endpoint=GeminiAPIEndpoint(
|
|
111
|
+
base_url="http://localhost:4010",
|
|
112
|
+
http_headers={"X-AIMock-Context": "my-app"},
|
|
113
|
+
),
|
|
114
|
+
)
|
|
115
|
+
```
|
|
116
|
+
|
|
117
|
+
The adapter pins both the text model and the image model to the endpoint, so no
|
|
118
|
+
call falls back to Google's API. The harness still requires a Gemini API key on
|
|
119
|
+
this path, wherever `base_url` points: pass `api_key` (or set `GEMINI_API_KEY`),
|
|
120
|
+
and against a mock any non-empty value works. `VertexEndpoint` works the same
|
|
121
|
+
way for Vertex AI.
|
|
122
|
+
|
|
123
|
+
### Local OpenAI-compatible servers
|
|
124
|
+
|
|
125
|
+
```python
|
|
126
|
+
agent = AntigravityAgent(model="gemma3", base_url="http://localhost:11434")
|
|
127
|
+
```
|
|
128
|
+
|
|
129
|
+
This is the harness's path for unauthenticated local servers such as Ollama or
|
|
130
|
+
LM Studio. It cannot reach hosted OpenAI: see
|
|
131
|
+
[Known gaps](#known-gaps-in-google-antigravity-018019-openai-compatible-path).
|
|
132
|
+
Pass the **root** URL, not `.../v1` — the harness appends
|
|
133
|
+
`/v1/chat/completions` itself. (The SDK's own docstring example is misleading
|
|
134
|
+
on this point.) `base_url` and `endpoint` are mutually exclusive.
|
|
135
|
+
|
|
136
|
+
## Event mapping
|
|
137
|
+
|
|
138
|
+
| Antigravity `Step` / signal | AG-UI event(s) |
|
|
139
|
+
|---|---|
|
|
140
|
+
| `run()` entry | `RUN_STARTED`, declaring `protocolVersion: "1.0"` |
|
|
141
|
+
| first `content_delta` on a step | `TEXT_MESSAGE_START` |
|
|
142
|
+
| subsequent `content_delta` | `TEXT_MESSAGE_CONTENT` (the delta, not `content`) |
|
|
143
|
+
| same step reaches `DONE` | `TEXT_MESSAGE_END` |
|
|
144
|
+
| `thinking_delta` | `REASONING_START`, `REASONING_MESSAGE_*`, `REASONING_END` (one id per span) |
|
|
145
|
+
| `TOOL_CALL` (built-in / MCP) | `TOOL_CALL_START` / `ARGS` / `END` / `RESULT` |
|
|
146
|
+
| `start_subagent` | `STEP_STARTED` / `STEP_FINISHED` around the delegated work |
|
|
147
|
+
| `FINISH.structured_output` | `STATE_SNAPSHOT` (or `CUSTOM`) |
|
|
148
|
+
| `experimental_set_state()` inside a server tool | `STATE_SNAPSHOT` |
|
|
149
|
+
| iterator exhaustion | `RUN_FINISHED` |
|
|
150
|
+
| raised `Antigravity*Error` | `RUN_ERROR` (except a mid-stream cancel, which ends `RUN_FINISHED`) |
|
|
151
|
+
| parked hook (question / approval) | `RUN_FINISHED` with an interrupt outcome |
|
|
152
|
+
| parked frontend tool | `RUN_FINISHED`, no outcome — the client replies with a `ToolMessage` |
|
|
153
|
+
| step with `status=ERROR` | `RUN_ERROR` |
|
|
154
|
+
|
|
155
|
+
Two details that only show up against a live harness:
|
|
156
|
+
|
|
157
|
+
* Steps whose `source` is `USER` are the harness echoing the prompt back. They
|
|
158
|
+
are never translated — doing so would replay the user's own message as
|
|
159
|
+
assistant output.
|
|
160
|
+
* **Failures usually arrive as a step, not an exception.** `receive_steps()`
|
|
161
|
+
raises only for `source=SYSTEM` errors carrying HTTP 400/401/403. Rate limits,
|
|
162
|
+
5xx and model-side failures are *yielded* with `status=ERROR`, so a loop that
|
|
163
|
+
only catches exceptions reports them to the client as an empty success. The
|
|
164
|
+
run loop inspects `status` for exactly this reason.
|
|
165
|
+
* **Built-in tool results have no fixed key.** The harness reports a tool's
|
|
166
|
+
outcome by *growing* the `args` dict at DONE, under a tool-specific name
|
|
167
|
+
(`list_directory` adds `results`). Some tools — `view_file` — add nothing at
|
|
168
|
+
all, because their output goes to the model out of band. The translator
|
|
169
|
+
therefore takes whatever keys appeared after the call was first seen. A
|
|
170
|
+
failure is reported in words (`TOOL_CALL_RESULT` has no error channel), and
|
|
171
|
+
so is "completed with no output" — an empty string would make the two
|
|
172
|
+
indistinguishable and render a failed call as a successful one.
|
|
173
|
+
* **`TOOL_CALL_ARGS` deltas are concatenated by the client**
|
|
174
|
+
(`function.arguments += delta`), so everything sent for one call must join
|
|
175
|
+
into a single JSON document. Antigravity hands over the whole args dict each
|
|
176
|
+
time rather than streaming fragments, and *grows* it with the result at DONE
|
|
177
|
+
— and a grown JSON object is not a string extension of the smaller one. The
|
|
178
|
+
args are therefore sent once; the result travels on `TOOL_CALL_RESULT`.
|
|
179
|
+
|
|
180
|
+
## Human-in-the-loop
|
|
181
|
+
|
|
182
|
+
Three cases, one primitive (emit → park a Future → resolve from the next run):
|
|
183
|
+
|
|
184
|
+
* **Frontend tools** — every `RunAgentInput.tools` entry becomes a custom async
|
|
185
|
+
Antigravity tool built from its JSON Schema. The client answers with a
|
|
186
|
+
`ToolMessage` carrying the `tool_call_id`.
|
|
187
|
+
* **Model questions** — `OnInteractionHook` maps `AskQuestionInteractionSpec`
|
|
188
|
+
onto a `RunFinishedInterruptOutcome`; the client answers via
|
|
189
|
+
`RunAgentInput.resume`.
|
|
190
|
+
* **Tool approval** (`tool_approval=True`) — `PreToolCallDecideHook` round-trips
|
|
191
|
+
an approval interrupt. Registering it also satisfies the SDK's mandatory
|
|
192
|
+
safety guard, so write and MCP tools stop raising without a separate policy.
|
|
193
|
+
It needs a client that implements the interrupt protocol — the dojo answers
|
|
194
|
+
`ToolMessage`s, not `resume` entries, so the demos leave it off.
|
|
195
|
+
|
|
196
|
+
### Answering an interrupt
|
|
197
|
+
|
|
198
|
+
Two wire shapes are accepted, because clients disagree:
|
|
199
|
+
|
|
200
|
+
* **AG-UI `resume`** — `RunAgentInput.resume` entries carrying an
|
|
201
|
+
`interrupt_id`. What the dojo and the protocol itself use.
|
|
202
|
+
* **`forwardedProps.command`** — what CopilotKit Channels sends. Its run loop
|
|
203
|
+
re-enters an interrupted run with
|
|
204
|
+
`runAgent({ forwardedProps: { command: resume } })` and attaches no interrupt
|
|
205
|
+
id, because it tracks a single outstanding interrupt per thread.
|
|
206
|
+
|
|
207
|
+
For the second shape the answer is matched to an explicit id in the payload if
|
|
208
|
+
there is one, otherwise to the single parked *interrupt*. With several parked
|
|
209
|
+
and no id it is refused and logged: resolving the wrong request is
|
|
210
|
+
unrecoverable, a warning is not. A parked frontend tool never counts, and an
|
|
211
|
+
interrupt answer can never resolve one: its result is the `ToolMessage` that
|
|
212
|
+
carries its `tool_call_id`, and letting a bare command stand in for it would
|
|
213
|
+
hand the model the user's reply as the tool's return value. Supporting only
|
|
214
|
+
`resume` left a channel-driven approval parked forever, which looks like a bot
|
|
215
|
+
that has silently gone quiet.
|
|
216
|
+
|
|
217
|
+
### Interrupts from your own tools
|
|
218
|
+
|
|
219
|
+
*Experimental.* A server tool can pause itself on an interrupt of its own design with
|
|
220
|
+
`experimental_interrupt()`, then carry on with the user's answer:
|
|
221
|
+
|
|
222
|
+
```python
|
|
223
|
+
from ag_ui_antigravity import experimental_interrupt
|
|
224
|
+
|
|
225
|
+
async def schedule_meeting(topic: str, attendee: str) -> str:
|
|
226
|
+
"""Books a meeting once the user picks a time."""
|
|
227
|
+
answer = await experimental_interrupt(
|
|
228
|
+
"schedule_meeting",
|
|
229
|
+
message=f"Pick a time for {topic}",
|
|
230
|
+
metadata={"topic": topic, "attendee": attendee},
|
|
231
|
+
)
|
|
232
|
+
if not answer.resolved:
|
|
233
|
+
return f"Not scheduled ({answer.status})."
|
|
234
|
+
return f"Scheduled for {answer.payload['chosen_label']}."
|
|
235
|
+
```
|
|
236
|
+
|
|
237
|
+
The run ends with `RUN_FINISHED` carrying an interrupt outcome whose interrupt
|
|
238
|
+
has that `reason`, `message` and `metadata`, the calling tool's
|
|
239
|
+
`tool_call_id`, and any extra keyword fields at the top level (AG-UI's
|
|
240
|
+
`Interrupt` allows them). Put what the UI needs in `metadata`, though:
|
|
241
|
+
CopilotKit's runtime relays only the protocol's own interrupt fields, so extra
|
|
242
|
+
top-level ones never reach the browser. The resume payload comes back unchanged in
|
|
243
|
+
`answer.payload`. `answer.status` is `"resolved"`, `"cancelled"` (the user
|
|
244
|
+
declined) or `"abandoned"` (the user sent a new message instead); tell the model
|
|
245
|
+
which, so it does not report a decline that never happened. As with
|
|
246
|
+
`experimental_get_state()`, it only works inside a server tool.
|
|
247
|
+
|
|
248
|
+
### One turn, several runs
|
|
249
|
+
|
|
250
|
+
An Antigravity *turn* that parks on a human spans several AG-UI *runs*, and on
|
|
251
|
+
each later run the harness re-delivers the steps of that turn it has already
|
|
252
|
+
sent. Three pieces of state are therefore scoped to the turn, not the run, and
|
|
253
|
+
are retired together by `AntigravitySession.reset_stream()`:
|
|
254
|
+
|
|
255
|
+
* the `receive_steps()` iterator (and any in-flight `__anext__()`, which is
|
|
256
|
+
parked rather than cancelled — cancelling discards the step being delivered),
|
|
257
|
+
* the `EventTranslator`, which records which steps and tool calls it has
|
|
258
|
+
already finished,
|
|
259
|
+
* the bridge's per-turn frontend-tool results.
|
|
260
|
+
|
|
261
|
+
Without this, every run re-translates the same tool call, the client re-executes
|
|
262
|
+
it, and the conversation never converges.
|
|
263
|
+
|
|
264
|
+
### Repeated tool calls
|
|
265
|
+
|
|
266
|
+
The harness escalates a slow custom tool to a **background task** and lets the
|
|
267
|
+
model continue without waiting for it. The model then commonly re-issues the
|
|
268
|
+
call — sometimes with slightly different arguments — which would make the client
|
|
269
|
+
run a side-effecting action a second time.
|
|
270
|
+
|
|
271
|
+
So a frontend tool is dispatched to the client **at most once per turn**. An
|
|
272
|
+
identical repeat gets the cached result; a repeat with different arguments gets
|
|
273
|
+
a plain statement of what already ran, so the model reports the result instead
|
|
274
|
+
of retrying. Set `deduplicate_tool_calls=False` if a tool is genuinely meant to
|
|
275
|
+
run repeatedly within one turn.
|
|
276
|
+
|
|
277
|
+
### Server-side tools
|
|
278
|
+
|
|
279
|
+
Pass your own Python callables as `tools=[...]` and they run in this process,
|
|
280
|
+
with the call and its result streamed to the client — that is what the dojo's
|
|
281
|
+
`backend_tool_rendering` demo draws its weather card from.
|
|
282
|
+
|
|
283
|
+
The adapter emits those events itself rather than reading them off the step
|
|
284
|
+
stream, because the harness reports a custom tool as a **single**
|
|
285
|
+
`TOOL_CALL`/`ACTIVE` step: there is no DONE step, and `Step` carries no result
|
|
286
|
+
field at all, since the return value goes back over the WebSocket straight to
|
|
287
|
+
the model. A client waiting for a `TOOL_CALL_RESULT` from the step stream would
|
|
288
|
+
wait forever. Built-in tools are different — the harness re-reports those at
|
|
289
|
+
DONE with their output folded into the call arguments.
|
|
290
|
+
|
|
291
|
+
The wrapper preserves each function's signature and docstring, so the SDK still
|
|
292
|
+
derives the same tool schema. Return a JSON-serializable value (or a string);
|
|
293
|
+
a raised exception is reported to the client as
|
|
294
|
+
`There was an error executing <tool>: ...` and re-raised.
|
|
295
|
+
|
|
296
|
+
### Shared state from server tools
|
|
297
|
+
|
|
298
|
+
*Experimental.* A server tool can read and write the AG-UI shared state of the session it runs
|
|
299
|
+
in:
|
|
300
|
+
|
|
301
|
+
```python
|
|
302
|
+
from ag_ui_antigravity import experimental_get_state, experimental_set_state
|
|
303
|
+
|
|
304
|
+
async def research_agent(task: str) -> str:
|
|
305
|
+
"""Delegates a research task."""
|
|
306
|
+
facts = await run_research(task)
|
|
307
|
+
state = experimental_get_state()
|
|
308
|
+
delegations = state.get("delegations", []) + [facts]
|
|
309
|
+
experimental_set_state({**state, "delegations": delegations})
|
|
310
|
+
return facts
|
|
311
|
+
```
|
|
312
|
+
|
|
313
|
+
`experimental_get_state()` returns a copy of the session's state: what the
|
|
314
|
+
client sent with the run (`RunAgentInput.state`), or what the last
|
|
315
|
+
`experimental_set_state()` stored if the client has not received that yet.
|
|
316
|
+
`experimental_set_state()` replaces the whole state and emits a `STATE_SNAPSHOT` at once, before the tool's `TOOL_CALL_RESULT`, so a UI
|
|
317
|
+
bound to agent state updates while the turn is still running. The state must be
|
|
318
|
+
a JSON-serializable dict; anything else raises `TypeError` in the tool.
|
|
319
|
+
|
|
320
|
+
Both raise `RuntimeError` outside a server tool. The session is found through a
|
|
321
|
+
context variable set for the duration of the call, not through a tool
|
|
322
|
+
parameter, because the SDK would put such a parameter into the tool's schema.
|
|
323
|
+
|
|
324
|
+
A write made while the turn is parked (no run attached) is queued and delivered
|
|
325
|
+
on the next run. Until then the adapter keeps its own copy rather than the
|
|
326
|
+
client's, since the client's copy predates the write.
|
|
327
|
+
|
|
328
|
+
`experimental_get_context()` works the same way and returns the run's
|
|
329
|
+
`RunAgentInput.context` (what CopilotKit's `useAgentContext` shares) as a list of
|
|
330
|
+
`{"description", "value"}` entries.
|
|
331
|
+
|
|
332
|
+
### What the model sees: app context and shared state
|
|
333
|
+
|
|
334
|
+
*Experimental, off by default.* Antigravity fixes an agent's instructions when
|
|
335
|
+
the harness session starts, so per-run input cannot be folded into the prompt
|
|
336
|
+
the way other integrations do. Without help, the model never sees the run's
|
|
337
|
+
context or shared state. The adapter can give an agent two read-only tools the
|
|
338
|
+
model calls when it needs them:
|
|
339
|
+
|
|
340
|
+
| Tool | Returns | Turn it on with |
|
|
341
|
+
|---|---|---|
|
|
342
|
+
| `get_app_context` | the run's `RunAgentInput.context` | `experimental_app_context=True` |
|
|
343
|
+
| `get_shared_state` | the session's shared state, including the user's edits in the UI | `experimental_app_state=True` |
|
|
344
|
+
|
|
345
|
+
```python
|
|
346
|
+
agent = AntigravityAgent(
|
|
347
|
+
model="gemini-2.5-flash",
|
|
348
|
+
experimental_app_context=True,
|
|
349
|
+
experimental_app_state=True,
|
|
350
|
+
system_instructions=(
|
|
351
|
+
"Call get_app_context and get_shared_state before answering anything "
|
|
352
|
+
"about the user or the app."
|
|
353
|
+
),
|
|
354
|
+
)
|
|
355
|
+
```
|
|
356
|
+
|
|
357
|
+
Both run silently: the client gets no `TOOL_CALL_*` events for them, so no tool
|
|
358
|
+
card appears for what is, elsewhere, an invisible prompt update. Their
|
|
359
|
+
docstrings tell the model to call them whenever the answer may depend on the
|
|
360
|
+
user, the page or the app state; say so in `system_instructions` too when it
|
|
361
|
+
matters. A server or client tool with the same name takes precedence.
|
|
362
|
+
|
|
363
|
+
The difference from prompt injection is that the model has to ask. A preference
|
|
364
|
+
the user changed in the UI reaches the model on its next call to
|
|
365
|
+
`get_shared_state`, not before.
|
|
366
|
+
|
|
367
|
+
### Attachments
|
|
368
|
+
|
|
369
|
+
Image, document, audio and video parts of a user message reach the model when
|
|
370
|
+
their bytes travel inline: a `data` source, or a `data:` URL. They become the
|
|
371
|
+
SDK's `Image`/`Document`/`Audio`/`Video` objects, sent in order with the text.
|
|
372
|
+
The SDK accepts PNG, JPEG, WebP and BMP images and PDF, plain-text, CSV, JSON,
|
|
373
|
+
HTML and XML documents, among others.
|
|
374
|
+
|
|
375
|
+
The harness cannot fetch anything itself, so an `https://` URL, a provider file
|
|
376
|
+
reference or an unsupported type such as GIF is replaced by a one-line note in
|
|
377
|
+
the prompt (`[Attached image 'x.png' was not forwarded: ...]`). The model then
|
|
378
|
+
says it could not see the file rather than answering as if nothing was attached.
|
|
379
|
+
`/capabilities` advertises `multimodal.input` accordingly.
|
|
380
|
+
|
|
381
|
+
### Built-in tools worth disabling
|
|
382
|
+
|
|
383
|
+
The harness exposes its whole built-in toolset by default. `search_web` returns
|
|
384
|
+
an *empty* summary unless the harness has Google credentials — the model then
|
|
385
|
+
retries it indefinitely and the conversation never settles. Pass a
|
|
386
|
+
`CapabilitiesConfig(enabled_tools=[...])` naming only what the agent needs;
|
|
387
|
+
`BuiltinTools.FINISH` must stay, since the harness uses it to end a turn. The
|
|
388
|
+
chat demos in `examples/` enable nothing else.
|
|
389
|
+
|
|
390
|
+
## Sessions
|
|
391
|
+
|
|
392
|
+
`SessionManager` keys a live `Conversation` by `thread_id`.
|
|
393
|
+
|
|
394
|
+
* **Hot resume** — a session with a parked coroutine stays in memory, because a
|
|
395
|
+
suspended coroutine cannot be serialized. It gets a longer grace period than
|
|
396
|
+
an idle session rather than an exemption: `parked_timeout_seconds` (2 h)
|
|
397
|
+
instead of `session_timeout_seconds` (30 min). A run in flight holds the
|
|
398
|
+
session lock and is never reclaimed.
|
|
399
|
+
|
|
400
|
+
A session occupies memory for its whole life, not only while parked — from
|
|
401
|
+
the first message on a `thread_id` until it times out, is evicted at
|
|
402
|
+
`max_sessions`, or is rebuilt because the client's tool contract changed —
|
|
403
|
+
a tool's name, description *or* parameter schema, since Antigravity fixes
|
|
404
|
+
the tool configuration when it connects. Parking does not allocate anything;
|
|
405
|
+
it extends how long the allocation is held.
|
|
406
|
+
* **Cold resume** — a recycled session with nothing parked is rebuilt from
|
|
407
|
+
`conversation_id` + `session_continuation_mode` + `save_dir`. This covers a
|
|
408
|
+
thread that comes back **after** its session was swept, not only one rebuilt
|
|
409
|
+
in place: the manager keeps `thread_id -> (conversation_id, forwarded
|
|
410
|
+
prompts)` for closed sessions, so a user returning past the idle timeout
|
|
411
|
+
continues where they left off rather than meeting an agent with amnesia while
|
|
412
|
+
their history sits unreachable in `save_dir`.
|
|
413
|
+
|
|
414
|
+
That map is currently unbounded — one short string and a small set per thread
|
|
415
|
+
the process has ever seen. Cap it (LRU or TTL) before running at a scale where
|
|
416
|
+
that matters.
|
|
417
|
+
|
|
418
|
+
Because Antigravity fixes the tool list in the harness config at connect time,
|
|
419
|
+
a client that changes its `tools` between runs forces a cold-resume rebuild
|
|
420
|
+
rather than running against a stale list.
|
|
421
|
+
|
|
422
|
+
## Harness pooling
|
|
423
|
+
|
|
424
|
+
Sessions do **not** get a subprocess each. A `HarnessPool` shares one
|
|
425
|
+
`localharness` process between up to `max_conversations_per_process` (default 8)
|
|
426
|
+
conversations, because Antigravity configures a harness twice:
|
|
427
|
+
|
|
428
|
+
| sent | when | contains |
|
|
429
|
+
|---|---|---|
|
|
430
|
+
| `InputConfig` | process stdin at startup | `save_dir`, `env` |
|
|
431
|
+
| `HarnessConfig` | WebSocket, **per conversation** | tools, model, system instructions, capabilities, MCP servers, hooks, subagents, `response_schema`, `conversation_id`, **`workspaces`** |
|
|
432
|
+
|
|
433
|
+
Almost everything varying per thread — including `workspaces`, so per-thread
|
|
434
|
+
filesystem isolation is unaffected — is per-conversation. Only `save_dir` and
|
|
435
|
+
`env` are process-wide, and they form the pool's partition key.
|
|
436
|
+
|
|
437
|
+
Measured with 8 concurrent conversations, one turn each:
|
|
438
|
+
|
|
439
|
+
| | pooled (1 process) | one process each |
|
|
440
|
+
|---|---|---|
|
|
441
|
+
| idle | **101 MB** | 752 MB |
|
|
442
|
+
| mid-turn | **154 MB** | 1040 MB |
|
|
443
|
+
| wall clock | 12.7 s | 12.3 s |
|
|
444
|
+
|
|
445
|
+
So an extra idle conversation costs ~1 MB rather than ~95 MB. Throughput is
|
|
446
|
+
unchanged; per-turn p50 rises ~1.3× only when all 8 turn simultaneously. A
|
|
447
|
+
20-second tool call in one conversation was measured **not** to delay its
|
|
448
|
+
neighbours (median inflation 0.95× against a control).
|
|
449
|
+
|
|
450
|
+
Pooling is configured on the agent:
|
|
451
|
+
|
|
452
|
+
| Option | Default | Effect |
|
|
453
|
+
|---|---|---|
|
|
454
|
+
| `max_conversations_per_process` | `8` | conversations sharing one harness process; `1` gives every conversation its own process |
|
|
455
|
+
| `harness_idle_grace_seconds` | `30.0` | how long an empty process stays up before it is stopped |
|
|
456
|
+
| `harness_pool` | a pool per agent | pass one `HarnessPool` to several agents to share processes between them (see below) |
|
|
457
|
+
|
|
458
|
+
Pooled conversations stay isolated at the harness level (each has its own
|
|
459
|
+
workspaces, tools and instructions), but they share one OS process. If your
|
|
460
|
+
deployment needs process-level isolation between users, set
|
|
461
|
+
`max_conversations_per_process=1`.
|
|
462
|
+
|
|
463
|
+
### Several agents in one server
|
|
464
|
+
|
|
465
|
+
Each `AntigravityAgent` owns a pool, so a server hosting four agents gets four
|
|
466
|
+
harness processes — a ~95 MB floor per agent, independent of traffic. It does
|
|
467
|
+
not affect scaling (threads of one agent still share), and it affects nothing
|
|
468
|
+
about correctness, but it is silent: you find out by counting processes.
|
|
469
|
+
|
|
470
|
+
Sharing needs **both** a pool and a `save_dir`. Passing only `harness_pool=`
|
|
471
|
+
changes nothing, because each agent otherwise mints its own `tempfile.mkdtemp()`
|
|
472
|
+
save directory and `save_dir` is half the pool's partition key:
|
|
473
|
+
|
|
474
|
+
```python
|
|
475
|
+
from ag_ui_antigravity.harness_pool import HarnessPool
|
|
476
|
+
|
|
477
|
+
pool = HarnessPool()
|
|
478
|
+
save_dir = "/var/lib/myapp/antigravity" # both, or you still get a process each
|
|
479
|
+
|
|
480
|
+
chat = AntigravityAgent(model=..., harness_pool=pool, save_dir=save_dir)
|
|
481
|
+
research = AntigravityAgent(model=..., harness_pool=pool, save_dir=save_dir)
|
|
482
|
+
```
|
|
483
|
+
|
|
484
|
+
Agents that must not share process-level storage should keep separate
|
|
485
|
+
`save_dir` values — they will land on separate processes by design. See
|
|
486
|
+
`examples/server/api/_common.py`.
|
|
487
|
+
|
|
488
|
+
* **Blast radius.** One dead process fails every conversation on it. They raise
|
|
489
|
+
promptly rather than hanging (`test_process_death_raises_rather_than_hangs`),
|
|
490
|
+
but this is the tradeoff pooling buys — hence the modest default.
|
|
491
|
+
* **Parked sessions.** Parking and pooling compose: a conversation parked on a
|
|
492
|
+
human does not block its co-tenants
|
|
493
|
+
(`test_a_parked_conversation_does_not_block_its_siblings`). But a parked
|
|
494
|
+
session pins its process, and the pool cannot know in advance which
|
|
495
|
+
conversations will park, so under park-heavy load the saving is bounded by
|
|
496
|
+
fragmentation rather than by the ~1 MB marginal figure.
|
|
497
|
+
* **Not a tenancy boundary.** `save_dir` is shared by every conversation on a
|
|
498
|
+
process. Use `workspaces` for isolation, and give tenants separate `save_dir`
|
|
499
|
+
values (which partitions them onto separate processes) if storage must be
|
|
500
|
+
isolated too.
|
|
501
|
+
|
|
502
|
+
## Operational notes
|
|
503
|
+
|
|
504
|
+
* **Sandboxing.** Real filesystem and shell access, scoped per conversation by
|
|
505
|
+
`workspaces`. Multi-tenant hosting needs per-thread workspace isolation,
|
|
506
|
+
resource caps, and reliable cleanup. Always set `workspaces=[...]`.
|
|
507
|
+
* **SDK churn.** `google-antigravity` is young; the dependency is pinned to
|
|
508
|
+
`<0.2.0` deliberately.
|
|
509
|
+
|
|
510
|
+
### A crashed harness loses its history
|
|
511
|
+
|
|
512
|
+
A conversation is pinned to one harness process for its whole life -- there is
|
|
513
|
+
no migration -- so if that process dies the session is gone with it. The run
|
|
514
|
+
that meets the corpse reports `RUN_ERROR`, and the next run on that thread
|
|
515
|
+
rebuilds the session automatically and carries on.
|
|
516
|
+
|
|
517
|
+
What does not survive is the conversation history. The harness writes
|
|
518
|
+
trajectories through SQLite's WAL and reopens them with `immutable=1`, which
|
|
519
|
+
ignores WAL files, so a killed process leaves everything uncheckpointed:
|
|
520
|
+
measured after a `SIGKILL`, the main database showed **0 steps while its
|
|
521
|
+
465 KB WAL held all 4**. The resume therefore sees an empty conversation and
|
|
522
|
+
`CREATE_OR_RESUME` quietly starts a new one. This is upstream and not
|
|
523
|
+
something the integration can work around; a clean shutdown checkpoints
|
|
524
|
+
normally and resumes fine.
|
|
525
|
+
|
|
526
|
+
The rebuild logs a warning naming the thread, so silent context loss is at
|
|
527
|
+
least visible in the logs.
|
|
528
|
+
|
|
529
|
+
### Keep workspace paths short
|
|
530
|
+
|
|
531
|
+
A long workspace path makes runs fail intermittently, and the failure looks
|
|
532
|
+
nothing like its cause.
|
|
533
|
+
|
|
534
|
+
If a prompt leads the model to write an absolute path into a tool call — "read
|
|
535
|
+
`/var/folders/0t/pq2_7rn97834lcsvc_qy4t8r0000gn/T/ag-ui-antigravity-wz_4xjbd/notes.txt`"
|
|
536
|
+
— it sometimes reproduces that path wrongly, truncating it or repeating a chunk
|
|
537
|
+
of it. macOS temp directories are ~75 characters of high-entropy text, which is
|
|
538
|
+
about the worst case. Measured: **0/14 runs failed with a
|
|
539
|
+
9-character workspace path, 2/14 with a 75-character one.**
|
|
540
|
+
|
|
541
|
+
The harness treats the resulting bad path as a fatal
|
|
542
|
+
`AntigravityExecutionError` rather than returning the error for the model to
|
|
543
|
+
retry, so the whole run dies mid-tool-call with
|
|
544
|
+
`RUN_ERROR: The model produced an invalid tool call`. Nothing in that message
|
|
545
|
+
points at path length.
|
|
546
|
+
|
|
547
|
+
Set `ANTIGRAVITY_WORKSPACE` (or `workspaces=[...]`) to something short and
|
|
548
|
+
stable — `/srv/agents/w1`, not a generated temp directory. This is a model
|
|
549
|
+
limitation rather than an integration bug: it reproduces identically on older
|
|
550
|
+
commits.
|
|
551
|
+
|
|
552
|
+
### Known gaps in `google-antigravity` 0.1.8–0.1.9 (OpenAI-compatible path)
|
|
553
|
+
|
|
554
|
+
These are upstream, not integration bugs. They affect only `base_url` usage:
|
|
555
|
+
|
|
556
|
+
1. **No API-key field.** `GemmaEndpoint` carries only `base_url`, and the Go
|
|
557
|
+
harness reads no `OPENAI_API_KEY`. The path targets unauthenticated local
|
|
558
|
+
servers (Ollama, LM Studio), so hosted OpenAI rejects every request.
|
|
559
|
+
2. **Gemini-shaped tool schemas.** Custom-tool schemas are generated with
|
|
560
|
+
`api_option="GEMINI_API"`, emitting proto-style uppercase types (`"STRING"`)
|
|
561
|
+
that OpenAI rejects. Tools registered via `ToolWithSchema` — which is how
|
|
562
|
+
this integration builds *frontend* tools — pass their schema through
|
|
563
|
+
untouched and are unaffected.
|
|
564
|
+
3. **`session_continuation_mode` dropped.** `LocalOpenAIAgentConfig.create_strategy`
|
|
565
|
+
does not forward it, disabling cold resume. Worked around by
|
|
566
|
+
`_ResumableOpenAIConfig` in `agent.py`.
|
|
567
|
+
|
|
568
|
+
### Tool calls do not stream their arguments
|
|
569
|
+
|
|
570
|
+
Not specific to the OpenAI path, and not workable around: the harness hands over
|
|
571
|
+
a tool call **fully formed, in a single step**. Measured with a custom tool
|
|
572
|
+
whose argument was 1850 characters — still one `TOOL_CALL` step carrying the
|
|
573
|
+
complete arguments. Presumably the Go side parses the model's tool-call JSON
|
|
574
|
+
before dispatching, since it needs valid JSON to invoke anything, and surfaces
|
|
575
|
+
the parsed call rather than the token stream.
|
|
576
|
+
|
|
577
|
+
So `TOOL_CALL_ARGS` arrives as one delta rather than filling in progressively.
|
|
578
|
+
There are no fragments to forward; this would need the harness to expose partial
|
|
579
|
+
tool-call deltas. What *does* stream is the call lifecycle: the call is emitted
|
|
580
|
+
as soon as the harness dispatches it, so a client renders its pending state
|
|
581
|
+
while the tool runs and swaps in the result when it returns.
|
|
582
|
+
|
|
583
|
+
## Not implemented yet
|
|
584
|
+
|
|
585
|
+
Deliberate gaps, so the surface above is not mistaken for more than it is:
|
|
586
|
+
|
|
587
|
+
* **Triggers** (async inbound messages) — the SDK supports them; nothing here
|
|
588
|
+
maps them to AG-UI yet.
|
|
589
|
+
* **`STATE_DELTA`** — structured output and `experimental_set_state()` are emitted as whole
|
|
590
|
+
snapshots only.
|
|
591
|
+
* **`forwardedProps`** — apart from interrupt answers, not passed to the model
|
|
592
|
+
or to tools.
|
|
593
|
+
* **MCP servers** — passed through to the SDK config and covered by the
|
|
594
|
+
approval hook, but not exercised by a live test.
|
|
595
|
+
* **`predictive_state_updates`** dojo feature — it streams a tool's
|
|
596
|
+
arguments into state while the model writes them, and the harness hands over
|
|
597
|
+
tool calls whole. (`shared_state` is in the menu, backed by
|
|
598
|
+
`examples/server/api/shared_state.py`.)
|
|
599
|
+
* Subagent bracketing is unit-tested against recorded step shapes, not against
|
|
600
|
+
a live multi-agent run.
|
|
601
|
+
|
|
602
|
+
## Development
|
|
603
|
+
|
|
604
|
+
```bash
|
|
605
|
+
uv sync
|
|
606
|
+
uv run pytest # 315 unit tests; live tests are deselected by default
|
|
607
|
+
```
|
|
608
|
+
|
|
609
|
+
The live checks start a real harness subprocess and call Gemini:
|
|
610
|
+
|
|
611
|
+
```bash
|
|
612
|
+
export GEMINI_API_KEY=...
|
|
613
|
+
uv run pytest tests/ -m live
|
|
614
|
+
```
|
|
615
|
+
|
|
616
|
+
`ANTIGRAVITY_TEST_MODEL` picks the model (default: the SDK's default model), and
|
|
617
|
+
`GOOGLE_GEMINI_BASE_URL` points the same tests at a Gemini-compatible gateway.
|
|
618
|
+
|
|
619
|
+
`tests/test_parking_gate.py` is the important one — it re-verifies the Go-side
|
|
620
|
+
no-timeout property the whole HITL design depends on. Raise the park duration
|
|
621
|
+
to reproduce the long soak:
|
|
622
|
+
|
|
623
|
+
```bash
|
|
624
|
+
PARK_SECONDS=180 uv run pytest tests/test_parking_gate.py -m live
|
|
625
|
+
```
|
|
626
|
+
|
|
627
|
+
### Dojo
|
|
628
|
+
|
|
629
|
+
```bash
|
|
630
|
+
# terminal 1
|
|
631
|
+
cd examples
|
|
632
|
+
GEMINI_API_KEY=... uv run dev # serves on :8027
|
|
633
|
+
|
|
634
|
+
# terminal 2
|
|
635
|
+
cd apps/dojo && pnpm dev
|
|
636
|
+
```
|
|
637
|
+
|
|
638
|
+
`pnpm run-dojo-everything` starts it alongside the other integrations.
|
|
639
|
+
|
|
640
|
+
To replay aimock fixtures instead of calling Gemini, point the server at aimock.
|
|
641
|
+
The harness insists on a key, but aimock ignores its value:
|
|
642
|
+
`GOOGLE_GEMINI_BASE_URL=http://localhost:4010 AIMOCK_CONTEXT=<fixture-set> GEMINI_API_KEY=unused uv run dev`.
|
|
643
|
+
|
|
644
|
+
Then open `/antigravity/feature/agentic_chat`.
|
|
645
|
+
|
|
646
|
+
## Verification status
|
|
647
|
+
|
|
648
|
+
Verified on 2026-09-30 against `google-antigravity` 0.1.9 on the native Gemini
|
|
649
|
+
path, with no proxy in between:
|
|
650
|
+
|
|
651
|
+
| Check | Result |
|
|
652
|
+
|---|---|
|
|
653
|
+
| 315 unit tests (translator, bridge, sessions, endpoint, config) | pass |
|
|
654
|
+
| 19 live tests against Gemini (streaming, multi-turn, frontend-tool park/resume, built-in tools, SSE, server tools, cold resume, pooling) | 15 pass; 4 were cut short by the free-tier key's quota (429) or Gemini overload (503) errors |
|
|
655
|
+
| Harness parked with the stream closed, then resumed | pass in the CopilotKit showcase's human-in-the-loop cells; the 30 s `test_parking_gate.py` soak ran into the quota limit |
|
|
656
|
+
| `endpoint=GeminiAPIEndpoint(...)` against aimock: text, server-tool round trip, reasoning | pass (reasoning streams as `REASONING_*` events) |
|
|
657
|
+
| Example server against aimock via `GOOGLE_GEMINI_BASE_URL` | pass |
|
|
658
|
+
| `/capabilities` payload against `AgentCapabilitiesSchema` (zod) | valid |
|