@kindgi/sdk 0.0.0-bootstrap.0 → 0.1.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +201 -0
- package/README.md +187 -1
- package/dist/build.d.ts +10 -0
- package/dist/build.d.ts.map +1 -0
- package/dist/build.js +11 -0
- package/dist/build.js.map +1 -0
- package/dist/client.d.ts +33 -0
- package/dist/client.d.ts.map +1 -0
- package/dist/client.js +55 -0
- package/dist/client.js.map +1 -0
- package/dist/define.d.ts +25 -0
- package/dist/define.d.ts.map +1 -0
- package/dist/define.js +37 -0
- package/dist/define.js.map +1 -0
- package/dist/index.d.ts +17 -0
- package/dist/index.d.ts.map +1 -0
- package/dist/index.js +19 -0
- package/dist/index.js.map +1 -0
- package/dist/runtime-config.d.ts +53 -0
- package/dist/runtime-config.d.ts.map +1 -0
- package/dist/runtime-config.js +114 -0
- package/dist/runtime-config.js.map +1 -0
- package/dist/types.d.ts +16 -0
- package/dist/types.d.ts.map +1 -0
- package/dist/types.js +4 -0
- package/dist/types.js.map +1 -0
- package/dist/webhooks.d.ts +15 -0
- package/dist/webhooks.d.ts.map +1 -0
- package/dist/webhooks.js +16 -0
- package/dist/webhooks.js.map +1 -0
- package/package.json +88 -4
- package/skills/kindgi-authoring-agents/SKILL.md +252 -0
- package/skills/kindgi-authoring-flows/SKILL.md +302 -0
- package/skills/kindgi-authoring-guardrails/SKILL.md +297 -0
- package/skills/kindgi-authoring-mcp-servers/SKILL.md +289 -0
- package/skills/kindgi-authoring-providers/SKILL.md +705 -0
- package/skills/kindgi-authoring-tools/SKILL.md +298 -0
- package/skills/kindgi-framework-feedback/SKILL.md +211 -0
- package/skills/kindgi-getting-started/SKILL.md +189 -0
- package/skills/kindgi-python-authoring-agents/SKILL.md +205 -0
- package/skills/kindgi-python-authoring-flows/SKILL.md +325 -0
- package/skills/kindgi-python-authoring-guardrails/SKILL.md +176 -0
- package/skills/kindgi-python-authoring-tools/SKILL.md +305 -0
- package/skills/kindgi-python-getting-started/SKILL.md +242 -0
- package/src/build.ts +18 -0
- package/src/client.ts +177 -0
- package/src/define.ts +75 -0
- package/src/index.ts +20 -0
- package/src/runtime-config.ts +180 -0
- package/src/types.ts +86 -0
- package/src/webhooks.ts +34 -0
|
@@ -0,0 +1,325 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: kindgi-python-authoring-flows
|
|
3
|
+
description: >
|
|
4
|
+
Covers writing flows for a Kindgi pack in Python (the `kindgi`
|
|
5
|
+
package): declaring a `Flow(...)` at module level, tool and agent
|
|
6
|
+
steps (Tool and Agent objects as refs), edges and `when` conditions,
|
|
7
|
+
branches that join again, inputMapping from runInput / nodeOutputs,
|
|
8
|
+
typed agent output in a flow, the flow's declared output, loops
|
|
9
|
+
(foreach / while) and fanout, per-edge retry and timeout, and running
|
|
10
|
+
a flow (kindgi runs start --flow, in the background, as a dry run) and
|
|
11
|
+
reading its journal. Load this whenever you are authoring or editing
|
|
12
|
+
code inside a Python pack's flows/ directory (a pack whose config is
|
|
13
|
+
`[tool.kindgi]` in pyproject.toml), defining a flow, or when the user
|
|
14
|
+
asks to add, change or debug one. Python tools are covered by
|
|
15
|
+
kindgi-python-authoring-tools, Python agents by
|
|
16
|
+
kindgi-python-authoring-agents.
|
|
17
|
+
type: core
|
|
18
|
+
library: "kindgi (Python)"
|
|
19
|
+
version: "0.1.1"
|
|
20
|
+
sdk_version: "0.0.0"
|
|
21
|
+
pack_languages: [python]
|
|
22
|
+
sources:
|
|
23
|
+
- sdks/python/src/kindgi/pack/define.py
|
|
24
|
+
- sdks/python/src/kindgi/pack/index.py
|
|
25
|
+
- packages/specs/schemas/flow.schema.json
|
|
26
|
+
---
|
|
27
|
+
|
|
28
|
+
# Authoring Kindgi flows in Python
|
|
29
|
+
|
|
30
|
+
> **Running `kindgi`:** a Python pack has no Node project, so the
|
|
31
|
+
> `kindgi` CLI is the one on `PATH`. Python commands run in the pack's
|
|
32
|
+
> environment: `uv run …` (or `.venv/bin/python …`).
|
|
33
|
+
|
|
34
|
+
A **flow** is a versioned, durable graph of steps: tools (your code) and
|
|
35
|
+
agents (a model's judgment), joined by edges that can carry conditions.
|
|
36
|
+
It is **data**, not code: a `Flow(...)` at module level in
|
|
37
|
+
`flows/<name>.py`. The runtime runs it step by step, journals every
|
|
38
|
+
step, and can resume a run that was interrupted. A run pins the flow
|
|
39
|
+
version it started on.
|
|
40
|
+
|
|
41
|
+
Use a flow when the order of the work is known: parse, then classify,
|
|
42
|
+
then branch, then write. Use a single agent when the model should decide
|
|
43
|
+
the order.
|
|
44
|
+
|
|
45
|
+
## Ask before building
|
|
46
|
+
|
|
47
|
+
- **What goes in, and what comes out?** The run input's shape and the
|
|
48
|
+
output the caller reads. They become `runInput.*` paths and `output`.
|
|
49
|
+
- **Which steps are code, which are judgment?** Deterministic work
|
|
50
|
+
(parse, rank, look up, write) is a tool; judgment (classify, draft,
|
|
51
|
+
summarize) is an agent with a typed `output`.
|
|
52
|
+
- **Where does it branch?** Every branch needs a condition, and the
|
|
53
|
+
steps after a branch must cope with the branch that didn't run.
|
|
54
|
+
- **What does it change outside Kindgi?** Know which tools write: a dry
|
|
55
|
+
run stops before them (see "Running a flow").
|
|
56
|
+
|
|
57
|
+
## A flow
|
|
58
|
+
|
|
59
|
+
```python
|
|
60
|
+
# flows/triage_ticket.py
|
|
61
|
+
from kindgi import Flow
|
|
62
|
+
|
|
63
|
+
from ..agents.ticket_classifier import ticket_classifier
|
|
64
|
+
from ..tools.tickets import draft_reply, lookup_invoice, parse_ticket
|
|
65
|
+
|
|
66
|
+
IS_BILLING = {
|
|
67
|
+
"op": "eq",
|
|
68
|
+
"left": {"path": "nodeOutputs.classify.output.category"},
|
|
69
|
+
"right": {"literal": "billing"},
|
|
70
|
+
}
|
|
71
|
+
|
|
72
|
+
triage_ticket = Flow(
|
|
73
|
+
id="acme.triage-ticket",
|
|
74
|
+
version="0.1.0",
|
|
75
|
+
name="Triage a support ticket",
|
|
76
|
+
description="Parses a ticket, classifies it, looks up billing when needed, drafts a reply.",
|
|
77
|
+
nodes=[
|
|
78
|
+
{
|
|
79
|
+
"id": "parse",
|
|
80
|
+
"kind": "tool",
|
|
81
|
+
"ref": parse_ticket,
|
|
82
|
+
"inputMapping": {"ticket": {"path": "runInput.ticket"}},
|
|
83
|
+
},
|
|
84
|
+
{
|
|
85
|
+
"id": "classify",
|
|
86
|
+
"kind": "agent",
|
|
87
|
+
"ref": ticket_classifier, # an Agent with output=
|
|
88
|
+
"inputMapping": {"text": {"path": "nodeOutputs.parse.text"}},
|
|
89
|
+
"config": {"parameters": {"product": "acme-cloud"}},
|
|
90
|
+
},
|
|
91
|
+
{
|
|
92
|
+
"id": "billing",
|
|
93
|
+
"kind": "tool",
|
|
94
|
+
"ref": lookup_invoice,
|
|
95
|
+
"inputMapping": {"customer_id": {"path": "runInput.ticket.customer_id"}},
|
|
96
|
+
},
|
|
97
|
+
{
|
|
98
|
+
"id": "reply",
|
|
99
|
+
"kind": "tool",
|
|
100
|
+
"ref": draft_reply,
|
|
101
|
+
"inputMapping": {
|
|
102
|
+
"category": {"path": "nodeOutputs.classify.output.category"},
|
|
103
|
+
"invoice": {"path": "nodeOutputs.billing.invoice"}, # absent when billing didn't run
|
|
104
|
+
},
|
|
105
|
+
},
|
|
106
|
+
],
|
|
107
|
+
edges=[
|
|
108
|
+
{"id": "e0", "from": "$start", "to": "parse"},
|
|
109
|
+
{"id": "e1", "from": "parse", "to": "classify"},
|
|
110
|
+
{"id": "e2", "from": "classify", "to": "billing", "when": IS_BILLING},
|
|
111
|
+
{"id": "e3", "from": "classify", "to": "reply", "when": {"op": "not", "child": IS_BILLING}},
|
|
112
|
+
{"id": "e4", "from": "billing", "to": "reply"},
|
|
113
|
+
{"id": "e5", "from": "reply", "to": "$end"},
|
|
114
|
+
],
|
|
115
|
+
output={
|
|
116
|
+
"mapping": {
|
|
117
|
+
"category": {"path": "nodeOutputs.classify.output.category"},
|
|
118
|
+
"reply": {"path": "nodeOutputs.reply.text"},
|
|
119
|
+
},
|
|
120
|
+
"schema": {
|
|
121
|
+
"type": "object",
|
|
122
|
+
"properties": {"category": {"type": "string"}, "reply": {"type": "string"}},
|
|
123
|
+
"required": ["category", "reply"],
|
|
124
|
+
},
|
|
125
|
+
},
|
|
126
|
+
)
|
|
127
|
+
```
|
|
128
|
+
|
|
129
|
+
- **Refs** are the `Tool` / `Agent` (or `Flow`) objects, imported
|
|
130
|
+
relatively from the pack's own modules — anywhere in the flow, loop
|
|
131
|
+
bodies and fanout branches included. A primitive from another pack is
|
|
132
|
+
its id string. Ids are plain strings; no casts.
|
|
133
|
+
- **The dicts are the wire shape**, so their keys are camelCase
|
|
134
|
+
(`inputMapping`, `loopKind`, `maxIterations`); `Flow`'s own keywords
|
|
135
|
+
are snake_case (`max_parallelism`).
|
|
136
|
+
- **Where it's checked:** `Flow(...)` checks the id when the module
|
|
137
|
+
loads. The indexer checks the rest against `flow.schema.json` —
|
|
138
|
+
version, node and edge shapes, `$start` / `$end` — and `kindgi dev`
|
|
139
|
+
reports a mistake as a file error that says what and where
|
|
140
|
+
(`flow 'acme.triage-ticket' is invalid at /nodes/0: must NOT have
|
|
141
|
+
additional properties ('input_mapping')`), while the pack's other
|
|
142
|
+
primitives keep serving. That a ref's tool or agent
|
|
143
|
+
exists is checked when a run starts (see "Running a flow").
|
|
144
|
+
|
|
145
|
+
## Nodes
|
|
146
|
+
|
|
147
|
+
- **`"kind": "tool"`** runs the tool `ref`. The tool's input is what the
|
|
148
|
+
node's `inputMapping` builds, else the output of the node's single
|
|
149
|
+
upstream node (the run input after `$start`). It is validated against
|
|
150
|
+
the tool's input schema, so a mismatch fails the step with
|
|
151
|
+
`input-validation-failed`. The node's output is the tool's return
|
|
152
|
+
value, as JSON (field aliases, if the model has any).
|
|
153
|
+
- **`"kind": "agent"`** runs one turn of the agent `ref` as a child run
|
|
154
|
+
of the flow run. The agent gets the node's input in two ways:
|
|
155
|
+
- as **structured input**: `{{ input.text }}` in its instructions;
|
|
156
|
+
- as its user message (the input as JSON).
|
|
157
|
+
|
|
158
|
+
`"config": {"parameters": {…}}` fills the agent's `parameters`
|
|
159
|
+
(string, number or boolean values). `"config": {"version": "1.2.0"}`
|
|
160
|
+
pins an agent version; without it the latest active version runs.
|
|
161
|
+
|
|
162
|
+
The node's output:
|
|
163
|
+
- `output` is the agent's typed answer (its `output=` model);
|
|
164
|
+
- `text` is the answer as text;
|
|
165
|
+
- `runId` and `conversationId` belong to the child run.
|
|
166
|
+
|
|
167
|
+
Read a field as `nodeOutputs.<node>.output.<field>`. An answer that
|
|
168
|
+
doesn't fit the agent's `output`, after its repairs, fails the step
|
|
169
|
+
with `output-schema-violation`. An approval inside the agent's turn
|
|
170
|
+
parks the flow until it's decided.
|
|
171
|
+
- **`"kind": "loop"`** repeats a body: `"loopKind": "foreach"` once per
|
|
172
|
+
element of `iterateOver` (`concurrency` up to 32 in parallel), or
|
|
173
|
+
`"loopKind": "while"` until `exitCondition`.
|
|
174
|
+
- The body (`"body": {"nodes": [...], "edges": [...]}`) has its own
|
|
175
|
+
nodes and edges, with `$loop-start` / `$loop-end`; the element is
|
|
176
|
+
the body's input.
|
|
177
|
+
- `maxIterations` and `outputSchema` are required.
|
|
178
|
+
- The loop's output is `finalOutput`, plus `outputs` with
|
|
179
|
+
`"collectAllIterations": True`.
|
|
180
|
+
- Node ids must be unique across the whole flow, bodies included.
|
|
181
|
+
- **`"kind": "fanout"`** runs several handlers on the same input at
|
|
182
|
+
once, each a branch (`branchId`, `handler` — a `Tool` or an id —
|
|
183
|
+
and `outputSchema`). `convergence` decides the result: `all-succeed`,
|
|
184
|
+
`any-succeed` (the first success wins), or `settle-all` (wait for
|
|
185
|
+
every branch and report each).
|
|
186
|
+
- **`"kind": "subgraph"`** (a sub-flow) is part of the flow schema, but a
|
|
187
|
+
run refuses it today (`flow-unbound`). Inline the steps instead.
|
|
188
|
+
|
|
189
|
+
## Edges and conditions
|
|
190
|
+
|
|
191
|
+
An edge goes from a node (or `$start`) to a node (or `$end`). Without
|
|
192
|
+
`when` it fires when its source completes; with `when` it fires only if
|
|
193
|
+
the condition is true. Conditions are dicts:
|
|
194
|
+
|
|
195
|
+
| Operator | Shape |
|
|
196
|
+
|---|---|
|
|
197
|
+
| `eq` `ne` `lt` `lte` `gt` `gte` | `{"op", "left", "right"}` |
|
|
198
|
+
| `in` `notIn` | `{"op", "value", "set"}` |
|
|
199
|
+
| `exists` `notExists` `truthy` `falsy` | `{"op", "value"}` |
|
|
200
|
+
| `and` `or` | `{"op", "children": [...]}` |
|
|
201
|
+
| `not` | `{"op", "child"}` |
|
|
202
|
+
|
|
203
|
+
Each operand is `{"literal": …}` or `{"path": …}`. When a path doesn't
|
|
204
|
+
resolve, `eq`, `lt`, `lte`, `gt` and `gte` are false and `ne` is true. So
|
|
205
|
+
for the "otherwise" branch, write `not` around the condition (as above),
|
|
206
|
+
rather than a second comparison: it covers exactly what the first edge
|
|
207
|
+
doesn't. A condition
|
|
208
|
+
used twice is easiest as a module-level constant (`IS_BILLING`).
|
|
209
|
+
|
|
210
|
+
**Joining branches.** A node with several incoming edges runs once every
|
|
211
|
+
one of them is decided and at least one fired. In the example, `reply`
|
|
212
|
+
runs after `billing` on the billing branch, and straight after
|
|
213
|
+
`classify` otherwise. A node none of whose incoming edges fired is
|
|
214
|
+
skipped, and so is everything only it leads to.
|
|
215
|
+
|
|
216
|
+
**Edge policy** (`"policy"` on the edge into a node with a single
|
|
217
|
+
incoming edge):
|
|
218
|
+
- `"retry": {"maxAttempts", "delayMs"?, "backoff"?, "maxDelayMs"?}`: up
|
|
219
|
+
to 10 attempts in all;
|
|
220
|
+
- `"timeoutMs"`: a step that takes longer fails with `reason: timeout`;
|
|
221
|
+
- `"concurrencyKey"`: at most one such step at a time in the tenant;
|
|
222
|
+
- `"priority"`: −100 to 100.
|
|
223
|
+
|
|
224
|
+
A node with several incoming edges ignores them.
|
|
225
|
+
|
|
226
|
+
## Inputs and the output
|
|
227
|
+
|
|
228
|
+
`inputMapping` maps each key to a `{"literal": …}` or a `{"path": …}`.
|
|
229
|
+
Its keys are the tool's input **as it travels**: a pydantic field's name,
|
|
230
|
+
or its alias if it has one (a `customer_id` field is the key
|
|
231
|
+
`customer_id`; with `alias="customerId"`, it's `customerId`). Paths are
|
|
232
|
+
dot-separated (a number segment indexes an array: `items.0.sku`), rooted at:
|
|
233
|
+
- `runInput.…`: the input the run was started with;
|
|
234
|
+
- `nodeOutputs.<nodeId>.…`: a step's output. For an agent step, add
|
|
235
|
+
`.output.<field>` to read its typed answer;
|
|
236
|
+
- `state.…`: values written by the runtime's own handlers. Pack tools
|
|
237
|
+
don't write it, so use `nodeOutputs`.
|
|
238
|
+
|
|
239
|
+
A path that doesn't resolve leaves its key out. A step after a branch
|
|
240
|
+
that didn't run gets no `invoice` key at all, rather than `None`. Give
|
|
241
|
+
that field a default in the tool's input model
|
|
242
|
+
(`invoice: Invoice | None = None`).
|
|
243
|
+
|
|
244
|
+
`output` is what the run returns: a `mapping` resolved when the run
|
|
245
|
+
finishes, checked against `schema` if you give one. A run whose output
|
|
246
|
+
doesn't match fails. Without `output`, the run returns the output of the
|
|
247
|
+
step that reached `$end`.
|
|
248
|
+
|
|
249
|
+
## Running a flow
|
|
250
|
+
|
|
251
|
+
From another terminal in the pack directory, while `kindgi dev` runs:
|
|
252
|
+
|
|
253
|
+
```sh
|
|
254
|
+
kindgi runs start --flow=acme.triage-ticket --input='{"ticket":{"customer_id":"c-1","body":"Charged twice"}}'
|
|
255
|
+
kindgi runs start --flow=acme.triage-ticket --input=@ticket.json --no-wait # the run id now; it finishes in the background
|
|
256
|
+
kindgi runs start --flow=acme.triage-ticket --input=@ticket.json --dry-run # stops before a tool that may write
|
|
257
|
+
kindgi runs get <run-id> # status, output, failureMessage
|
|
258
|
+
kindgi runs journal <run-id> # every step.started / step.completed / edge.evaluated
|
|
259
|
+
kindgi runs stream <run-id> # follow a running one
|
|
260
|
+
kindgi runs cancel <run-id>
|
|
261
|
+
```
|
|
262
|
+
|
|
263
|
+
Or from Python, with `kindgi.client`:
|
|
264
|
+
`Kindgi().runs.start(flow="acme.triage-ticket", input={...})`.
|
|
265
|
+
|
|
266
|
+
- **A refusal before the run exists:** `422 flow-unbound` names the
|
|
267
|
+
nodes a run can't bind: a tool or agent id the tenant doesn't have, or
|
|
268
|
+
a sub-flow. Fix the ids; nothing ran.
|
|
269
|
+
- **A failed step fails the run**, and `failureMessage` says which step
|
|
270
|
+
and why. Retry it on its edge with `policy.retry` only if running the
|
|
271
|
+
step twice is safe.
|
|
272
|
+
- **`--no-wait`** is how an application starts runs
|
|
273
|
+
(`options={"wait": False}`). It answers with the run id at once; poll
|
|
274
|
+
the run or follow its stream.
|
|
275
|
+
- **`--dry-run`** runs a tool only if it's declared read-only:
|
|
276
|
+
`@tool(..., mutating=False)` (or `http_tool(..., mutating=False)`),
|
|
277
|
+
and no `writes`, `deletes`, `spawns-run`, `emits-event` or
|
|
278
|
+
`external-side-effect` effect. The first other tool stops the run with
|
|
279
|
+
`dry-run-effectful-tool`, and everything before it really ran. That's
|
|
280
|
+
useful for checking the wiring without the writes.
|
|
281
|
+
|
|
282
|
+
## Iterating on a flow
|
|
283
|
+
|
|
284
|
+
Save the file and `kindgi dev` re-indexes; the next run uses the new
|
|
285
|
+
definition, with no restart. A run already in flight keeps the version it
|
|
286
|
+
started on. Bump `version` when callers' contract changes (the input or
|
|
287
|
+
the output), not on every save.
|
|
288
|
+
|
|
289
|
+
## Common mistakes
|
|
290
|
+
|
|
291
|
+
1. **Building a flow without asking what goes in and comes out.** The
|
|
292
|
+
pack's `echo_flow` proves the runtime works. It isn't a template for
|
|
293
|
+
the user's flow.
|
|
294
|
+
2. **A second comparison for "otherwise".** On a path that may be
|
|
295
|
+
missing, `eq` is false and `ne` is true, and `lt`/`gt` are both false,
|
|
296
|
+
so a hand-written opposite can miss a case or overlap. Use `not` around
|
|
297
|
+
the positive condition: it covers exactly what the first edge doesn't.
|
|
298
|
+
3. **Reading an agent step's answer at `nodeOutputs.<step>.<field>`.**
|
|
299
|
+
The typed answer is under `.output`: `nodeOutputs.<step>.output.<field>`.
|
|
300
|
+
An agent without `output=` has only `text`.
|
|
301
|
+
4. **A required input field fed by a branch that may not run.** The key
|
|
302
|
+
is omitted, the tool's input check fails, and so does the step. Give
|
|
303
|
+
the field a default.
|
|
304
|
+
5. **`inputMapping` keys in the wrong spelling.** They're the input's
|
|
305
|
+
wire names (field names, or aliases): `customerId` doesn't fill a
|
|
306
|
+
`customer_id` field that has no alias.
|
|
307
|
+
6. **snake_case keys inside the dicts** (`input_mapping`,
|
|
308
|
+
`max_iterations`). The indexer refuses them, naming the key; the
|
|
309
|
+
dicts are camelCase.
|
|
310
|
+
7. **A sub-flow node.** A run refuses it (`flow-unbound`) until
|
|
311
|
+
sub-flows are supported.
|
|
312
|
+
8. **Duplicate node ids inside a loop body.** Ids are unique across the
|
|
313
|
+
whole flow, bodies included.
|
|
314
|
+
9. **`mutating=False` on a tool that writes.** A dry run then runs it for
|
|
315
|
+
real, and an agent's tool approval gate (when it falls back on
|
|
316
|
+
`mutating`) won't ask before it.
|
|
317
|
+
10. **A read-only tool without `mutating=False`.** A dry run stops at it,
|
|
318
|
+
and an agent's tool approval gate asks before it on first use.
|
|
319
|
+
|
|
320
|
+
## When the framework itself is the problem
|
|
321
|
+
|
|
322
|
+
If the bug is in Kindgi or the `kindgi` package (a step's output missing
|
|
323
|
+
a field, a condition that evaluates wrongly, a misleading error) and not
|
|
324
|
+
in the pack's code, load `kindgi-framework-feedback` and file it with
|
|
325
|
+
`kindgi feedback write`.
|
|
@@ -0,0 +1,176 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: kindgi-python-authoring-guardrails
|
|
3
|
+
description: >
|
|
4
|
+
Covers writing guardrails (safety checks on an agent's turn) for a
|
|
5
|
+
Kindgi pack in Python (the `kindgi` package): the `@guardrail`
|
|
6
|
+
decorator over a `(config, trace)` check, `RunTrace` and
|
|
7
|
+
`CheckResult`, config models, actions (halt / retry / escalate /
|
|
8
|
+
log-only / compensate), severity and scope, unit tests, and wiring a
|
|
9
|
+
guardrail onto an agent. Load this whenever you are authoring or
|
|
10
|
+
editing code inside a Python pack's guardrails/ directory (a pack
|
|
11
|
+
whose config is `[tool.kindgi]` in pyproject.toml), defining a check,
|
|
12
|
+
or wiring a guardrail onto an agent. Python agents are covered by
|
|
13
|
+
kindgi-python-authoring-agents, Python tools by
|
|
14
|
+
kindgi-python-authoring-tools.
|
|
15
|
+
type: core
|
|
16
|
+
library: "kindgi (Python)"
|
|
17
|
+
version: "0.1.1"
|
|
18
|
+
sdk_version: "0.0.0"
|
|
19
|
+
pack_languages: [python]
|
|
20
|
+
sources:
|
|
21
|
+
- sdks/python/src/kindgi/pack/define.py
|
|
22
|
+
- sdks/python/src/kindgi/pack/trace.py
|
|
23
|
+
- sdks/python/src/kindgi/pack/service.py
|
|
24
|
+
---
|
|
25
|
+
|
|
26
|
+
# Authoring Kindgi guardrails in Python
|
|
27
|
+
|
|
28
|
+
> **Running `kindgi`:** a Python pack has no Node project, so the
|
|
29
|
+
> `kindgi` CLI is the one on `PATH`. Python commands run in the pack's
|
|
30
|
+
> environment: `uv run …` (or `.venv/bin/python …`).
|
|
31
|
+
|
|
32
|
+
A **guardrail** is a rule an agent's turn must satisfy: a **check** (a
|
|
33
|
+
function over the turn's trace) plus an **action** (what happens when it
|
|
34
|
+
fails). In a Python pack, `@guardrail(...)` on a function at module
|
|
35
|
+
level in a file under `guardrails/` declares both. For an agent turn,
|
|
36
|
+
the runtime evaluates every guardrail the agent lists once, on the final
|
|
37
|
+
answer, before it is stored.
|
|
38
|
+
|
|
39
|
+
Ask what the rule should catch before writing one; the sample
|
|
40
|
+
`response-not-empty` guardrail is a demonstration, not a template.
|
|
41
|
+
|
|
42
|
+
## A guardrail
|
|
43
|
+
|
|
44
|
+
```python
|
|
45
|
+
# guardrails/citations.py
|
|
46
|
+
from pydantic import BaseModel, Field
|
|
47
|
+
|
|
48
|
+
from kindgi import CheckResult, RunTrace, guardrail
|
|
49
|
+
|
|
50
|
+
|
|
51
|
+
class Config(BaseModel):
|
|
52
|
+
min_lookups: int = Field(1, alias="minLookups", ge=0)
|
|
53
|
+
|
|
54
|
+
|
|
55
|
+
@guardrail(
|
|
56
|
+
id="acme.no-fabricated-quotes",
|
|
57
|
+
name="No fabricated quotations",
|
|
58
|
+
on_violation="halt",
|
|
59
|
+
severity="critical",
|
|
60
|
+
config={"minLookups": 2}, # what the check runs with — keyed as on the wire
|
|
61
|
+
)
|
|
62
|
+
def no_fabricated_quotes(config: Config, trace: RunTrace) -> CheckResult:
|
|
63
|
+
lookups = [c for c in trace.tool_calls if c.tool_name == "acme.verify-citation"]
|
|
64
|
+
if len(lookups) < config.min_lookups:
|
|
65
|
+
return CheckResult(
|
|
66
|
+
passed=False,
|
|
67
|
+
reason=f"Only {len(lookups)} citation lookups (need {config.min_lookups}+).",
|
|
68
|
+
)
|
|
69
|
+
return CheckResult(passed=True)
|
|
70
|
+
```
|
|
71
|
+
|
|
72
|
+
- **The check** is `(config, trace)`, `def` or `async def`, and returns a
|
|
73
|
+
`CheckResult`, a dict with a boolean `"passed"`, or a `bool`. A failed
|
|
74
|
+
result's `reason` is what the violation reports — make it say what was
|
|
75
|
+
wrong.
|
|
76
|
+
- **`trace`** is a `RunTrace` (snake_case here, camelCase on the wire):
|
|
77
|
+
`output` (the final answer text), `tool_calls` (`tool_id`,
|
|
78
|
+
`tool_name`, `arguments`), `tool_results` (`tool_call_id`, `output`),
|
|
79
|
+
`model_calls` (`provider_id`, `model`, tokens), `user_input`,
|
|
80
|
+
`agent_id`, `conversation_id`, `turn_number`, `total_cost_usd`,
|
|
81
|
+
`duration_ms`, `mode` (`"runtime"` or `"ci"`). Annotate it `dict` to get
|
|
82
|
+
the raw wire dict instead.
|
|
83
|
+
- **`config`** — its type comes from the first parameter's annotation
|
|
84
|
+
(or `config_type=`) and becomes the guardrail's config schema. The
|
|
85
|
+
values are `config=` on the decorator, keyed as on the wire (the
|
|
86
|
+
model's aliases): `@guardrail(..., config={"minLookups": 2})`. They are
|
|
87
|
+
checked against the type where declared, go into the index, and the
|
|
88
|
+
check runs with them. Without `config=` the check runs with `{}` — so
|
|
89
|
+
give every field a default; a required field without a value fails
|
|
90
|
+
every evaluation (`input-validation-failed`).
|
|
91
|
+
- An exception in the check fails the evaluation (`handler-throw`);
|
|
92
|
+
return a failed `CheckResult` for a rule that isn't met.
|
|
93
|
+
- A check gets no model and no provider: it can't call an LLM. Keep it a
|
|
94
|
+
pure function of the trace (fast, deterministic, free).
|
|
95
|
+
|
|
96
|
+
## `@guardrail(...)`
|
|
97
|
+
|
|
98
|
+
- **`id`** — `<pack-id>.<guardrail-name>`, kebab-case. Name the rule as
|
|
99
|
+
a positive assertion: `no-fabricated-quotes`, `response-not-empty`.
|
|
100
|
+
- **`on_violation`** — the action: `"halt"`, `"retry"`, `"escalate"`,
|
|
101
|
+
`"log-only"`, `"compensate"`. For one that needs settings pass the
|
|
102
|
+
whole object with `action=` instead (exactly one of the two):
|
|
103
|
+
`action={"on-violation": "retry", "retry": {"maxAttempts": 2}}`,
|
|
104
|
+
`{"on-violation": "escalate", "escalateTo": …}`,
|
|
105
|
+
`{"on-violation": "compensate", "compensateWith": "<tool id>"}`.
|
|
106
|
+
In an agent turn a failed `halt` guardrail fails the turn
|
|
107
|
+
(`guardrail-violation`) and the answer is not stored; any other action
|
|
108
|
+
reports the failure in the turn result's `violations` and the turn
|
|
109
|
+
completes. In 0.1 the runtime acts only on `halt`: `retry`, `escalate`
|
|
110
|
+
and `compensate` are recorded on the violation, with no second attempt,
|
|
111
|
+
escalation or compensating call.
|
|
112
|
+
- **`severity`** — `"info"`, `"warn"`, `"error"` (default), `"critical"`.
|
|
113
|
+
Independent of the action: dashboards group by severity, execution
|
|
114
|
+
follows the action.
|
|
115
|
+
- **`scope`** — when it applies: `{"when": "always" | "ci-only" |
|
|
116
|
+
"runtime-only"}`, narrowed by `agents`, `flows`, `tenants` lists.
|
|
117
|
+
- **`kind`** — `"zero-llm"` (default): a check over the trace — what a
|
|
118
|
+
pack writes.
|
|
119
|
+
- **`config`** — the values the check runs with (above); **`config_type`**
|
|
120
|
+
— the config's type when the check's first parameter isn't annotated
|
|
121
|
+
with it.
|
|
122
|
+
- **`name`** — a display name. **`check_id`** — defaults to the id.
|
|
123
|
+
- `sandbox=`, `limits=`, `network=` are recorded in the index.
|
|
124
|
+
|
|
125
|
+
## Testing
|
|
126
|
+
|
|
127
|
+
A `Guardrail` is still callable:
|
|
128
|
+
|
|
129
|
+
```python
|
|
130
|
+
# tests/test_guardrails.py
|
|
131
|
+
from kindgi import RunTrace
|
|
132
|
+
from guardrails.citations import Config, no_fabricated_quotes
|
|
133
|
+
|
|
134
|
+
|
|
135
|
+
def test_no_lookups_fails():
|
|
136
|
+
trace = RunTrace(run_id="r", tenant_id="t", output="As held in Smith v. Jones…")
|
|
137
|
+
assert not no_fabricated_quotes(Config(), trace).passed
|
|
138
|
+
```
|
|
139
|
+
|
|
140
|
+
`RunTrace(...)` takes snake_case fields; `tool_calls` entries are
|
|
141
|
+
`ToolCallRecord(tool_id=…, tool_name=…, arguments={…}, at="…")`.
|
|
142
|
+
|
|
143
|
+
## Wiring onto an agent
|
|
144
|
+
|
|
145
|
+
```python
|
|
146
|
+
from ..guardrails.citations import no_fabricated_quotes
|
|
147
|
+
|
|
148
|
+
brief_writer = Agent(..., guardrails=[no_fabricated_quotes])
|
|
149
|
+
```
|
|
150
|
+
|
|
151
|
+
The `Guardrail` object (or its id). `kindgi dev` registers the pack's
|
|
152
|
+
guardrails; an agent naming an id with no registered guardrail fails the
|
|
153
|
+
turn before the model is called (`Error [invalid-request]: Agent "…"
|
|
154
|
+
references guardrails not in the registry: <id>`).
|
|
155
|
+
|
|
156
|
+
## Common mistakes
|
|
157
|
+
|
|
158
|
+
1. **A required config field with no `config=` value.** Without
|
|
159
|
+
`config=` the check runs with `{}`; give the field a default or the
|
|
160
|
+
guardrail its values.
|
|
161
|
+
2. **Snake_case keys in `config=`.** It is keyed like the wire — the
|
|
162
|
+
model's aliases (`{"minLookups": 2}`), not the field names.
|
|
163
|
+
3. **Both `on_violation=` and `action=`, or neither** — `DefinitionError`.
|
|
164
|
+
4. **Expecting another attempt.** `halt` stops the turn, and in 0.1
|
|
165
|
+
`retry` doesn't run the turn again: it's only recorded.
|
|
166
|
+
5. **Calling a model from the check.** Not available; keep checks pure.
|
|
167
|
+
6. **Raising for a broken rule.** Return `CheckResult(passed=False,
|
|
168
|
+
reason=…)`; an exception is an evaluation error, not a violation.
|
|
169
|
+
7. **Defining the guardrail inside a function** — only module-level
|
|
170
|
+
primitives are indexed.
|
|
171
|
+
|
|
172
|
+
## When the framework itself is the problem
|
|
173
|
+
|
|
174
|
+
If the bug is in Kindgi or the `kindgi` package (a trace field missing,
|
|
175
|
+
a misleading error) and not in the check, load
|
|
176
|
+
`kindgi-framework-feedback` and file it with `kindgi feedback write`.
|