@kindgi/api 0.1.2 → 0.1.4-rc.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/agent-binding.d.ts +18 -4
- package/dist/agent-binding.d.ts.map +1 -1
- package/dist/agent-pins.d.ts +48 -0
- package/dist/agent-pins.d.ts.map +1 -0
- package/dist/agent-pins.js +102 -0
- package/dist/agent-pins.js.map +1 -0
- package/dist/app.d.ts +22 -0
- package/dist/app.d.ts.map +1 -1
- package/dist/app.js +22 -3
- package/dist/app.js.map +1 -1
- package/dist/block-binding.d.ts +132 -0
- package/dist/block-binding.d.ts.map +1 -0
- package/dist/block-binding.js +4 -0
- package/dist/block-binding.js.map +1 -0
- package/dist/block-pins.d.ts +22 -0
- package/dist/block-pins.d.ts.map +1 -0
- package/dist/block-pins.js +112 -0
- package/dist/block-pins.js.map +1 -0
- package/dist/cost-binding.d.ts +100 -5
- package/dist/cost-binding.d.ts.map +1 -1
- package/dist/cost-binding.js +10 -0
- package/dist/cost-binding.js.map +1 -1
- package/dist/deploy-versions.d.ts +58 -0
- package/dist/deploy-versions.d.ts.map +1 -0
- package/dist/deploy-versions.js +91 -0
- package/dist/deploy-versions.js.map +1 -0
- package/dist/deployment-binding.d.ts +24 -3
- package/dist/deployment-binding.d.ts.map +1 -1
- package/dist/derive-agent-version.d.ts +69 -0
- package/dist/derive-agent-version.d.ts.map +1 -0
- package/dist/derive-agent-version.js +139 -0
- package/dist/derive-agent-version.js.map +1 -0
- package/dist/errors.d.ts.map +1 -1
- package/dist/errors.js +17 -0
- package/dist/errors.js.map +1 -1
- package/dist/eval-case-binding.d.ts +65 -0
- package/dist/eval-case-binding.d.ts.map +1 -0
- package/dist/eval-case-binding.js +4 -0
- package/dist/eval-case-binding.js.map +1 -0
- package/dist/eval-run-binding.d.ts +30 -0
- package/dist/eval-run-binding.d.ts.map +1 -1
- package/dist/eval-run-dispatcher.d.ts +38 -3
- package/dist/eval-run-dispatcher.d.ts.map +1 -1
- package/dist/eval-run-dispatcher.js +21 -15
- package/dist/eval-run-dispatcher.js.map +1 -1
- package/dist/eval-suite-binding.d.ts +1 -1
- package/dist/eval-suite-binding.d.ts.map +1 -1
- package/dist/eval-suite-binding.js +2 -0
- package/dist/eval-suite-binding.js.map +1 -1
- package/dist/flow-binding.d.ts +10 -4
- package/dist/flow-binding.d.ts.map +1 -1
- package/dist/flow-pins.d.ts +36 -0
- package/dist/flow-pins.d.ts.map +1 -0
- package/dist/flow-pins.js +81 -0
- package/dist/flow-pins.js.map +1 -0
- package/dist/hitl-binding.d.ts +20 -5
- package/dist/hitl-binding.d.ts.map +1 -1
- package/dist/index.d.ts +19 -8
- package/dist/index.d.ts.map +1 -1
- package/dist/index.js +8 -2
- package/dist/index.js.map +1 -1
- package/dist/judged-dispatcher.d.ts +134 -0
- package/dist/judged-dispatcher.d.ts.map +1 -0
- package/dist/judged-dispatcher.js +297 -0
- package/dist/judged-dispatcher.js.map +1 -0
- package/dist/judged-items.d.ts +86 -0
- package/dist/judged-items.d.ts.map +1 -0
- package/dist/judged-items.js +184 -0
- package/dist/judged-items.js.map +1 -0
- package/dist/judgment-binding.d.ts +316 -0
- package/dist/judgment-binding.d.ts.map +1 -0
- package/dist/judgment-binding.js +19 -0
- package/dist/judgment-binding.js.map +1 -0
- package/dist/middleware/auth.d.ts +3 -2
- package/dist/middleware/auth.d.ts.map +1 -1
- package/dist/middleware/auth.js.map +1 -1
- package/dist/middleware/idempotency.d.ts +5 -1
- package/dist/middleware/idempotency.d.ts.map +1 -1
- package/dist/middleware/idempotency.js +8 -1
- package/dist/middleware/idempotency.js.map +1 -1
- package/dist/openapi/generate.d.ts.map +1 -1
- package/dist/openapi/generate.js +4 -1
- package/dist/openapi/generate.js.map +1 -1
- package/dist/openapi/operations.d.ts.map +1 -1
- package/dist/openapi/operations.js +543 -24
- package/dist/openapi/operations.js.map +1 -1
- package/dist/openapi/schemas.d.ts +58 -0
- package/dist/openapi/schemas.d.ts.map +1 -1
- package/dist/openapi/schemas.js +1058 -27
- package/dist/openapi/schemas.js.map +1 -1
- package/dist/provenance-binding.d.ts +27 -1
- package/dist/provenance-binding.d.ts.map +1 -1
- package/dist/provenance-binding.js.map +1 -1
- package/dist/provider-binding.d.ts +12 -7
- package/dist/provider-binding.d.ts.map +1 -1
- package/dist/reviewer-binding.d.ts +10 -2
- package/dist/reviewer-binding.d.ts.map +1 -1
- package/dist/reviewer-role.d.ts +13 -0
- package/dist/reviewer-role.d.ts.map +1 -0
- package/dist/reviewer-role.js +27 -0
- package/dist/reviewer-role.js.map +1 -0
- package/dist/routes/agents.d.ts +9 -1
- package/dist/routes/agents.d.ts.map +1 -1
- package/dist/routes/agents.js +175 -11
- package/dist/routes/agents.js.map +1 -1
- package/dist/routes/approvals.d.ts +5 -4
- package/dist/routes/approvals.d.ts.map +1 -1
- package/dist/routes/approvals.js +52 -16
- package/dist/routes/approvals.js.map +1 -1
- package/dist/routes/blocks.d.ts +19 -0
- package/dist/routes/blocks.d.ts.map +1 -0
- package/dist/routes/blocks.js +281 -0
- package/dist/routes/blocks.js.map +1 -0
- package/dist/routes/conversations.d.ts +7 -1
- package/dist/routes/conversations.d.ts.map +1 -1
- package/dist/routes/conversations.js +41 -1
- package/dist/routes/conversations.js.map +1 -1
- package/dist/routes/cost.d.ts.map +1 -1
- package/dist/routes/cost.js +142 -11
- package/dist/routes/cost.js.map +1 -1
- package/dist/routes/deployments.d.ts +3 -0
- package/dist/routes/deployments.d.ts.map +1 -1
- package/dist/routes/deployments.js +208 -55
- package/dist/routes/deployments.js.map +1 -1
- package/dist/routes/eval-comparison.d.ts +14 -0
- package/dist/routes/eval-comparison.d.ts.map +1 -0
- package/dist/routes/eval-comparison.js +87 -0
- package/dist/routes/eval-comparison.js.map +1 -0
- package/dist/routes/eval-runs.d.ts.map +1 -1
- package/dist/routes/eval-runs.js +8 -0
- package/dist/routes/eval-runs.js.map +1 -1
- package/dist/routes/flows.d.ts +13 -1
- package/dist/routes/flows.d.ts.map +1 -1
- package/dist/routes/flows.js +44 -3
- package/dist/routes/flows.js.map +1 -1
- package/dist/routes/hierarchy-errors.d.ts +35 -0
- package/dist/routes/hierarchy-errors.d.ts.map +1 -0
- package/dist/routes/hierarchy-errors.js +39 -0
- package/dist/routes/hierarchy-errors.js.map +1 -0
- package/dist/routes/identity.d.ts +7 -0
- package/dist/routes/identity.d.ts.map +1 -1
- package/dist/routes/identity.js +3 -2
- package/dist/routes/identity.js.map +1 -1
- package/dist/routes/judged-suites.d.ts +20 -0
- package/dist/routes/judged-suites.d.ts.map +1 -0
- package/dist/routes/judged-suites.js +272 -0
- package/dist/routes/judged-suites.js.map +1 -0
- package/dist/routes/judgment-context.d.ts +22 -0
- package/dist/routes/judgment-context.d.ts.map +1 -0
- package/dist/routes/judgment-context.js +88 -0
- package/dist/routes/judgment-context.js.map +1 -0
- package/dist/routes/judgment-flow-context.d.ts +32 -0
- package/dist/routes/judgment-flow-context.d.ts.map +1 -0
- package/dist/routes/judgment-flow-context.js +195 -0
- package/dist/routes/judgment-flow-context.js.map +1 -0
- package/dist/routes/judgments.d.ts +41 -0
- package/dist/routes/judgments.d.ts.map +1 -0
- package/dist/routes/judgments.js +566 -0
- package/dist/routes/judgments.js.map +1 -0
- package/dist/routes/orgs.d.ts +5 -2
- package/dist/routes/orgs.d.ts.map +1 -1
- package/dist/routes/orgs.js +38 -22
- package/dist/routes/orgs.js.map +1 -1
- package/dist/routes/policies.d.ts.map +1 -1
- package/dist/routes/policies.js +12 -1
- package/dist/routes/policies.js.map +1 -1
- package/dist/routes/projects.d.ts +10 -2
- package/dist/routes/projects.d.ts.map +1 -1
- package/dist/routes/projects.js +87 -79
- package/dist/routes/projects.js.map +1 -1
- package/dist/routes/provenance.d.ts.map +1 -1
- package/dist/routes/provenance.js +32 -1
- package/dist/routes/provenance.js.map +1 -1
- package/dist/routes/providers.d.ts.map +1 -1
- package/dist/routes/providers.js +6 -1
- package/dist/routes/providers.js.map +1 -1
- package/dist/routes/runs.d.ts +0 -7
- package/dist/routes/runs.d.ts.map +1 -1
- package/dist/routes/runs.js +65 -20
- package/dist/routes/runs.js.map +1 -1
- package/dist/routes/scope-params.d.ts +16 -1
- package/dist/routes/scope-params.d.ts.map +1 -1
- package/dist/routes/scope-params.js +28 -0
- package/dist/routes/scope-params.js.map +1 -1
- package/dist/routes/teams.d.ts +6 -2
- package/dist/routes/teams.d.ts.map +1 -1
- package/dist/routes/teams.js +77 -73
- package/dist/routes/teams.js.map +1 -1
- package/dist/types.d.ts +4 -3
- package/dist/types.d.ts.map +1 -1
- package/dist/webhook-endpoint-binding.d.ts +11 -0
- package/dist/webhook-endpoint-binding.d.ts.map +1 -1
- package/dist/webhook-endpoint-binding.js.map +1 -1
- package/openapi.json +13316 -9299
- package/package.json +21 -21
- package/src/agent-binding.ts +19 -4
- package/src/agent-pins.ts +147 -0
- package/src/app.ts +76 -3
- package/src/block-binding.ts +137 -0
- package/src/block-pins.ts +148 -0
- package/src/cost-binding.ts +116 -5
- package/src/deploy-versions.ts +157 -0
- package/src/deployment-binding.ts +27 -3
- package/src/derive-agent-version.ts +206 -0
- package/src/errors.ts +17 -0
- package/src/eval-case-binding.ts +71 -0
- package/src/eval-run-binding.ts +33 -0
- package/src/eval-run-dispatcher.ts +57 -16
- package/src/eval-suite-binding.ts +2 -0
- package/src/flow-binding.ts +11 -4
- package/src/flow-pins.ts +113 -0
- package/src/hitl-binding.ts +20 -4
- package/src/index.ts +88 -2
- package/src/judged-dispatcher.ts +507 -0
- package/src/judged-items.ts +263 -0
- package/src/judgment-binding.ts +349 -0
- package/src/middleware/auth.ts +3 -2
- package/src/middleware/idempotency.ts +7 -1
- package/src/openapi/generate.ts +7 -1
- package/src/openapi/operations.ts +615 -24
- package/src/openapi/schemas.ts +1157 -22
- package/src/provenance-binding.ts +42 -1
- package/src/provider-binding.ts +12 -7
- package/src/reviewer-binding.ts +10 -2
- package/src/reviewer-role.ts +35 -0
- package/src/routes/agents.ts +243 -19
- package/src/routes/approvals.ts +70 -19
- package/src/routes/blocks.ts +362 -0
- package/src/routes/conversations.ts +57 -1
- package/src/routes/cost.ts +159 -21
- package/src/routes/deployments.ts +266 -56
- package/src/routes/eval-comparison.ts +101 -0
- package/src/routes/eval-runs.ts +11 -0
- package/src/routes/flows.ts +63 -5
- package/src/routes/hierarchy-errors.ts +51 -0
- package/src/routes/identity.ts +10 -2
- package/src/routes/judged-suites.ts +363 -0
- package/src/routes/judgment-context.ts +128 -0
- package/src/routes/judgment-flow-context.ts +245 -0
- package/src/routes/judgments.ts +743 -0
- package/src/routes/orgs.ts +44 -27
- package/src/routes/policies.ts +19 -0
- package/src/routes/projects.ts +106 -95
- package/src/routes/provenance.ts +48 -1
- package/src/routes/providers.ts +5 -0
- package/src/routes/runs.ts +79 -22
- package/src/routes/scope-params.ts +35 -1
- package/src/routes/teams.ts +96 -90
- package/src/types.ts +4 -3
- package/src/webhook-endpoint-binding.ts +11 -0
|
@@ -0,0 +1,263 @@
|
|
|
1
|
+
// SPDX-License-Identifier: Apache-2.0
|
|
2
|
+
// Copyright (C) 2026 Kindgi Inc.
|
|
3
|
+
|
|
4
|
+
/**
|
|
5
|
+
* The items of a run's output that people judge, and how a new output is
|
|
6
|
+
* scored against the judgments of an old one. Pure, with no dependencies,
|
|
7
|
+
* so a console can derive items the same way.
|
|
8
|
+
*
|
|
9
|
+
* Items:
|
|
10
|
+
* - an agent turn's answer (its last agent message) is `answer`;
|
|
11
|
+
* - a flow run's whole output is `output`;
|
|
12
|
+
* - each element of a list in the output (an agent's typed result sits
|
|
13
|
+
* under `output`) is an item, keyed by its own `id` or `key` when it
|
|
14
|
+
* has one, else by where it sits (a JSON Pointer), with its rank.
|
|
15
|
+
*
|
|
16
|
+
* A judgment carries over to a new output's item when it's the same item:
|
|
17
|
+
* an item keyed by its own id is the same item when the key matches; any
|
|
18
|
+
* other (the answer, an element keyed by its place) only when its content
|
|
19
|
+
* is the same too. So a changed answer is a new item, with no judgments.
|
|
20
|
+
*/
|
|
21
|
+
|
|
22
|
+
export interface OutputItem {
|
|
23
|
+
/** The key judgments use for this item. */
|
|
24
|
+
readonly key: string;
|
|
25
|
+
/** Where it sits in the output (JSON Pointer). */
|
|
26
|
+
readonly pointer: string;
|
|
27
|
+
/** Its place in its list (0 = first). */
|
|
28
|
+
readonly rank?: number;
|
|
29
|
+
/** `true` when the key is the item's own id, so it identifies the item by itself. */
|
|
30
|
+
readonly ownKey: boolean;
|
|
31
|
+
readonly value: unknown;
|
|
32
|
+
}
|
|
33
|
+
|
|
34
|
+
const MAX_DEPTH = 3;
|
|
35
|
+
const MAX_LIST = 100;
|
|
36
|
+
|
|
37
|
+
function obj(value: unknown): Record<string, unknown> | undefined {
|
|
38
|
+
return typeof value === 'object' && value !== null && !Array.isArray(value)
|
|
39
|
+
? (value as Record<string, unknown>)
|
|
40
|
+
: undefined;
|
|
41
|
+
}
|
|
42
|
+
|
|
43
|
+
function escapeToken(token: string): string {
|
|
44
|
+
return token.replace(/~/g, '~0').replace(/\//g, '~1');
|
|
45
|
+
}
|
|
46
|
+
|
|
47
|
+
function unescapeToken(token: string): string {
|
|
48
|
+
return token.replace(/~1/g, '/').replace(/~0/g, '~');
|
|
49
|
+
}
|
|
50
|
+
|
|
51
|
+
/** An element's own id, if it has a usable one. */
|
|
52
|
+
function ownId(element: unknown): string | undefined {
|
|
53
|
+
const o = obj(element);
|
|
54
|
+
for (const field of ['id', 'key']) {
|
|
55
|
+
const v = o?.[field];
|
|
56
|
+
if (typeof v === 'string' && v.length > 0) return v;
|
|
57
|
+
if (typeof v === 'number') return String(v);
|
|
58
|
+
}
|
|
59
|
+
return undefined;
|
|
60
|
+
}
|
|
61
|
+
|
|
62
|
+
function listItems(value: unknown, pointer: string, depth: number): OutputItem[] {
|
|
63
|
+
if (depth > MAX_DEPTH) return [];
|
|
64
|
+
if (Array.isArray(value)) {
|
|
65
|
+
if (value.length === 0 || value.length > MAX_LIST) return [];
|
|
66
|
+
const seen = new Set<string>();
|
|
67
|
+
return value.map((element, i) => {
|
|
68
|
+
const id = ownId(element);
|
|
69
|
+
// Duplicate ids within one list fall back to the element's place.
|
|
70
|
+
const own = id !== undefined && !seen.has(id);
|
|
71
|
+
const key = own ? id : `${pointer}/${i}`;
|
|
72
|
+
seen.add(key);
|
|
73
|
+
return { key, pointer: `${pointer}/${i}`, rank: i, ownKey: own, value: element };
|
|
74
|
+
});
|
|
75
|
+
}
|
|
76
|
+
const o = obj(value);
|
|
77
|
+
if (o === undefined) return [];
|
|
78
|
+
return Object.entries(o).flatMap(([field, child]) =>
|
|
79
|
+
listItems(child, `${pointer}/${escapeToken(field)}`, depth + 1),
|
|
80
|
+
);
|
|
81
|
+
}
|
|
82
|
+
|
|
83
|
+
function turnAnswer(output: Record<string, unknown>): OutputItem | undefined {
|
|
84
|
+
const appended = Array.isArray(output.appended) ? output.appended : [];
|
|
85
|
+
for (let i = appended.length - 1; i >= 0; i--) {
|
|
86
|
+
const m = obj(appended[i]);
|
|
87
|
+
if (m?.role === 'agent' && m.content !== undefined && !hasToolCalls(m.content)) {
|
|
88
|
+
return { key: 'answer', pointer: `/appended/${i}/content`, ownKey: false, value: m.content };
|
|
89
|
+
}
|
|
90
|
+
}
|
|
91
|
+
return undefined;
|
|
92
|
+
}
|
|
93
|
+
|
|
94
|
+
function hasToolCalls(content: unknown): boolean {
|
|
95
|
+
return Array.isArray(obj(content)?.toolCalls);
|
|
96
|
+
}
|
|
97
|
+
|
|
98
|
+
/** The items of a run's output (`agentTurn`: an agent turn's result). */
|
|
99
|
+
export function outputItems(output: unknown, agentTurn: boolean): readonly OutputItem[] {
|
|
100
|
+
const o = obj(output);
|
|
101
|
+
if (agentTurn) {
|
|
102
|
+
if (o === undefined) return [];
|
|
103
|
+
const answer = turnAnswer(o);
|
|
104
|
+
return [
|
|
105
|
+
...(answer !== undefined ? [answer] : []),
|
|
106
|
+
...(o.output !== undefined ? listItems(o.output, '/output', 0) : []),
|
|
107
|
+
];
|
|
108
|
+
}
|
|
109
|
+
if (output === undefined || output === null) return [];
|
|
110
|
+
return [
|
|
111
|
+
{ key: 'output', pointer: '', ownKey: false, value: output },
|
|
112
|
+
...listItems(output, '', 0),
|
|
113
|
+
];
|
|
114
|
+
}
|
|
115
|
+
|
|
116
|
+
/** The value at a JSON Pointer in `doc`, or `undefined` when there's none. */
|
|
117
|
+
export function valueAt(doc: unknown, pointer: string): unknown {
|
|
118
|
+
if (pointer === '') return doc;
|
|
119
|
+
if (!pointer.startsWith('/')) return undefined;
|
|
120
|
+
let current: unknown = doc;
|
|
121
|
+
for (const raw of pointer.slice(1).split('/')) {
|
|
122
|
+
const token = unescapeToken(raw);
|
|
123
|
+
if (Array.isArray(current)) {
|
|
124
|
+
const i = Number(token);
|
|
125
|
+
current = Number.isInteger(i) ? current[i] : undefined;
|
|
126
|
+
} else {
|
|
127
|
+
current = obj(current)?.[token];
|
|
128
|
+
}
|
|
129
|
+
if (current === undefined) return undefined;
|
|
130
|
+
}
|
|
131
|
+
return current;
|
|
132
|
+
}
|
|
133
|
+
|
|
134
|
+
/** JSON with sorted keys, so equal values compare equal. */
|
|
135
|
+
function canonical(value: unknown): string {
|
|
136
|
+
if (Array.isArray(value)) return `[${value.map(canonical).join(',')}]`;
|
|
137
|
+
const o = obj(value);
|
|
138
|
+
if (o !== undefined) {
|
|
139
|
+
return `{${Object.keys(o)
|
|
140
|
+
.sort()
|
|
141
|
+
.filter((k) => o[k] !== undefined)
|
|
142
|
+
.map((k) => `${JSON.stringify(k)}:${canonical(o[k])}`)
|
|
143
|
+
.join(',')}}`;
|
|
144
|
+
}
|
|
145
|
+
return JSON.stringify(value) ?? 'null';
|
|
146
|
+
}
|
|
147
|
+
|
|
148
|
+
/** What people said about one judged item (as a judged test-set case keeps it). */
|
|
149
|
+
export interface ItemJudgments {
|
|
150
|
+
readonly key: string;
|
|
151
|
+
readonly pointer?: string;
|
|
152
|
+
readonly rank?: number;
|
|
153
|
+
readonly yesWeight: number;
|
|
154
|
+
readonly totalWeight: number;
|
|
155
|
+
}
|
|
156
|
+
|
|
157
|
+
/** A new output's item and the judged item it is the same as, if any. */
|
|
158
|
+
export interface MatchedItem {
|
|
159
|
+
readonly item: OutputItem;
|
|
160
|
+
readonly judged?: ItemJudgments;
|
|
161
|
+
}
|
|
162
|
+
|
|
163
|
+
/**
|
|
164
|
+
* Match a new output's items to the items judged in a past output
|
|
165
|
+
* (`judgedOutput`, where the judged items' pointers point).
|
|
166
|
+
*/
|
|
167
|
+
export function matchJudged(
|
|
168
|
+
items: readonly OutputItem[],
|
|
169
|
+
judged: readonly ItemJudgments[],
|
|
170
|
+
judgedOutput: unknown,
|
|
171
|
+
): readonly MatchedItem[] {
|
|
172
|
+
const byKey = new Map(judged.map((j) => [j.key, j]));
|
|
173
|
+
return items.map((item) => {
|
|
174
|
+
const j = byKey.get(item.key);
|
|
175
|
+
if (j === undefined) return { item };
|
|
176
|
+
if (item.ownKey) return { item, judged: j };
|
|
177
|
+
const before = j.pointer === undefined ? undefined : valueAt(judgedOutput, j.pointer);
|
|
178
|
+
return before !== undefined && canonical(before) === canonical(item.value)
|
|
179
|
+
? { item, judged: j }
|
|
180
|
+
: { item };
|
|
181
|
+
});
|
|
182
|
+
}
|
|
183
|
+
|
|
184
|
+
/** An output scored against judgments: the sums the metrics are made of. */
|
|
185
|
+
export interface OutputScore {
|
|
186
|
+
/** Σ yesWeight and Σ totalWeight over the output's judged items. */
|
|
187
|
+
readonly yesWeight: number;
|
|
188
|
+
readonly totalWeight: number;
|
|
189
|
+
readonly items: number;
|
|
190
|
+
readonly judgedItems: number;
|
|
191
|
+
/** The same sums over the judged items among the first `k` ranked items. */
|
|
192
|
+
readonly topK: { readonly yesWeight: number; readonly totalWeight: number };
|
|
193
|
+
}
|
|
194
|
+
|
|
195
|
+
export function scoreItems(matched: readonly MatchedItem[], k: number): OutputScore {
|
|
196
|
+
let yesWeight = 0;
|
|
197
|
+
let totalWeight = 0;
|
|
198
|
+
let judgedItems = 0;
|
|
199
|
+
for (const m of matched) {
|
|
200
|
+
if (m.judged === undefined) continue;
|
|
201
|
+
judgedItems += 1;
|
|
202
|
+
yesWeight += m.judged.yesWeight;
|
|
203
|
+
totalWeight += m.judged.totalWeight;
|
|
204
|
+
}
|
|
205
|
+
const top = matched
|
|
206
|
+
.filter((m) => m.item.rank !== undefined)
|
|
207
|
+
.sort((a, b) => (a.item.rank ?? 0) - (b.item.rank ?? 0))
|
|
208
|
+
.slice(0, k);
|
|
209
|
+
const topK = { yesWeight: 0, totalWeight: 0 };
|
|
210
|
+
for (const m of top) {
|
|
211
|
+
if (m.judged === undefined) continue;
|
|
212
|
+
topK.yesWeight += m.judged.yesWeight;
|
|
213
|
+
topK.totalWeight += m.judged.totalWeight;
|
|
214
|
+
}
|
|
215
|
+
return { yesWeight, totalWeight, items: matched.length, judgedItems, topK };
|
|
216
|
+
}
|
|
217
|
+
|
|
218
|
+
/** How a new output's items moved against the judged ones. */
|
|
219
|
+
export interface ItemChanges {
|
|
220
|
+
/** Judged items the new output still has, with their rank before and now. */
|
|
221
|
+
readonly kept: readonly {
|
|
222
|
+
readonly key: string;
|
|
223
|
+
readonly rankBefore?: number;
|
|
224
|
+
readonly rank?: number;
|
|
225
|
+
}[];
|
|
226
|
+
/** Judged items the new output no longer has. */
|
|
227
|
+
readonly dropped: readonly { readonly key: string; readonly rankBefore?: number }[];
|
|
228
|
+
/** The new output's items no one judged (for experts to judge). */
|
|
229
|
+
readonly new: readonly {
|
|
230
|
+
readonly key: string;
|
|
231
|
+
readonly pointer: string;
|
|
232
|
+
readonly rank?: number;
|
|
233
|
+
}[];
|
|
234
|
+
}
|
|
235
|
+
|
|
236
|
+
export function itemChanges(
|
|
237
|
+
matched: readonly MatchedItem[],
|
|
238
|
+
judged: readonly ItemJudgments[],
|
|
239
|
+
): ItemChanges {
|
|
240
|
+
const keptKeys = new Set<string>();
|
|
241
|
+
const kept: ItemChanges['kept'][number][] = [];
|
|
242
|
+
const fresh: ItemChanges['new'][number][] = [];
|
|
243
|
+
for (const m of matched) {
|
|
244
|
+
if (m.judged !== undefined) {
|
|
245
|
+
keptKeys.add(m.judged.key);
|
|
246
|
+
kept.push({
|
|
247
|
+
key: m.judged.key,
|
|
248
|
+
...(m.judged.rank !== undefined && { rankBefore: m.judged.rank }),
|
|
249
|
+
...(m.item.rank !== undefined && { rank: m.item.rank }),
|
|
250
|
+
});
|
|
251
|
+
} else {
|
|
252
|
+
fresh.push({
|
|
253
|
+
key: m.item.key,
|
|
254
|
+
pointer: m.item.pointer,
|
|
255
|
+
...(m.item.rank !== undefined && { rank: m.item.rank }),
|
|
256
|
+
});
|
|
257
|
+
}
|
|
258
|
+
}
|
|
259
|
+
const dropped = judged
|
|
260
|
+
.filter((j) => !keptKeys.has(j.key))
|
|
261
|
+
.map((j) => ({ key: j.key, ...(j.rank !== undefined && { rankBefore: j.rank }) }));
|
|
262
|
+
return { kept, dropped, new: fresh };
|
|
263
|
+
}
|
|
@@ -0,0 +1,349 @@
|
|
|
1
|
+
// SPDX-License-Identifier: Apache-2.0
|
|
2
|
+
// Copyright (C) 2026 Kindgi Inc.
|
|
3
|
+
|
|
4
|
+
import type { Cursor, ListScope, ProjectId, TenantId } from '@kindgi/types';
|
|
5
|
+
|
|
6
|
+
/**
|
|
7
|
+
* Caller-plugged surface for judgments: a person says yes or no, with an
|
|
8
|
+
* optional reason, about one item of a run's output. A judgment may name a
|
|
9
|
+
* **judge class** ("expert", "user", …: the deployment's own names) that
|
|
10
|
+
* carries a weight, so later metrics can weigh judgments. Classes are
|
|
11
|
+
* optional: an unclassified judgment counts with weight 1, so judging
|
|
12
|
+
* needs no setup.
|
|
13
|
+
*
|
|
14
|
+
* A judgment keeps **copies** of what was judged: the run's input and
|
|
15
|
+
* output (once per run) and the judged item. The copies outlive the run's
|
|
16
|
+
* own retention, so a set of judged cases stays stable when the run is
|
|
17
|
+
* purged.
|
|
18
|
+
*
|
|
19
|
+
* Who judged: `assertedBy` is the authenticated caller (a user, or a
|
|
20
|
+
* service token such as an app's API key). It is never taken from the
|
|
21
|
+
* request body. When an app judges on behalf of one of its own users, it
|
|
22
|
+
* passes that user's opaque id as `participantId`.
|
|
23
|
+
*
|
|
24
|
+
* One live judgment per (run, item key, asserting principal, participant):
|
|
25
|
+
* recording another supersedes the previous one, which stays as history.
|
|
26
|
+
*
|
|
27
|
+
* Every method is tenant-scoped. Deletes are soft (`unregister`); purging
|
|
28
|
+
* follows the retention policies for the `judgment` and `judge_class`
|
|
29
|
+
* domains.
|
|
30
|
+
*/
|
|
31
|
+
export interface JudgmentRegistryBinding {
|
|
32
|
+
/** Create a judge class. `name-taken` when a live class of that scope has the name. */
|
|
33
|
+
createClass(input: JudgeClassCreateInput): Promise<JudgeClassCreateOutcome>;
|
|
34
|
+
/** Live classes, newest first, cursor-paginated; `scope` narrows to one scope. */
|
|
35
|
+
listClasses(input: JudgeClassListInput): Promise<JudgeClassPage>;
|
|
36
|
+
/**
|
|
37
|
+
* One class, or `null` when unknown. `includeUnregistered` also returns
|
|
38
|
+
* a retired class (judgments keep naming theirs).
|
|
39
|
+
*/
|
|
40
|
+
getClass(input: JudgeClassGetInput): Promise<JudgeClass | null>;
|
|
41
|
+
/** Change a live class's weight or description; `null` when unknown or retired. */
|
|
42
|
+
updateClass(input: JudgeClassUpdateInput): Promise<JudgeClass | null>;
|
|
43
|
+
/** Retire a class: no new judgments may name it. `false` when unknown or already retired. */
|
|
44
|
+
unregisterClass(input: JudgeClassGetInput): Promise<{ readonly unregistered: boolean }>;
|
|
45
|
+
|
|
46
|
+
/**
|
|
47
|
+
* Record a judgment. The first judgment of a run also stores the run's
|
|
48
|
+
* copy (`run`); later ones keep the stored copy. Supersedes the live
|
|
49
|
+
* judgment with the same run, item key, `assertedBy` and `participantId`.
|
|
50
|
+
*/
|
|
51
|
+
record(input: JudgmentRecordInput): Promise<Judgment>;
|
|
52
|
+
/** Live judgments, newest first, cursor-paginated. */
|
|
53
|
+
list(input: JudgmentListInput): Promise<JudgmentPage>;
|
|
54
|
+
/** One judgment (live or not) with its copies, or `null` when unknown. */
|
|
55
|
+
get(input: JudgmentGetInput): Promise<JudgmentWithCopies | null>;
|
|
56
|
+
/** Soft-delete a live judgment. `false` when unknown or already removed. */
|
|
57
|
+
unregister(input: JudgmentGetInput): Promise<{ readonly unregistered: boolean }>;
|
|
58
|
+
/**
|
|
59
|
+
* Judged runs with their copies and live judgments, newest first:
|
|
60
|
+
* what a test set is built from. Optional (a binding without it can't
|
|
61
|
+
* build test sets from judgments).
|
|
62
|
+
*/
|
|
63
|
+
listJudgedRuns?(input: JudgedRunListInput): Promise<JudgedRunPage>;
|
|
64
|
+
}
|
|
65
|
+
|
|
66
|
+
// ---------- judge classes ----------
|
|
67
|
+
|
|
68
|
+
/**
|
|
69
|
+
* Where a class applies: the whole tenant, one project, or one agent in a
|
|
70
|
+
* project. A judgment may name a class whose scope covers its run.
|
|
71
|
+
*/
|
|
72
|
+
export type JudgeClassScope =
|
|
73
|
+
| { readonly kind: 'tenant' }
|
|
74
|
+
| { readonly kind: 'project'; readonly projectId: ProjectId }
|
|
75
|
+
| { readonly kind: 'agent'; readonly projectId: ProjectId; readonly agentId: string };
|
|
76
|
+
|
|
77
|
+
export const JUDGE_CLASS_SCOPE_KINDS = ['tenant', 'project', 'agent'] as const;
|
|
78
|
+
export type JudgeClassScopeKind = (typeof JUDGE_CLASS_SCOPE_KINDS)[number];
|
|
79
|
+
|
|
80
|
+
export interface JudgeClass {
|
|
81
|
+
readonly id: string;
|
|
82
|
+
readonly tenantId: TenantId;
|
|
83
|
+
readonly scope: JudgeClassScope;
|
|
84
|
+
/** The deployment's own word for the class: "expert", "user", "arbitrator"… */
|
|
85
|
+
readonly name: string;
|
|
86
|
+
/** How much a judgment of this class counts (≥ 0); an unclassified judgment counts 1. */
|
|
87
|
+
readonly weight: number;
|
|
88
|
+
readonly description?: string;
|
|
89
|
+
readonly createdAt: string;
|
|
90
|
+
readonly updatedAt: string;
|
|
91
|
+
/** Set when the class was retired. */
|
|
92
|
+
readonly unregisteredAt?: string;
|
|
93
|
+
}
|
|
94
|
+
|
|
95
|
+
export interface JudgeClassCreateInput {
|
|
96
|
+
readonly tenantId: TenantId;
|
|
97
|
+
readonly scope: JudgeClassScope;
|
|
98
|
+
readonly name: string;
|
|
99
|
+
readonly weight: number;
|
|
100
|
+
readonly description?: string;
|
|
101
|
+
}
|
|
102
|
+
|
|
103
|
+
export type JudgeClassCreateOutcome =
|
|
104
|
+
| { readonly kind: 'created'; readonly judgeClass: JudgeClass }
|
|
105
|
+
| { readonly kind: 'name-taken' };
|
|
106
|
+
|
|
107
|
+
export interface JudgeClassListInput {
|
|
108
|
+
readonly tenantId: TenantId;
|
|
109
|
+
readonly scope?: JudgeClassScope;
|
|
110
|
+
readonly cursor?: Cursor;
|
|
111
|
+
readonly limit: number;
|
|
112
|
+
}
|
|
113
|
+
|
|
114
|
+
export interface JudgeClassPage {
|
|
115
|
+
readonly data: readonly JudgeClass[];
|
|
116
|
+
readonly hasMore: boolean;
|
|
117
|
+
readonly nextCursor?: Cursor;
|
|
118
|
+
}
|
|
119
|
+
|
|
120
|
+
export interface JudgeClassGetInput {
|
|
121
|
+
readonly tenantId: TenantId;
|
|
122
|
+
readonly judgeClassId: string;
|
|
123
|
+
readonly includeUnregistered?: boolean;
|
|
124
|
+
}
|
|
125
|
+
|
|
126
|
+
export interface JudgeClassUpdateInput {
|
|
127
|
+
readonly tenantId: TenantId;
|
|
128
|
+
readonly judgeClassId: string;
|
|
129
|
+
readonly weight?: number;
|
|
130
|
+
readonly description?: string;
|
|
131
|
+
}
|
|
132
|
+
|
|
133
|
+
/** Whether a class's scope covers a run in `projectId` whose subject is `subject`. */
|
|
134
|
+
export function judgeClassApplies(
|
|
135
|
+
scope: JudgeClassScope,
|
|
136
|
+
run: { readonly projectId: ProjectId; readonly subject: JudgedSubject },
|
|
137
|
+
): boolean {
|
|
138
|
+
switch (scope.kind) {
|
|
139
|
+
case 'tenant':
|
|
140
|
+
return true;
|
|
141
|
+
case 'project':
|
|
142
|
+
return scope.projectId === run.projectId;
|
|
143
|
+
case 'agent':
|
|
144
|
+
return (
|
|
145
|
+
scope.projectId === run.projectId &&
|
|
146
|
+
run.subject.kind === 'agent' &&
|
|
147
|
+
run.subject.id === scope.agentId
|
|
148
|
+
);
|
|
149
|
+
}
|
|
150
|
+
}
|
|
151
|
+
|
|
152
|
+
// ---------- judgments ----------
|
|
153
|
+
|
|
154
|
+
export const VERDICTS = ['yes', 'no'] as const;
|
|
155
|
+
export type Verdict = (typeof VERDICTS)[number];
|
|
156
|
+
|
|
157
|
+
/** What a judged run ran: an agent at a version, or a flow at a version. */
|
|
158
|
+
export interface JudgedSubject {
|
|
159
|
+
readonly kind: 'agent' | 'flow';
|
|
160
|
+
readonly id: string;
|
|
161
|
+
readonly version: string;
|
|
162
|
+
}
|
|
163
|
+
|
|
164
|
+
/** The judged item of a run's output. */
|
|
165
|
+
export interface JudgedItem {
|
|
166
|
+
/** The caller's stable id for the item, e.g. a matched case's id. */
|
|
167
|
+
readonly key: string;
|
|
168
|
+
/** Where the item is in the run's output, as a JSON Pointer (RFC 6901), e.g. `/matches/2`. */
|
|
169
|
+
readonly pointer?: string;
|
|
170
|
+
/** The item's position in a ranked list (0 = first). */
|
|
171
|
+
readonly rank?: number;
|
|
172
|
+
}
|
|
173
|
+
|
|
174
|
+
/** Who asserted a judgment: the authenticated caller, never a typed name. */
|
|
175
|
+
export interface JudgmentAssertedBy {
|
|
176
|
+
readonly kind: 'user' | 'service';
|
|
177
|
+
/** A user id, or for a service token its token or session id. */
|
|
178
|
+
readonly id: string;
|
|
179
|
+
}
|
|
180
|
+
|
|
181
|
+
export interface Judgment {
|
|
182
|
+
readonly id: string;
|
|
183
|
+
readonly tenantId: TenantId;
|
|
184
|
+
readonly projectId: ProjectId;
|
|
185
|
+
readonly runId: string;
|
|
186
|
+
readonly subject: JudgedSubject;
|
|
187
|
+
readonly item: JudgedItem;
|
|
188
|
+
readonly verdict: Verdict;
|
|
189
|
+
readonly reason?: string;
|
|
190
|
+
/** The judge's class; absent for an unclassified judgment (weight 1). */
|
|
191
|
+
readonly judgeClassId?: string;
|
|
192
|
+
readonly assertedBy: JudgmentAssertedBy;
|
|
193
|
+
/** The app's opaque id for its end user who judged, when an app judged on their behalf. */
|
|
194
|
+
readonly participantId?: string;
|
|
195
|
+
readonly createdAt: string;
|
|
196
|
+
/** Set when the judgment was removed or superseded. */
|
|
197
|
+
readonly unregisteredAt?: string;
|
|
198
|
+
/** The judgment that replaced this one. */
|
|
199
|
+
readonly supersededBy?: string;
|
|
200
|
+
}
|
|
201
|
+
|
|
202
|
+
/**
|
|
203
|
+
* What a judged agent turn read besides its input, captured when it was
|
|
204
|
+
* first judged, so the turn can later be replayed faithfully: the
|
|
205
|
+
* conversation before it and the context its retrievals returned. (Its
|
|
206
|
+
* tool calls and results, provider and model are in its output.)
|
|
207
|
+
*/
|
|
208
|
+
export interface JudgedRunContext {
|
|
209
|
+
/** The conversation's messages before the turn, oldest first (at most the last 200). */
|
|
210
|
+
readonly history?: readonly unknown[];
|
|
211
|
+
/** Whether older messages were left out of `history`. */
|
|
212
|
+
readonly historyTruncated?: boolean;
|
|
213
|
+
/** What the turn's retrievals returned. */
|
|
214
|
+
readonly retrieved?: unknown;
|
|
215
|
+
/**
|
|
216
|
+
* The reviewer's decision at the turn's session approval gate, when the
|
|
217
|
+
* turn waited on one. A replay of the turn follows it.
|
|
218
|
+
*/
|
|
219
|
+
readonly sessionApproval?: { readonly approved: boolean; readonly rationale?: string };
|
|
220
|
+
/** For a flow run: what it did (see `JudgedFlowContext`). */
|
|
221
|
+
readonly flow?: JudgedFlowContext;
|
|
222
|
+
}
|
|
223
|
+
|
|
224
|
+
/** One tool call a judged run made, and its result. */
|
|
225
|
+
export interface JudgedToolCall {
|
|
226
|
+
/** The run that made it: the flow run, a sub-flow's run, or an agent step's turn. */
|
|
227
|
+
readonly runId: string;
|
|
228
|
+
/** The tool node that made it, or the agent step whose turn did. */
|
|
229
|
+
readonly nodeId?: string;
|
|
230
|
+
/** The loop iteration, when the node is in a loop body. */
|
|
231
|
+
readonly scope?: string;
|
|
232
|
+
readonly toolId: string;
|
|
233
|
+
readonly arguments: unknown;
|
|
234
|
+
readonly result: unknown;
|
|
235
|
+
}
|
|
236
|
+
|
|
237
|
+
/** One agent step of a judged flow run: the turn it started. */
|
|
238
|
+
export interface JudgedFlowStep {
|
|
239
|
+
readonly runId: string;
|
|
240
|
+
readonly nodeId?: string;
|
|
241
|
+
readonly scope?: string;
|
|
242
|
+
readonly agentId: string;
|
|
243
|
+
readonly agentVersion: string;
|
|
244
|
+
/** What the turn's retrievals returned. */
|
|
245
|
+
readonly retrieved?: unknown;
|
|
246
|
+
}
|
|
247
|
+
|
|
248
|
+
/**
|
|
249
|
+
* What a judged flow run did, kept at its first judgment: every tool
|
|
250
|
+
* call it made with its result (at its tool nodes, in its agent steps'
|
|
251
|
+
* turns and in its sub-flows), at most 500, and its agent steps.
|
|
252
|
+
*/
|
|
253
|
+
export interface JudgedFlowContext {
|
|
254
|
+
readonly calls: readonly JudgedToolCall[];
|
|
255
|
+
readonly steps: readonly JudgedFlowStep[];
|
|
256
|
+
/** More calls were made than were kept. */
|
|
257
|
+
readonly truncated?: boolean;
|
|
258
|
+
}
|
|
259
|
+
|
|
260
|
+
/** The stored copies of a judged run. */
|
|
261
|
+
export interface JudgedRunCopy {
|
|
262
|
+
readonly runId: string;
|
|
263
|
+
readonly subject: JudgedSubject;
|
|
264
|
+
readonly input: unknown;
|
|
265
|
+
readonly output: unknown;
|
|
266
|
+
/** Absent for runs judged before context was captured, and for flow runs. */
|
|
267
|
+
readonly context?: JudgedRunContext;
|
|
268
|
+
readonly capturedAt: string;
|
|
269
|
+
}
|
|
270
|
+
|
|
271
|
+
export interface JudgmentWithCopies extends Judgment {
|
|
272
|
+
/** The run's input and output as they were when it was first judged. */
|
|
273
|
+
readonly run: JudgedRunCopy;
|
|
274
|
+
/** The judged item's value, when the judgment pointed at it. */
|
|
275
|
+
readonly itemValue?: unknown;
|
|
276
|
+
}
|
|
277
|
+
|
|
278
|
+
export interface JudgmentRecordInput {
|
|
279
|
+
readonly tenantId: TenantId;
|
|
280
|
+
readonly projectId: ProjectId;
|
|
281
|
+
readonly runId: string;
|
|
282
|
+
/** The run's copy, stored on its first judgment. */
|
|
283
|
+
readonly run: {
|
|
284
|
+
readonly subject: JudgedSubject;
|
|
285
|
+
readonly input: unknown;
|
|
286
|
+
readonly output: unknown;
|
|
287
|
+
readonly context?: JudgedRunContext;
|
|
288
|
+
};
|
|
289
|
+
readonly item: JudgedItem;
|
|
290
|
+
/** The item's value at `item.pointer`, resolved by the route from the run's output. */
|
|
291
|
+
readonly itemValue?: unknown;
|
|
292
|
+
readonly verdict: Verdict;
|
|
293
|
+
readonly reason?: string;
|
|
294
|
+
readonly judgeClassId?: string;
|
|
295
|
+
readonly assertedBy: JudgmentAssertedBy;
|
|
296
|
+
readonly participantId?: string;
|
|
297
|
+
}
|
|
298
|
+
|
|
299
|
+
export interface JudgmentListInput {
|
|
300
|
+
readonly tenantId: TenantId;
|
|
301
|
+
readonly scope?: ListScope;
|
|
302
|
+
readonly runId?: string;
|
|
303
|
+
readonly agentId?: string;
|
|
304
|
+
readonly agentVersion?: string;
|
|
305
|
+
readonly flowId?: string;
|
|
306
|
+
readonly verdict?: Verdict;
|
|
307
|
+
readonly judgeClassId?: string;
|
|
308
|
+
readonly participantId?: string;
|
|
309
|
+
readonly cursor?: Cursor;
|
|
310
|
+
readonly limit: number;
|
|
311
|
+
}
|
|
312
|
+
|
|
313
|
+
export interface JudgmentPage {
|
|
314
|
+
readonly data: readonly Judgment[];
|
|
315
|
+
readonly hasMore: boolean;
|
|
316
|
+
readonly nextCursor?: Cursor;
|
|
317
|
+
}
|
|
318
|
+
|
|
319
|
+
export interface JudgmentGetInput {
|
|
320
|
+
readonly tenantId: TenantId;
|
|
321
|
+
readonly judgmentId: string;
|
|
322
|
+
}
|
|
323
|
+
|
|
324
|
+
export interface JudgedRunListInput {
|
|
325
|
+
readonly tenantId: TenantId;
|
|
326
|
+
readonly projectId?: ProjectId;
|
|
327
|
+
readonly agentId?: string;
|
|
328
|
+
readonly agentVersion?: string;
|
|
329
|
+
readonly flowId?: string;
|
|
330
|
+
/** Runs first judged at or after this time (ISO 8601). */
|
|
331
|
+
readonly since?: string;
|
|
332
|
+
/** Runs first judged before this time (ISO 8601). */
|
|
333
|
+
readonly until?: string;
|
|
334
|
+
readonly cursor?: Cursor;
|
|
335
|
+
readonly limit: number;
|
|
336
|
+
}
|
|
337
|
+
|
|
338
|
+
export interface JudgedRunWithJudgments {
|
|
339
|
+
readonly projectId: ProjectId;
|
|
340
|
+
readonly run: JudgedRunCopy;
|
|
341
|
+
/** The run's live judgments. */
|
|
342
|
+
readonly judgments: readonly Judgment[];
|
|
343
|
+
}
|
|
344
|
+
|
|
345
|
+
export interface JudgedRunPage {
|
|
346
|
+
readonly data: readonly JudgedRunWithJudgments[];
|
|
347
|
+
readonly hasMore: boolean;
|
|
348
|
+
readonly nextCursor?: Cursor;
|
|
349
|
+
}
|
package/src/middleware/auth.ts
CHANGED
|
@@ -37,8 +37,9 @@ export interface TokenResolution {
|
|
|
37
37
|
* middleware surfaces it as `c.get('reviewerRole')` and approvals
|
|
38
38
|
* routes gate visibility / decisions on the `standard < senior < admin`
|
|
39
39
|
* hierarchy. Absent for tokens that were never provisioned with a
|
|
40
|
-
* reviewer role
|
|
41
|
-
*
|
|
40
|
+
* reviewer role: the approvals routes then take the role the reviewer
|
|
41
|
+
* roster gives the token's user (`ReviewerBinding.resolveReviewerRole`),
|
|
42
|
+
* and return `403 permission-denied` when it gives none.
|
|
42
43
|
*/
|
|
43
44
|
readonly reviewerRole?: ReviewerRole;
|
|
44
45
|
/**
|
|
@@ -63,7 +63,11 @@ const DEFAULT_TTL_MS = 24 * 60 * 60 * 1000;
|
|
|
63
63
|
* byte-identical (status, content-type, body).
|
|
64
64
|
* - Cached hit with different body hash → 409
|
|
65
65
|
* `idempotency-key-body-mismatch`.
|
|
66
|
-
* - Cache miss → run the handler, then store the response
|
|
66
|
+
* - Cache miss → run the handler, then store the response when the
|
|
67
|
+
* request took effect (a status below 400). A refusal (4xx) changed
|
|
68
|
+
* nothing, and a failure (5xx) may be transient: neither is stored, so
|
|
69
|
+
* a retry with the same key after fixing the cause runs again instead
|
|
70
|
+
* of replaying the old error.
|
|
67
71
|
*
|
|
68
72
|
* The middleware sits AFTER auth so `tenantId` is available; the
|
|
69
73
|
* cache key is namespaced by tenant so callers can't collide across
|
|
@@ -122,6 +126,8 @@ export function idempotencyMiddleware(
|
|
|
122
126
|
await next();
|
|
123
127
|
|
|
124
128
|
const res = c.res;
|
|
129
|
+
// Only what took effect is replayed; see the doc above.
|
|
130
|
+
if (res.status >= 400) return;
|
|
125
131
|
const savedBody = await res.clone().text();
|
|
126
132
|
const contentType = res.headers.get('Content-Type') ?? 'application/octet-stream';
|
|
127
133
|
await store.set(cacheKey, {
|
package/src/openapi/generate.ts
CHANGED
|
@@ -80,7 +80,13 @@ const TAG_DESCRIPTIONS: Readonly<Record<string, string>> = {
|
|
|
80
80
|
policies:
|
|
81
81
|
'Tenant policy catalog (list, get, versions, publish, unregister) — part of the admin control plane. Full versioned CRUD, mirrors `flows` 1:1. Registry-only: enforcement is out of scope. `policyKind` selects the shape of `spec` — closed enum, extended additively (`access-control`, `model-routing`, `adapter-allowlist`, `rate-limit`, `retention`, `compliance`). Runtime consumers (e.g. model routing for `model-routing`, retention sweeps for `retention`) read policies from this store and apply them at their own boundary. Caller-plugged via `PolicyRegistryBinding`.',
|
|
82
82
|
'eval-suites':
|
|
83
|
-
'Evaluation suite catalog (list, get, versions, publish, unregister) — part of the admin control plane. Full versioned CRUD, mirrors `policies` 1:1. Registry-only: eval-run execution + per-kind grader dispatch live on the `eval-runs` tag. `evalKind` selects the shape of `spec` — closed enum, extended additively (`accuracy`, `pairwise`, `regression`, `human-review`, `benchmark`, `custom`). Runs against these suites are started through the `eval-runs` tag. Caller-plugged via `EvalSuiteRegistryBinding`.',
|
|
83
|
+
'Evaluation suite catalog (list, get, versions, publish, unregister) — part of the admin control plane. Full versioned CRUD, mirrors `policies` 1:1. Registry-only: eval-run execution + per-kind grader dispatch live on the `eval-runs` tag. `evalKind` selects the shape of `spec` — closed enum, extended additively (`accuracy`, `pairwise`, `regression`, `human-review`, `benchmark`, `custom`, `judged`). The cases of a `judged` suite are copies of judged runs built from judgments (`POST .../versions/from-judgments`, listed at `GET .../versions/{version}/cases`; `EvalCaseStoreBinding`). Runs against these suites are started through the `eval-runs` tag. Caller-plugged via `EvalSuiteRegistryBinding`.',
|
|
84
|
+
blocks:
|
|
85
|
+
"Data blocks: versioned prompts and settings an agent version pins when it's published (list, get, versions, publish, unregister, reinstate). A version never changes, nor does a block's kind; a settings block's values satisfy its schema and the latest version's. A block belongs to one project and is authorized through it: `read` on the project to read, `write` to publish, unregister or reinstate. Caller-plugged via `BlockRegistryBinding`.",
|
|
86
|
+
judgments:
|
|
87
|
+
"Judgments: yes or no, with an optional reason, about one item of a finished run's output, optionally recorded under a judge class (list, get, create, unregister). Each judgment keeps copies of what was judged (the run's input and output, and the item) so they outlive the run's own retention. Who judged comes from the authenticated caller, never the body; an app judging for one of its users passes that user's opaque id as `participantId`. One live judgment per run, item key, caller and participant: judging again supersedes the earlier one, which stays as history. Caller-plugged via `JudgmentRegistryBinding`.",
|
|
88
|
+
'judge-classes':
|
|
89
|
+
'Judge classes (list, get, create, update, unregister): the deployment\'s named kinds of judge ("expert", "user", ...), each with a weight, scoped to the tenant, a project, or an agent in a project. A judgment may name a class; an unclassified judgment counts with weight 1. Caller-plugged via `JudgmentRegistryBinding`.',
|
|
84
90
|
'eval-runs':
|
|
85
91
|
'Eval-run data plane (start, list, get, cancel, events SSE). Dispatches an eval run against a registered suite. Dispatch is per-`EvalKind`; kinds without a registered dispatcher return `422 dispatcher-not-registered`. `result` is kind-specific opaque JSON on the wire — for `accuracy` it is `{ passCount, totalCount, meanScore, perCase[] }`. SSE events mirror run streaming: per-case progress followed by a terminal frame carrying the aggregate. Caller-plugged via `EvalRunBinding`.',
|
|
86
92
|
auth: 'OAuth 2.0 / OIDC identity providers + session lifecycle. Layers browser-based auth on top of the static bearer-token surface: bearer tokens continue to work byte-shape-identical; session tokens use the `kgi_sk_*` prefix so the same middleware routes both flavors. Providers are caller-plugged via `IdentityProviderBinding` (no baked-in list). Sessions persist via `SessionStoreBinding`. Code exchange is caller-supplied via `exchangeCode`. PKCE (S256) is mandatory. `clientSecretRef` is a REFERENCE — the plaintext secret never crosses the wire.',
|