scout-ai 1.2.3 → 2.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/.vimproject +138 -50
- data/README.md +171 -290
- data/Rakefile +17 -1
- data/VERSION +1 -1
- data/doc/Improvements.md +325 -0
- data/doc/StartHere.md +110 -0
- data/doc/developer/Architecture.md +126 -0
- data/doc/developer/Backends.md +199 -0
- data/doc/developer/ChatLifecycle.md +183 -0
- data/doc/developer/DelegationInternals.md +295 -0
- data/doc/developer/DesignPrinciples.md +245 -0
- data/doc/developer/PromptProcessing.md +292 -0
- data/doc/developer/Provenance.md +317 -0
- data/doc/user/BuildingAgents.md +345 -0
- data/doc/user/Cookbook.md +333 -0
- data/doc/user/CoreConcepts.md +181 -0
- data/doc/user/Delegation.md +191 -0
- data/doc/user/GettingStarted.md +159 -0
- data/doc/user/ManagingContext.md +163 -0
- data/doc/user/MultiAgentWorkflows.md +256 -0
- data/doc/user/Python.md +159 -0
- data/doc/user/RunningInference.md +200 -0
- data/doc/user/ToolCalling.md +193 -0
- data/doc/user/WritingChats.md +197 -0
- data/lib/scout/llm/agent/chat.rb +61 -11
- data/lib/scout/llm/agent/delegate.rb +274 -65
- data/lib/scout/llm/agent/iterate.rb +2 -2
- data/lib/scout/llm/agent/save.rb +273 -0
- data/lib/scout/llm/agent/workflow.rb +164 -0
- data/lib/scout/llm/agent.rb +86 -61
- data/lib/scout/llm/ask.rb +62 -17
- data/lib/scout/llm/backends/anthropic.rb +9 -2
- data/lib/scout/llm/backends/bedrock.rb +15 -3
- data/lib/scout/llm/backends/default.rb +183 -99
- data/lib/scout/llm/backends/glm.rb +58 -0
- data/lib/scout/llm/backends/huggingface.rb +196 -26
- data/lib/scout/llm/backends/ollama.rb +13 -1
- data/lib/scout/llm/backends/openai.rb +0 -2
- data/lib/scout/llm/backends/openwebui.rb +20 -13
- data/lib/scout/llm/backends/relay.rb +22 -22
- data/lib/scout/llm/backends/responses.rb +1 -1
- data/lib/scout/llm/chat/agent_meta.rb +264 -0
- data/lib/scout/llm/chat/annotation.rb +39 -10
- data/lib/scout/llm/chat/parse.rb +28 -6
- data/lib/scout/llm/chat/persist.rb +25 -0
- data/lib/scout/llm/chat/process/clear.rb +41 -6
- data/lib/scout/llm/chat/process/files.rb +21 -6
- data/lib/scout/llm/chat/process/meta.rb +421 -34
- data/lib/scout/llm/chat/process/options.rb +21 -1
- data/lib/scout/llm/chat/process/tools.rb +56 -15
- data/lib/scout/llm/chat/process.rb +4 -0
- data/lib/scout/llm/chat/prompt/shorten_tools.rb +125 -0
- data/lib/scout/llm/chat/prompt/shorten_tools_epoch.rb +365 -0
- data/lib/scout/llm/chat/prompt.rb +48 -0
- data/lib/scout/llm/chat/provenance.rb +775 -0
- data/lib/scout/llm/chat/tool_calls.rb +76 -0
- data/lib/scout/llm/chat.rb +18 -2
- data/lib/scout/llm/embed.rb +11 -3
- data/lib/scout/llm/image.rb +86 -0
- data/lib/scout/llm/mcp.rb +10 -2
- data/lib/scout/llm/rag.rb +3 -3
- data/lib/scout/llm/tools/call.rb +160 -11
- data/lib/scout/llm/tools/knowledge_base.rb +1 -1
- data/lib/scout/llm/tools/workflow.rb +32 -16
- data/lib/scout/model/python/huggingface/causal.rb +23 -5
- data/lib/scout/model/python/huggingface.rb +2 -1
- data/lib/scout-ai.rb +1 -0
- data/python/README.md +197 -14
- data/python/scout_ai/huggingface/eval.py +245 -34
- data/python/tests/test_huggingface_eval.py +58 -0
- data/research/ChatAnalyst-required-changes.md +167 -0
- data/research/agent-delegation-analysis.md +810 -0
- data/research/agent-meta-provenance-integration-plan.md +622 -0
- data/research/agent-workflow-analysis.md +1120 -0
- data/research/backends-analysis.md +836 -0
- data/research/chat-core-analysis.md +946 -0
- data/research/chatanalyst-provenance/00-baseline.md +30 -0
- data/research/chatanalyst-provenance/01-repo-map.md +60 -0
- data/research/chatanalyst-provenance/02-event-reconstruction.md +55 -0
- data/research/chatanalyst-provenance/03-duplication-evidence.md +45 -0
- data/research/chatanalyst-provenance/04-tooling-root-cause.md +57 -0
- data/research/chatanalyst-provenance/05-fix-plan.md +46 -0
- data/research/chatanalyst-provenance/07-critic-review.md +25 -0
- data/research/chatanalyst-provenance/final-report.md +45 -0
- data/research/chatanalyst-provenance/resumption.md +37 -0
- data/research/coding-philosophy-analysis.md +928 -0
- data/research/commands-analysis.md +947 -0
- data/research/multi-agent-patterns-analysis.md +853 -0
- data/research/prompt-strategies-analysis.md +630 -0
- data/research/prov-verbosity-fix-notes.md +77 -0
- data/research/provenance-analysis.md +469 -0
- data/research/provenance-navigation-design.md +640 -0
- data/research/synthesis-report.md +487 -0
- data/research/tools-system-analysis.md +779 -0
- data/scout-ai.gemspec +100 -11
- data/scout_commands/agent/ask +13 -3
- data/scout_commands/agent/kb +2 -0
- data/scout_commands/llm/ask +11 -4
- data/scout_commands/llm/md +76 -0
- data/scout_commands/llm/process_queries +48 -0
- data/scout_commands/llm/prov +602 -0
- data/scout_commands/llm/word +71 -0
- data/scout_commands/workflow/mcp +43 -0
- data/share/word/reference.docx +0 -0
- data/test/etc/AI/mock.yaml +11 -0
- data/test/fixtures/backends/anthropic.json +19 -0
- data/test/fixtures/backends/anthropic_tool_use.json +24 -0
- data/test/fixtures/backends/bedrock.json +8 -0
- data/test/fixtures/backends/bedrock_embedding.json +3 -0
- data/test/fixtures/backends/bedrock_tool_use.json +17 -0
- data/test/fixtures/backends/ollama.json +16 -0
- data/test/fixtures/backends/ollama_tool_call.json +27 -0
- data/test/fixtures/backends/openai_chat.json +21 -0
- data/test/fixtures/backends/openai_chat_tool_call.json +31 -0
- data/test/fixtures/backends/responses.json +33 -0
- data/test/fixtures/backends/responses_tool_call.json +28 -0
- data/test/integration/README.md +32 -0
- data/test/integration/scout/llm/backends/test_endpoints.rb +34 -0
- data/test/integration/scout/llm/backends/test_openwebui.rb +61 -0
- data/test/integration/scout/llm/backends/test_relay.rb +52 -0
- data/test/integration/scout/llm/test_infrastructure.rb +74 -0
- data/test/{scout → integration/scout}/llm/test_mcp.rb +1 -1
- data/test/integration/scout/llm/tools/test_mcp.rb +42 -0
- data/test/integration/scout/model/test_base.rb +91 -0
- data/test/scout/llm/agent/test_chat.rb +8 -2
- data/test/scout/llm/agent/test_save.rb +413 -0
- data/test/scout/llm/agent/test_workflow.rb +110 -0
- data/test/scout/llm/backends/test_anthropic.rb +93 -10
- data/test/scout/llm/backends/test_bedrock.rb +118 -2
- data/test/scout/llm/backends/test_huggingface.rb +137 -42
- data/test/scout/llm/backends/test_ollama.rb +70 -20
- data/test/scout/llm/backends/test_openwebui.rb +42 -40
- data/test/scout/llm/backends/test_relay.rb +4 -2
- data/test/scout/llm/chat/agent_meta_fixtures.rb +131 -0
- data/test/scout/llm/chat/process/test_meta.rb +518 -0
- data/test/scout/llm/chat/process/test_normalize_usage.rb +183 -0
- data/test/scout/llm/chat/test_agent_meta.rb +357 -0
- data/test/scout/llm/chat/test_agent_meta_provenance.rb +467 -0
- data/test/scout/llm/chat/test_agent_meta_tokens.rb +594 -0
- data/test/scout/llm/chat/test_parse.rb +70 -15
- data/test/scout/llm/chat/test_prov_cli.rb +274 -0
- data/test/scout/llm/chat/test_provenance.rb +240 -0
- data/test/scout/llm/chat/test_tool_calls.rb +38 -0
- data/test/scout/llm/test_agent.rb +13 -36
- data/test/scout/llm/test_ask.rb +75 -52
- data/test/scout/llm/test_chat.rb +107 -13
- data/test/scout/llm/test_embed.rb +48 -0
- data/test/scout/llm/test_rag.rb +23 -16
- data/test/scout/llm/test_tools.rb +12 -1
- data/test/scout/llm/tools/test_knowledge_base.rb +0 -1
- data/test/scout/llm/tools/test_mcp.rb +5 -3
- data/test/scout/llm/tools/test_workflow.rb +23 -2
- data/test/scout/model/python/huggingface/causal/test_next_token.rb +11 -5
- data/test/scout/model/python/huggingface/test_causal.rb +9 -3
- data/test/scout/model/python/huggingface/test_classification.rb +11 -2
- data/test/scout/model/python/test_torch.rb +2 -0
- data/test/scout/model/python/torch/test_helpers.rb +4 -0
- data/test/scout/model/test_base.rb +4 -2
- data/test/support/availability.rb +231 -0
- data/test/support/fake_clients.rb +138 -0
- data/test/support/fixtures.rb +21 -0
- data/test/support/infrastructure_probes.rb +136 -0
- data/test/support/mock_backend.rb +215 -0
- data/test/test_helper.rb +32 -2
- metadata +99 -10
- data/doc/Agent.md +0 -327
- data/doc/Chat.md +0 -458
- data/doc/LLM.md +0 -340
- data/doc/RAG.md +0 -129
- data/scout_commands/documenter +0 -148
- data/test/scout/llm/backends/test_openai.rb +0 -192
- data/test/scout/llm/backends/test_responses.rb +0 -238
- data/test/scout/llm/test_parse.rb +0 -98
|
@@ -0,0 +1,622 @@
|
|
|
1
|
+
# Agent-meta provenance integration: implementation and test plan
|
|
2
|
+
|
|
3
|
+
> **Superseded (historical document).** Commit `efd8ebbc` ("Change how meta is
|
|
4
|
+
> passed in function_call_output. Auto-save agent chats in .files") replaced
|
|
5
|
+
> the serialized `agent_meta` writer shape described here with a `meta` key
|
|
6
|
+
> holding an Array of already-deserialized field Hashes. The reader still
|
|
7
|
+
> accepts both formats; see `doc/developer/Provenance.md` for current
|
|
8
|
+
> behavior. The rest of this file is kept verbatim as the design record.
|
|
9
|
+
|
|
10
|
+
> This is an implementation report for Scout-AI and ChatAnalyst coding agents.
|
|
11
|
+
> It describes a change to provenance handling for `agent_meta` receipts that
|
|
12
|
+
> are embedded in `function_call_output` records. It does not introduce a
|
|
13
|
+
> Session, ChatGraph, ProvenanceContext, or database abstraction.
|
|
14
|
+
|
|
15
|
+
## Decision summary
|
|
16
|
+
|
|
17
|
+
Scout-AI should treat `agent_meta` as **embedded provenance evidence**.
|
|
18
|
+
|
|
19
|
+
It must be included in two places:
|
|
20
|
+
|
|
21
|
+
1. **Structural traversal** must follow `job=` values found in an `agent_meta`
|
|
22
|
+
receipt, so a delegated agent whose answer is produced by a `chat_task`
|
|
23
|
+
exposes the actual producer Step, dependencies, logs, and result.
|
|
24
|
+
2. **Token accounting** must include direct token meta records found in
|
|
25
|
+
`agent_meta`, so socialized delegation remains auditable even when there is
|
|
26
|
+
no separately saved child chat.
|
|
27
|
+
|
|
28
|
+
It must not be represented as a third graph-node kind, converted into ordinary
|
|
29
|
+
parent-chat `meta:` messages, or assigned its own stateful context object.
|
|
30
|
+
|
|
31
|
+
The saved child chat and the receipt may both contain the same direct inference
|
|
32
|
+
metadata. They are two evidence locations for one inference event. Deduplicate
|
|
33
|
+
by `inference_id`, preserve both locations for auditing, and count the event
|
|
34
|
+
once.
|
|
35
|
+
|
|
36
|
+
## Current behavior and gap
|
|
37
|
+
|
|
38
|
+
`LLM.process_calls` already creates the required receipt. When a tool returns
|
|
39
|
+
an `LLM::Agent`, it:
|
|
40
|
+
|
|
41
|
+
1. runs `agent.chat(return_messages: true)`;
|
|
42
|
+
2. appends the resulting messages to the child agent's current chat;
|
|
43
|
+
3. extracts `Chat.find_role(res, :meta)`;
|
|
44
|
+
4. serializes those records as `agent_meta` in the parent
|
|
45
|
+
`function_call_output` JSON.
|
|
46
|
+
|
|
47
|
+
The receipt shape is therefore intentionally narrow and stable:
|
|
48
|
+
|
|
49
|
+
function_call_output: {
|
|
50
|
+
"name": "ask",
|
|
51
|
+
"content": "child answer",
|
|
52
|
+
"id": "call_...",
|
|
53
|
+
"agent_meta": [
|
|
54
|
+
{"role": "meta", "content": "pt=... tt=... inference_id=..."},
|
|
55
|
+
{"role": "meta", "content": "job=Workflow/ask/..."}
|
|
56
|
+
],
|
|
57
|
+
"start_timestamp": "...",
|
|
58
|
+
"timestamp": "..."
|
|
59
|
+
}
|
|
60
|
+
|
|
61
|
+
The current provenance walker only reads ordinary persisted Chat `meta:`
|
|
62
|
+
messages through `chat.jobs`. It does not parse `function_call_output` JSON,
|
|
63
|
+
so it misses `job=` records stored in `agent_meta`.
|
|
64
|
+
|
|
65
|
+
Likewise, `Chat.token_totals` and the current `llm prov` aggregate direct
|
|
66
|
+
messages found in discovered chat files. They do not see direct token events
|
|
67
|
+
inside a receipt. ChatAnalyst can recover this only through specialized agent
|
|
68
|
+
instructions, which is exactly the framework responsibility that should move
|
|
69
|
+
into Scout-AI primitives.
|
|
70
|
+
|
|
71
|
+
## Semantics and invariants
|
|
72
|
+
|
|
73
|
+
### The three relevant persisted facts
|
|
74
|
+
|
|
75
|
+
| Fact | Authoritative use |
|
|
76
|
+
|---|---|
|
|
77
|
+
| Parent function call/output | The parent asked a tool or agent and received this result. |
|
|
78
|
+
| `agent_meta` receipt | Portable evidence of direct child inference events and child producer jobs. It is especially important when no child log is discoverable. |
|
|
79
|
+
| Child chat / Step logs | Full child execution history: messages, tool calls, direct inference traces, dependencies, files, statuses, and exceptions. |
|
|
80
|
+
|
|
81
|
+
The receipt is not a replacement for a saved child log. It is a compact
|
|
82
|
+
projection of some of the same evidence.
|
|
83
|
+
|
|
84
|
+
### What is authoritative for token counting
|
|
85
|
+
|
|
86
|
+
The authoritative counted unit is a **direct inference event**, identified by
|
|
87
|
+
`inference_id`.
|
|
88
|
+
|
|
89
|
+
- A direct meta record in a saved child `agent.chat` and the same record in an
|
|
90
|
+
`agent_meta` receipt represent one event if their `inference_id` matches.
|
|
91
|
+
- Both source locations must remain visible in an audit report.
|
|
92
|
+
- The event's token fields are counted once.
|
|
93
|
+
- A `job=` meta record is not a token event. It is a producer reference and
|
|
94
|
+
must cause traversal into the referenced Step.
|
|
95
|
+
- `*_c` and `*_s` remain checkpoints and must never be summed.
|
|
96
|
+
|
|
97
|
+
### Why `inference_id` is specifically justified here
|
|
98
|
+
|
|
99
|
+
The receipt/log duplication is a real and common ambiguity, unlike a merely
|
|
100
|
+
hypothetical retry:
|
|
101
|
+
|
|
102
|
+
- parent output contains the Worker’s two direct metas;
|
|
103
|
+
- Worker `agent.chat`, when saved, contains the same two metas;
|
|
104
|
+
- both are valid persisted evidence;
|
|
105
|
+
- only the event identity says they are the same paid requests.
|
|
106
|
+
|
|
107
|
+
Without `inference_id`, no exact general rule can distinguish a receipt copy
|
|
108
|
+
from an independent request with matching metadata. The field is therefore a
|
|
109
|
+
join key between two persisted representations of one actual inference.
|
|
110
|
+
|
|
111
|
+
### Never inject receipt records into the parent Chat
|
|
112
|
+
|
|
113
|
+
Do not append `agent_meta` records to the parent Chat as ordinary `role: meta`
|
|
114
|
+
messages.
|
|
115
|
+
|
|
116
|
+
That would corrupt parent conversational bookkeeping:
|
|
117
|
+
|
|
118
|
+
- `Chat.meta` selects checkpoints used by later parent requests;
|
|
119
|
+
- child `*_c` values belong to the child conversation;
|
|
120
|
+
- parent and child session totals must remain separate;
|
|
121
|
+
- a delegated receipt is evidence attached to a tool output, not a new parent
|
|
122
|
+
response segment.
|
|
123
|
+
|
|
124
|
+
Extraction must be read-only and happen only in provenance analysis.
|
|
125
|
+
|
|
126
|
+
## Proposed Scout-AI primitives
|
|
127
|
+
|
|
128
|
+
The implementation should be small, procedural, and return ordinary Arrays
|
|
129
|
+
and Hashes.
|
|
130
|
+
|
|
131
|
+
### 1. Parse agent-meta evidence from tool outputs
|
|
132
|
+
|
|
133
|
+
Add a Chat helper near `lib/scout/llm/chat/tool_calls.rb`, reusing
|
|
134
|
+
`Chat.tool_calls` to avoid reparsing output pairing logic.
|
|
135
|
+
|
|
136
|
+
Suggested API:
|
|
137
|
+
|
|
138
|
+
Chat.agent_meta_evidence(chat, source: nil)
|
|
139
|
+
|
|
140
|
+
It returns one plain Hash per valid receipt meta record, for example:
|
|
141
|
+
|
|
142
|
+
{
|
|
143
|
+
origin: :agent_meta,
|
|
144
|
+
meta: <IndiferentHash parsed by Chat.parse_meta>,
|
|
145
|
+
source: "/path/to/parent.chat",
|
|
146
|
+
output_address: ["/path/to/parent.chat", 17],
|
|
147
|
+
evidence_address: ["/path/to/parent.chat", 17, :agent_meta, 0],
|
|
148
|
+
call_id: "call_123",
|
|
149
|
+
tool_name: "ask",
|
|
150
|
+
agent_meta_index: 0,
|
|
151
|
+
raw_message: {role: "meta", content: "..."}
|
|
152
|
+
}
|
|
153
|
+
|
|
154
|
+
`source` is optional, matching existing source-aware Chat APIs. If omitted,
|
|
155
|
+
the address can use only message indexes.
|
|
156
|
+
|
|
157
|
+
Extraction rules:
|
|
158
|
+
|
|
159
|
+
- inspect parsed `function_call_output` records, not raw text with regular
|
|
160
|
+
expressions;
|
|
161
|
+
- accept an `agent_meta` Array only;
|
|
162
|
+
- accept entries only when they are Hashes with `role == "meta"` and a String
|
|
163
|
+
`content`;
|
|
164
|
+
- parse content through `Chat.parse_meta`;
|
|
165
|
+
- retain malformed receipt entries as warnings when the caller asks for
|
|
166
|
+
warnings, but never silently reinterpret arbitrary output JSON as
|
|
167
|
+
provenance;
|
|
168
|
+
- do not limit the logic to tool name `ask`: the envelope is safe to support
|
|
169
|
+
generically, although reports may label agent-oriented tools specially.
|
|
170
|
+
|
|
171
|
+
### 2. Expose all meta evidence without altering ordinary Chat semantics
|
|
172
|
+
|
|
173
|
+
Add:
|
|
174
|
+
|
|
175
|
+
Chat.meta_evidence(chat, source: nil)
|
|
176
|
+
|
|
177
|
+
This returns normal persisted `role: meta` records plus
|
|
178
|
+
`Chat.agent_meta_evidence` records, with a shared shape that includes
|
|
179
|
+
`origin`.
|
|
180
|
+
|
|
181
|
+
Suggested origins:
|
|
182
|
+
|
|
183
|
+
- `:chat_meta` for a normal Chat message;
|
|
184
|
+
- `:agent_meta` for a receipt embedded in a function output.
|
|
185
|
+
|
|
186
|
+
Do not change `Chat#meta`, `chat.role_messages(:meta)`, or the current
|
|
187
|
+
meaning of `Chat.token_totals([chat])`. Those APIs describe the local Chat and
|
|
188
|
+
must not unexpectedly absorb delegated receipts.
|
|
189
|
+
|
|
190
|
+
### 3. Expose embedded producer-job references
|
|
191
|
+
|
|
192
|
+
Add:
|
|
193
|
+
|
|
194
|
+
Chat.agent_meta_job_references(chat, source: nil)
|
|
195
|
+
|
|
196
|
+
This filters `agent_meta_evidence` to records with `meta[:job]`. It returns the
|
|
197
|
+
reference and the evidence location.
|
|
198
|
+
|
|
199
|
+
The existing `chat.jobs` can remain a local-chat API. Do not silently redefine
|
|
200
|
+
it to include nested receipt jobs; callers need to distinguish direct visible
|
|
201
|
+
projection from delegated receipt provenance.
|
|
202
|
+
|
|
203
|
+
### 4. Add a provenance-aware token-event collector
|
|
204
|
+
|
|
205
|
+
Add a root-oriented API, for example:
|
|
206
|
+
|
|
207
|
+
Chat.provenance_token_events(root, **traversal_options)
|
|
208
|
+
Chat.provenance_token_totals(root, **traversal_options)
|
|
209
|
+
|
|
210
|
+
This API should:
|
|
211
|
+
|
|
212
|
+
1. invoke `Chat.traverse_provenance` to discover persisted chats and jobs;
|
|
213
|
+
2. load every discovered chat once;
|
|
214
|
+
3. collect ordinary direct events from source-aware chat tracing;
|
|
215
|
+
4. collect direct receipt events from `agent_meta_evidence`;
|
|
216
|
+
5. discard projection records with `meta[:job]` from token addition;
|
|
217
|
+
6. group direct events by `inference_id`;
|
|
218
|
+
7. return one event record with all evidence locations;
|
|
219
|
+
8. sum canonical direct fields only: `pt`, `ct`, `tt`, `cct`, `cwt`, and `rt`.
|
|
220
|
+
|
|
221
|
+
A returned event should include enough audit information to explain a total:
|
|
222
|
+
|
|
223
|
+
{
|
|
224
|
+
inference_id: "...",
|
|
225
|
+
meta: <canonical parsed direct meta>,
|
|
226
|
+
tokens: {pt: ..., ct: ..., tt: ...},
|
|
227
|
+
evidence: [
|
|
228
|
+
{origin: :agent_meta, evidence_address: [...]},
|
|
229
|
+
{origin: :chat_meta, meta_address: [...]}
|
|
230
|
+
],
|
|
231
|
+
deduplication: :inference_id
|
|
232
|
+
}
|
|
233
|
+
|
|
234
|
+
### Legacy records without `inference_id`
|
|
235
|
+
|
|
236
|
+
Do not claim exact receipt/log deduplication for old data.
|
|
237
|
+
|
|
238
|
+
For ordinary persisted chat meta, preserve the existing lineage-based fallback
|
|
239
|
+
already used by `Chat.trace_chats`.
|
|
240
|
+
|
|
241
|
+
For an embedded receipt meta without an inference ID, there is no full child
|
|
242
|
+
conversation in which to compute its lineage. Use a receipt-address-based
|
|
243
|
+
fallback and mark it clearly, for example:
|
|
244
|
+
|
|
245
|
+
deduplication: :receipt_unresolved
|
|
246
|
+
|
|
247
|
+
If a provider response ID is present and Scout-AI considers it stable enough,
|
|
248
|
+
it may be used as an explicit secondary identity. Do not silently use
|
|
249
|
+
heuristic equality of token counts, timestamps, or reasoning text as proof of
|
|
250
|
+
identity.
|
|
251
|
+
|
|
252
|
+
Reports should warn that legacy receipt/log records can be overcounted when no
|
|
253
|
+
shared inference or provider identity exists.
|
|
254
|
+
|
|
255
|
+
### Identity conflicts
|
|
256
|
+
|
|
257
|
+
If two records share an `inference_id` but disagree on direct token fields,
|
|
258
|
+
provider response ID, or other immutable event facts:
|
|
259
|
+
|
|
260
|
+
- do not sum both;
|
|
261
|
+
- retain both evidence records;
|
|
262
|
+
- emit a provenance conflict warning;
|
|
263
|
+
- make the conflict visible in `prov` and ChatAnalyst.
|
|
264
|
+
|
|
265
|
+
This detects a broken producer rather than hiding it behind deduplication.
|
|
266
|
+
|
|
267
|
+
## Traversal changes
|
|
268
|
+
|
|
269
|
+
### Add `:agent_job` as a relation
|
|
270
|
+
|
|
271
|
+
Extend `Chat::PROVENANCE_RELATIONS` with `:agent_job`.
|
|
272
|
+
|
|
273
|
+
When visiting a Chat node, the walker should enqueue both:
|
|
274
|
+
|
|
275
|
+
- ordinary direct `chat.jobs` as `:job` edges;
|
|
276
|
+
- receipt `agent_meta_job_references` as `:agent_job` edges.
|
|
277
|
+
|
|
278
|
+
The parent remains the enclosing Chat node and the child remains a native
|
|
279
|
+
Step. There is no receipt node.
|
|
280
|
+
|
|
281
|
+
This relation is useful because it preserves why the job was discovered:
|
|
282
|
+
|
|
283
|
+
| Relation | Meaning |
|
|
284
|
+
|---|---|
|
|
285
|
+
| `job` | A visible response segment in this chat was projected from the job. |
|
|
286
|
+
| `agent_job` | A delegated agent receipt inside a tool output says its response was projected from the job. |
|
|
287
|
+
|
|
288
|
+
The referenced Step then follows normal `dependency`, `log`, and `result`
|
|
289
|
+
relations. Existing cycle and shared-node handling applies unchanged.
|
|
290
|
+
|
|
291
|
+
### Error handling
|
|
292
|
+
|
|
293
|
+
Malformed receipt JSON or unreadable receipt job references must be reported
|
|
294
|
+
through the existing `on_error`/warning path with:
|
|
295
|
+
|
|
296
|
+
- enclosing chat path;
|
|
297
|
+
- tool output address;
|
|
298
|
+
- call ID and tool name when available;
|
|
299
|
+
- relation `:agent_job`;
|
|
300
|
+
- original reference.
|
|
301
|
+
|
|
302
|
+
A malformed receipt should not hide other normal provenance from the same
|
|
303
|
+
chat.
|
|
304
|
+
|
|
305
|
+
### Do not follow ordinary imports
|
|
306
|
+
|
|
307
|
+
The current traversal intentionally treats imports as compilation composition,
|
|
308
|
+
not persisted provenance edges. Preserve that policy. The agent-meta work does
|
|
309
|
+
not require reintroducing import traversal.
|
|
310
|
+
|
|
311
|
+
## `scout-ai llm prov` changes
|
|
312
|
+
|
|
313
|
+
The command already separates graph discovery from token computation. Update
|
|
314
|
+
only those layers.
|
|
315
|
+
|
|
316
|
+
### Graph discovery and rendering
|
|
317
|
+
|
|
318
|
+
1. Call the enhanced traversal.
|
|
319
|
+
2. Include `:agent_job` in adjacency sorting, after direct `:job` and before
|
|
320
|
+
ordinary dependency/log display as appropriate.
|
|
321
|
+
3. In default tree mode, show a job discovered through a receipt with an
|
|
322
|
+
explicit label, for example:
|
|
323
|
+
|
|
324
|
+
chat Manager.chat
|
|
325
|
+
delegated-job Worker/ask abc12345
|
|
326
|
+
|
|
327
|
+
The job remains a normal job node; `delegated-job` is the edge label.
|
|
328
|
+
|
|
329
|
+
4. In flow and DOT modes, reverse `:agent_job` in the same way as `:job`,
|
|
330
|
+
because the job produces content used by the parent chat. Use a distinct
|
|
331
|
+
visual style or label such as `delegated_result` so it is not confused with
|
|
332
|
+
a direct projected job result.
|
|
333
|
+
|
|
334
|
+
5. Do not make a Graphviz node for every receipt or every direct inference.
|
|
335
|
+
The graph should remain a Chat/Step structural graph.
|
|
336
|
+
|
|
337
|
+
### Token display
|
|
338
|
+
|
|
339
|
+
Replace direct uses of `Chat.token_totals(chats)` in `prov` aggregate logic
|
|
340
|
+
with `Chat.provenance_token_totals` or the equivalent event collector.
|
|
341
|
+
|
|
342
|
+
The default total must include receipt-only child usage. It must not add it
|
|
343
|
+
again when the same child log is reachable.
|
|
344
|
+
|
|
345
|
+
The current `--component` mode should become explicit about source scope. The
|
|
346
|
+
recommended output distinctions are:
|
|
347
|
+
|
|
348
|
+
- `local`: direct meta events physically stored in the chat/log;
|
|
349
|
+
- `receipt`: child direct events found only or also in `agent_meta` receipts;
|
|
350
|
+
- `aggregate`: deduplicated union of all reachable event identities.
|
|
351
|
+
|
|
352
|
+
For a compact tree, do not print a separate line for every receipt by default.
|
|
353
|
+
Instead, append a concise annotation to the parent chat or tool call summary,
|
|
354
|
+
for example:
|
|
355
|
+
|
|
356
|
+
delegated receipt: 2 events, total=17.0k, Worker/test_sum
|
|
357
|
+
|
|
358
|
+
When `--component` is enabled, print receipt components individually with call
|
|
359
|
+
ID and source address.
|
|
360
|
+
|
|
361
|
+
### New CLI option
|
|
362
|
+
|
|
363
|
+
Add a focused diagnostic option rather than overloading flow output. Suggested
|
|
364
|
+
name:
|
|
365
|
+
|
|
366
|
+
--evidence
|
|
367
|
+
|
|
368
|
+
It prints the deduplicated direct inference event table:
|
|
369
|
+
|
|
370
|
+
| Inference ID | Tokens | Evidence | Status |
|
|
371
|
+
|---|---:|---|---|
|
|
372
|
+
| `eed7eb...` | 8463 | parent ask output; Worker agent.chat | counted once |
|
|
373
|
+
| `a477d8...` | 8550 | parent ask output; Worker agent.chat | counted once |
|
|
374
|
+
|
|
375
|
+
It should also display:
|
|
376
|
+
|
|
377
|
+
- receipt-only events;
|
|
378
|
+
- legacy unresolved events;
|
|
379
|
+
- inference-ID conflicts;
|
|
380
|
+
- job projection references, but with no direct tokens.
|
|
381
|
+
|
|
382
|
+
This is the right command for diagnosing why a total contains child work. The
|
|
383
|
+
normal tree and flow should remain compact.
|
|
384
|
+
|
|
385
|
+
## ChatAnalyst changes
|
|
386
|
+
|
|
387
|
+
ChatAnalyst should consume the Scout-AI primitives and delete any bespoke
|
|
388
|
+
receipt-parsing logic it currently has. It should not need special agent
|
|
389
|
+
instructions to discover `agent_meta`.
|
|
390
|
+
|
|
391
|
+
### `chat_overview`
|
|
392
|
+
|
|
393
|
+
Add:
|
|
394
|
+
|
|
395
|
+
- `agent_job` structural edges;
|
|
396
|
+
- number of receipt records per chat;
|
|
397
|
+
- number of receipt-only direct events;
|
|
398
|
+
- provenance warnings for malformed receipts or conflicts.
|
|
399
|
+
|
|
400
|
+
### `chat_tool_calls`
|
|
401
|
+
|
|
402
|
+
For each function call output, add a compact receipt summary:
|
|
403
|
+
|
|
404
|
+
agent_meta: {
|
|
405
|
+
direct_events: 2,
|
|
406
|
+
job_references: ["Worker/ask/..."],
|
|
407
|
+
token_total: {tt: 17013},
|
|
408
|
+
event_ids: ["eed7...", "a477..."]
|
|
409
|
+
}
|
|
410
|
+
|
|
411
|
+
Keep the raw answer content and tool success status separate from token
|
|
412
|
+
attribution.
|
|
413
|
+
|
|
414
|
+
### `chat_tokens`
|
|
415
|
+
|
|
416
|
+
This task should switch from plain `Chat.token_totals` to the new provenance
|
|
417
|
+
event collector. Return:
|
|
418
|
+
|
|
419
|
+
- aggregate deduplicated totals;
|
|
420
|
+
- `events` or a compact per-event index;
|
|
421
|
+
- per-chat local totals;
|
|
422
|
+
- receipt-derived totals;
|
|
423
|
+
- receipt-only totals;
|
|
424
|
+
- duplicate evidence count;
|
|
425
|
+
- unresolved legacy receipt count;
|
|
426
|
+
- identity conflict warnings.
|
|
427
|
+
|
|
428
|
+
Do not describe a receipt contribution as a second paid inference when its ID
|
|
429
|
+
also appears in a saved child log.
|
|
430
|
+
|
|
431
|
+
### `chat_agents`
|
|
432
|
+
|
|
433
|
+
An agent interaction should report:
|
|
434
|
+
|
|
435
|
+
- call ID;
|
|
436
|
+
- target agent and conversation when present in arguments;
|
|
437
|
+
- receipt event IDs and totals;
|
|
438
|
+
- linked `agent_job` Steps, if any;
|
|
439
|
+
- whether child evidence was receipt-only, log-only, or both;
|
|
440
|
+
- whether the log association is structural, inferred by naming convention, or
|
|
441
|
+
absent.
|
|
442
|
+
|
|
443
|
+
### `message_index` and `message_content`
|
|
444
|
+
|
|
445
|
+
Do not pretend embedded receipt metas are ordinary child-chat messages.
|
|
446
|
+
|
|
447
|
+
Either:
|
|
448
|
+
|
|
449
|
+
1. add a dedicated `meta_evidence` task; or
|
|
450
|
+
2. allow `message_index --role meta` to include an `origin` field and a nested
|
|
451
|
+
receipt address.
|
|
452
|
+
|
|
453
|
+
A dedicated task is clearer. It can return normal and embedded meta evidence
|
|
454
|
+
without asserting that an embedded record has a full child message history.
|
|
455
|
+
|
|
456
|
+
Suggested address format:
|
|
457
|
+
|
|
458
|
+
[parent_chat_path, function_output_index, :agent_meta, meta_index]
|
|
459
|
+
|
|
460
|
+
`message_content` may resolve this address by re-parsing the parent output;
|
|
461
|
+
there is no separate physical child message at that address.
|
|
462
|
+
|
|
463
|
+
### `chat_report`
|
|
464
|
+
|
|
465
|
+
Include concise highlights only:
|
|
466
|
+
|
|
467
|
+
- total deduplicated direct tokens;
|
|
468
|
+
- local versus receipt-only contribution;
|
|
469
|
+
- number of event IDs with multiple evidence locations;
|
|
470
|
+
- count of `agent_job` edges;
|
|
471
|
+
- conflicts and unresolved legacy receipts.
|
|
472
|
+
|
|
473
|
+
## Test plan
|
|
474
|
+
|
|
475
|
+
All fixtures must be offline and use persisted chat text plus temporary job
|
|
476
|
+
layouts. Do not call model providers.
|
|
477
|
+
|
|
478
|
+
### A. Receipt-only socialized delegation
|
|
479
|
+
|
|
480
|
+
Fixture:
|
|
481
|
+
|
|
482
|
+
- parent chat has one `ask` function call/output;
|
|
483
|
+
- output includes two direct `agent_meta` records with distinct inference IDs;
|
|
484
|
+
- no saved Worker chat or Worker job exists.
|
|
485
|
+
|
|
486
|
+
Assertions:
|
|
487
|
+
|
|
488
|
+
- traversal still has only the parent Chat node;
|
|
489
|
+
- provenance token events include both Worker events;
|
|
490
|
+
- aggregate total includes parent plus Worker tokens;
|
|
491
|
+
- `prov --evidence` identifies both events as receipt-only;
|
|
492
|
+
- ChatAnalyst `chat_agents` reports the receipt and its total.
|
|
493
|
+
|
|
494
|
+
### B. Receipt plus saved Worker log
|
|
495
|
+
|
|
496
|
+
Fixture extends A with a discoverable Worker job/log containing the same two
|
|
497
|
+
metadata records and inference IDs.
|
|
498
|
+
|
|
499
|
+
Assertions:
|
|
500
|
+
|
|
501
|
+
- traversal discovers the Worker Step and log through normal structure;
|
|
502
|
+
- token total is identical to fixture A plus any additional known worker log
|
|
503
|
+
events, not doubled;
|
|
504
|
+
- event records retain both receipt and log source locations;
|
|
505
|
+
- `prov` default total and ChatAnalyst aggregate total agree;
|
|
506
|
+
- component output labels receipt evidence rather than charging it twice.
|
|
507
|
+
|
|
508
|
+
### C. Receipt with `job=` projection
|
|
509
|
+
|
|
510
|
+
Fixture:
|
|
511
|
+
|
|
512
|
+
- parent output contains `agent_meta` with `job=Worker/ask/...`;
|
|
513
|
+
- Worker Step has a normal agent log with direct token metadata.
|
|
514
|
+
|
|
515
|
+
Assertions:
|
|
516
|
+
|
|
517
|
+
- traversal emits an `:agent_job` edge;
|
|
518
|
+
- Worker dependencies and logs are recursively reached;
|
|
519
|
+
- the receipt job meta itself contributes zero direct tokens;
|
|
520
|
+
- actual Worker direct log tokens are counted once;
|
|
521
|
+
- tree, flow, and DOT label the delegated producer relationship.
|
|
522
|
+
|
|
523
|
+
### D. Nested receipt chain
|
|
524
|
+
|
|
525
|
+
Fixture:
|
|
526
|
+
|
|
527
|
+
- Manager receipt points to Worker;
|
|
528
|
+
- Worker log contains a second socialized receipt for Critic;
|
|
529
|
+
- Critic has either direct receipt-only events or a `chat_task` producer job.
|
|
530
|
+
|
|
531
|
+
Assertions:
|
|
532
|
+
|
|
533
|
+
- recursion terminates safely;
|
|
534
|
+
- all direct event IDs appear once in aggregate accounting;
|
|
535
|
+
- graph edges preserve both delegation paths;
|
|
536
|
+
- no Session-like recursive state object is needed.
|
|
537
|
+
|
|
538
|
+
### E. Malformed and incomplete receipts
|
|
539
|
+
|
|
540
|
+
Fixtures:
|
|
541
|
+
|
|
542
|
+
- invalid outer function-output JSON;
|
|
543
|
+
- `agent_meta` is not an Array;
|
|
544
|
+
- entry lacks role/content;
|
|
545
|
+
- malformed meta content;
|
|
546
|
+
- unresolved `job=` path.
|
|
547
|
+
|
|
548
|
+
Assertions:
|
|
549
|
+
|
|
550
|
+
- normal chat/job provenance remains available;
|
|
551
|
+
- warnings include source address and call ID when available;
|
|
552
|
+
- no malformed record is accidentally counted as tokens;
|
|
553
|
+
- strict and warning callback modes behave as documented.
|
|
554
|
+
|
|
555
|
+
### F. Identity conflict
|
|
556
|
+
|
|
557
|
+
Fixture contains two evidence locations with the same `inference_id` but
|
|
558
|
+
different `tt` or provider response ID.
|
|
559
|
+
|
|
560
|
+
Assertions:
|
|
561
|
+
|
|
562
|
+
- total does not silently add both;
|
|
563
|
+
- event records retain both facts;
|
|
564
|
+
- warning is emitted in `prov --evidence` and ChatAnalyst;
|
|
565
|
+
- automated tests make the chosen conflict policy explicit.
|
|
566
|
+
|
|
567
|
+
### G. Legacy records
|
|
568
|
+
|
|
569
|
+
Fixture has ordinary and receipt metadata without inference IDs.
|
|
570
|
+
|
|
571
|
+
Assertions:
|
|
572
|
+
|
|
573
|
+
- normal Chat records retain current lineage fallback behavior;
|
|
574
|
+
- receipt records are marked unresolved unless an explicit secondary identity
|
|
575
|
+
is available;
|
|
576
|
+
- reports do not claim exact receipt/log deduplication;
|
|
577
|
+
- no regression occurs for old chats without `agent_meta`.
|
|
578
|
+
|
|
579
|
+
### H. Output truncation
|
|
580
|
+
|
|
581
|
+
Fixture uses an oversized agent answer whose parent output content is replaced
|
|
582
|
+
with the standard truncation exception but whose `agent_meta` remains present.
|
|
583
|
+
|
|
584
|
+
Assertions:
|
|
585
|
+
|
|
586
|
+
- receipt token events remain discoverable;
|
|
587
|
+
- truncation state is reported separately from model cost;
|
|
588
|
+
- no child tokens are lost merely because parent context omitted the full text.
|
|
589
|
+
|
|
590
|
+
## Acceptance criteria
|
|
591
|
+
|
|
592
|
+
The implementation is complete when:
|
|
593
|
+
|
|
594
|
+
1. a root chat with only an `ask` receipt reports delegated direct token usage;
|
|
595
|
+
2. a receipt plus saved child log counts each `inference_id` once;
|
|
596
|
+
3. receipt `job=` entries create discoverable normal Steps through
|
|
597
|
+
`:agent_job` traversal edges;
|
|
598
|
+
4. parent Chat checkpoint semantics remain unchanged;
|
|
599
|
+
5. existing `Chat.token_totals([chat])` remains local-Chat compatible;
|
|
600
|
+
6. `llm prov`, ChatAnalyst, and the new core collector return the same
|
|
601
|
+
deduplicated aggregate total for shared fixtures;
|
|
602
|
+
7. every receipt-derived amount is traceable to a parent output address and
|
|
603
|
+
call ID;
|
|
604
|
+
8. malformed receipts and identity conflicts are warnings, not silent
|
|
605
|
+
undercounting or double counting;
|
|
606
|
+
9. no wrapper class is introduced merely to hold traversal or accounting state;
|
|
607
|
+
10. all new tests run offline.
|
|
608
|
+
|
|
609
|
+
## Recommended delivery order
|
|
610
|
+
|
|
611
|
+
1. Add receipt/meta evidence extraction and unit tests.
|
|
612
|
+
2. Add embedded job-reference extraction and `:agent_job` traversal tests.
|
|
613
|
+
3. Add root-oriented event collection and exact inference-ID deduplication.
|
|
614
|
+
4. Update `Chat.tokens` and any provenance-specific aggregate callers to use
|
|
615
|
+
the new collector, while retaining local `Chat.token_totals` semantics.
|
|
616
|
+
5. Update `scout-ai llm prov`, including `--evidence` and component labels.
|
|
617
|
+
6. Update ChatAnalyst to consume the core APIs and remove bespoke receipt
|
|
618
|
+
recovery logic.
|
|
619
|
+
7. Add cross-consumer fixtures asserting equal totals and equivalent job
|
|
620
|
+
discovery.
|
|
621
|
+
8. Update maintained developer documentation and preserve this plan as the
|
|
622
|
+
deeper design record.
|