scout-ai 2.0.0 → 2.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/.vimproject +5 -0
- data/VERSION +1 -1
- data/doc/developer/ChatLifecycle.md +1 -1
- data/doc/developer/DelegationInternals.md +1 -1
- data/doc/developer/Provenance.md +15 -5
- data/lib/scout/llm/agent/attach.rb +68 -0
- data/lib/scout/llm/agent/workflow.rb +2 -1
- data/lib/scout/llm/agent.rb +9 -1
- data/lib/scout/llm/ask.rb +11 -12
- data/lib/scout/llm/backends/default.rb +9 -3
- data/lib/scout/llm/chat/process/tools.rb +2 -5
- data/lib/scout/llm/chat/provenance.rb +61 -2
- data/research/provenance-continuation-accounting.md +255 -0
- data/scout-ai.gemspec +4 -2
- data/scout_commands/llm/prov +121 -21
- data/test/scout/llm/chat/agent_meta_fixtures.rb +91 -5
- data/test/scout/llm/chat/test_agent_meta_provenance.rb +133 -0
- data/test/scout/llm/chat/test_agent_meta_tokens.rb +54 -0
- data/test/scout/llm/chat/test_prov_cli.rb +406 -6
- metadata +3 -1
|
@@ -0,0 +1,255 @@
|
|
|
1
|
+
# Provenance token accounting for continued conversations (design, not implemented)
|
|
2
|
+
|
|
3
|
+
Step 3, Part 2 of the provenance-gap plan. Part 1 (provenance follows `job=`
|
|
4
|
+
receipt edges) is done and validated; this document designs token accounting
|
|
5
|
+
for continuation chains (Cortex `continue` jobs and any `chat_task` that
|
|
6
|
+
follows a previously projected chat). No library code was changed for this
|
|
7
|
+
design; probes live in `tmp/prov-cont/measure*.rb` and
|
|
8
|
+
`tmp/cortex-continuation-accounting-probes.md`.
|
|
9
|
+
|
|
10
|
+
All numbers below are measured on the real example chain unless marked
|
|
11
|
+
otherwise. Jobs live under `/home/mvazque2/.scout/var/jobs/Cortex/continue/`.
|
|
12
|
+
|
|
13
|
+
## Evidence
|
|
14
|
+
|
|
15
|
+
### The example chain
|
|
16
|
+
|
|
17
|
+
Three sequential continuations of one conversation, `fin_fixes_recon`
|
|
18
|
+
(workspace copy not found in this environment, see Open questions):
|
|
19
|
+
|
|
20
|
+
| job | result chat (`Default_<k>.chat`) | inference metas | pt sum |
|
|
21
|
+
|-----|----------------------------------|-----------------|--------|
|
|
22
|
+
| `c0b45189` | 445 messages, 1 `job=` marker, 147 metas | 147 | 6,297,210 |
|
|
23
|
+
| `68b5b7fa` | 711 messages, 1 `job=` marker, 237 metas | 237 | 10,707,400 |
|
|
24
|
+
| `80070846` | 60 messages, 1 `job=` marker, 20 metas | 20 | 1,124,684 |
|
|
25
|
+
|
|
26
|
+
Each result chat starts with its own `job=` meta marker followed by exactly
|
|
27
|
+
that job's own new inferences (this is `Chat.project(self.short_path, result)`
|
|
28
|
+
with `result = agent.current_chat - agent.start_chat`,
|
|
29
|
+
lib/scout/llm/agent/workflow.rb:127-139, projection at
|
|
30
|
+
lib/scout/llm/chat/process/meta.rb:410).
|
|
31
|
+
|
|
32
|
+
### What the agent log actually contains (replay measured)
|
|
33
|
+
|
|
34
|
+
`Default_80070846...chat.files/agent.chat` (the third, latest job):
|
|
35
|
+
|
|
36
|
+
- 404 inference metas, 404 *distinct* `inference_id`s, 0 duplicated ids
|
|
37
|
+
(`tmp/prov-cont/measure41.rb`).
|
|
38
|
+
- Organized in projected segments delimited by `job=` meta markers:
|
|
39
|
+
- segment `job=...c0b45189...chat`: 147 metas, pt 6,297,210,
|
|
40
|
+
timestamps 2026-09-02T20:32:48Z .. 21:53:59Z;
|
|
41
|
+
- segment `job=...68b5b7fa...chat`: 257 metas, pt 11,832,084,
|
|
42
|
+
timestamps 21:54:59Z .. 22:48:55Z
|
|
43
|
+
(`measure43.rb`, `measure47.rb`).
|
|
44
|
+
- The last segment is MIXED: 257 = 237 inferences authored by job
|
|
45
|
+
`68b5b7fa` + 20 authored by `80070846` itself. There is no marker between
|
|
46
|
+
the second job's projected result and the third job's own new turns; the
|
|
47
|
+
third job appended its turns after `start_chat.follow chat`
|
|
48
|
+
(lib/scout/llm/agent/workflow.rb:85-87) of the projected second result.
|
|
49
|
+
- `Default_68b5b7fa...files/agent.chat`: 384 metas, all inside the
|
|
50
|
+
`job=c0b45189` segment (147 replayed + 237 own new, unmarked).
|
|
51
|
+
- `Default_c0b45189...files/agent.chat`: 147 unmarked metas, nothing replayed.
|
|
52
|
+
|
|
53
|
+
So: a continuation log carries the entire grown conversation; every replayed
|
|
54
|
+
inference retains its original per-call token fields (`pt`, `ct`, `tt`,
|
|
55
|
+
`cct`, `cwt`, `rt`, plus cumulative `*_c` and session `*_s` keys, all 404
|
|
56
|
+
rows), original `inference_id`, `provider_response_id` and `timestamp`.
|
|
57
|
+
|
|
58
|
+
### Inference identity is stable across copies
|
|
59
|
+
|
|
60
|
+
Shared `inference_id`s between the second and third logs: 384; token-field
|
|
61
|
+
(`pt/ct/tt/cct`) mismatches among them: 0 (`measure39.rb`). The third log's
|
|
62
|
+
inherited rows sum to pt 17,004,610 / cct 16,288,896, byte-identical totals
|
|
63
|
+
to the second job's own log content. Identity, not position, is the reliable
|
|
64
|
+
key.
|
|
65
|
+
|
|
66
|
+
### Traversal already reaches the whole chain
|
|
67
|
+
|
|
68
|
+
`Chat.provenance_chat_files(third result chat)` returns 6 files: the three
|
|
69
|
+
result chats plus the three `.files/agent.chat` logs (`measure44.rb`), via
|
|
70
|
+
the `job=` marker edges (Part 1). Per-call meta keys observed in the third
|
|
71
|
+
log, all 404 rows: `inference_id`, `timestamp`, `provider_response_id`,
|
|
72
|
+
`pt/ct/tt/cct/cwt/rt`, `pt_c...rt_c`, `pt_s...rt_s`, plus `reas` on 12 rows.
|
|
73
|
+
|
|
74
|
+
## Current behavior (measured)
|
|
75
|
+
|
|
76
|
+
`Chat.provenance_token_totals(root)` (lib/scout/llm/chat/provenance.rb:729)
|
|
77
|
+
collects per-inference events (`provenance_token_events`, line 499),
|
|
78
|
+
deduplicates them by identity (`[:inference_id, id]`, else
|
|
79
|
+
`provider_response_id`, else lineage, else receipt address; lines 577-608),
|
|
80
|
+
then sums the canonical event's tokens (lines 706-708).
|
|
81
|
+
|
|
82
|
+
Measured per-root totals (`measure44.rb`):
|
|
83
|
+
|
|
84
|
+
| root | pt | interpretation |
|
|
85
|
+
|------|----|----------------|
|
|
86
|
+
| `c0b45189` | 6,297,210 | own only |
|
|
87
|
+
| `68b5b7fa` | 17,004,610 | cumulative: own 10,707,400 + first's 6,297,210 |
|
|
88
|
+
| `80070846` | 18,129,294 | cumulative: own 1,124,684 + 17,004,610 |
|
|
89
|
+
|
|
90
|
+
Verdicts:
|
|
91
|
+
|
|
92
|
+
1. Global deduplicated total does NOT double count. 18,129,294 =
|
|
93
|
+
6,297,210 + 10,707,400 + 1,124,684 exactly. Replayed copies of the same
|
|
94
|
+
inference in later logs collapse onto one event; the code's `next if
|
|
95
|
+
meta[:job]` filters (lines 534, 557) plus identity grouping do their job.
|
|
96
|
+
2. Per-job root totals ARE cumulative. A continuation job's root total
|
|
97
|
+
includes every ancestor inference, so summing per-root totals across a
|
|
98
|
+
chain double counts (6,297,210 would be counted three times), and a
|
|
99
|
+
naive per-job cost read overstates job `68b5b7fa` by 6.3M pt. Today
|
|
100
|
+
there is no per-job "owned" breakdown: scope symbols
|
|
101
|
+
(`:deduplicated_total`, `:chat_evidence`, `:receipt_evidence`,
|
|
102
|
+
`:receipt_only`, lines 737-750) partition evidence coverage, not authorship.
|
|
103
|
+
3. Re-send cost is present but unseparated. The third job's own 20 calls
|
|
104
|
+
cost pt 1,124,684 with cct 1,028,352 (91.5% of prompt tokens served from
|
|
105
|
+
cache; uncached prompt = pt - cct = 96,332). Per-call `pt` already
|
|
106
|
+
includes re-sent history for every call, which is the honest provider
|
|
107
|
+
view, but nothing at job level separates "what this job authored" from
|
|
108
|
+
"history this job re-paid for".
|
|
109
|
+
4. Degraded case (missing child `81e265cf`): root
|
|
110
|
+
`Planned/work/Default_de207a52...chat` totals pt 37,969,140 with 2
|
|
111
|
+
warnings, both `unresolved_job_reference` for
|
|
112
|
+
`Cortex/continue/Default_81e265cf...chat` (file absent on disk). The
|
|
113
|
+
parent's log carries that job's projected segment with 0 inference metas
|
|
114
|
+
(its result was an exception receipt), so nothing is silently counted or
|
|
115
|
+
lost; the missing job's own cost is simply unknown, reported only as a
|
|
116
|
+
warning.
|
|
117
|
+
|
|
118
|
+
## Attribution rules
|
|
119
|
+
|
|
120
|
+
Core rule: **an inference belongs to exactly one job - the job that issued
|
|
121
|
+
the model call - and its tokens are attributed to that job only; any other
|
|
122
|
+
copy of the same `inference_id` is evidence, never a second charge.**
|
|
123
|
+
|
|
124
|
+
1. Ownership set of a job J = the `inference_id`s of the inference metas in
|
|
125
|
+
J's persisted result chat (for `chat` result-type jobs, `job_result_chat_file`,
|
|
126
|
+
provenance.rb:83). Measured: this reproduces the exact deltas above
|
|
127
|
+
(147 / 237 / 20).
|
|
128
|
+
2. Fallback ownership (result chat absent, e.g. non-chat jobs or pruned
|
|
129
|
+
results): inference metas in J's own logs (`direct_job_chat_files`,
|
|
130
|
+
provenance.rb:44) whose `timestamp` >= J's start time. Measured: job
|
|
131
|
+
`68b5b7fa` issued 2026-09-02T21:54:53Z; the 237 rows with
|
|
132
|
+
`timestamp >= issued` equal its result-chat set exactly (set difference 0).
|
|
133
|
+
3. Rows in J's logs belonging to other jobs are labeled inherited, with
|
|
134
|
+
provenance `{owning job id -> count}` derived from the `job=` segment
|
|
135
|
+
markers around them; they never enter J's owned sums.
|
|
136
|
+
4. Global total = sum over distinct `inference_id`s (unchanged; already
|
|
137
|
+
correct today). Per-job owned totals must sum to the global total when
|
|
138
|
+
the chain is complete; the residual (global - sum of known owners) is
|
|
139
|
+
attributed to `unknown`, driven by missing jobs.
|
|
140
|
+
5. A missing job referenced by a receipt keeps its `unresolved_job_reference`
|
|
141
|
+
warning; its contribution is reported as `missing`, never as zero-owned
|
|
142
|
+
and never dropped from the residual.
|
|
143
|
+
|
|
144
|
+
## Continuation-point detection
|
|
145
|
+
|
|
146
|
+
Three independent signals, used in this priority:
|
|
147
|
+
|
|
148
|
+
1. **Inference identity** (primary). `inference_id` is unique per call and
|
|
149
|
+
stable across copies (0 token-field mismatches over 384 shared ids).
|
|
150
|
+
Ownership from the result chat (Attribution rule 1) needs no markers at
|
|
151
|
+
all and survives arbitrary re-projection.
|
|
152
|
+
2. **Timestamps** (secondary/corroborating). Every inference meta carries
|
|
153
|
+
`timestamp`; job issue time bounds the owned set. Use to cross-check rule
|
|
154
|
+
1 and to recover ownership when a result chat is missing. Measured
|
|
155
|
+
boundary in the chain: first's rows end 21:53:59Z, second starts 21:54:59Z,
|
|
156
|
+
second job issued 21:54:53Z.
|
|
157
|
+
3. **`job=` segment markers** (structural, not authoritative). Markers delimit
|
|
158
|
+
projected segments and are what the traversal follows, but the FINAL
|
|
159
|
+
segment of a continuation log mixes the parent's projected result with the
|
|
160
|
+
current job's own turns (measured: 257 = 237 + 20). Therefore markers
|
|
161
|
+
alone can only give an upper bound on inherited content; they must not be
|
|
162
|
+
the sole attribution mechanism. Markers are also where duplicate/cycle
|
|
163
|
+
safety and re-projection idempotence already live
|
|
164
|
+
(`Chat.project` `seen_inference` guard, meta.rb:410-431).
|
|
165
|
+
|
|
166
|
+
## Cost vs content (honest cost model)
|
|
167
|
+
|
|
168
|
+
Two quantities must both stay visible and stay distinct:
|
|
169
|
+
|
|
170
|
+
- **Authored (content) cost**: tokens of inferences the job itself issued.
|
|
171
|
+
Report per job: owned `{pt, ct, tt, cct, cwt, rt}` over its owned set.
|
|
172
|
+
- **Re-send (transmission) cost**: history re-paid on each of the job's own
|
|
173
|
+
calls. It is already inside per-call `pt`; surface it as a derived view per
|
|
174
|
+
job: `uncached_prompt = sum(pt) - sum(cct)` over owned calls, plus
|
|
175
|
+
`cache_read = sum(cct)`, `cache_write = sum(cwt)`. Example, job `80070846`:
|
|
176
|
+
pt 1,124,684, cache_read 1,028,352, uncached 96,332 - the continuation was
|
|
177
|
+
cheap precisely because history was cached, and that is now legible.
|
|
178
|
+
- Do NOT hide re-send by subtracting history tokens from per-call `pt`: each
|
|
179
|
+
call really did pay its prompt; the cache fields carry the discount.
|
|
180
|
+
- Do NOT add inherited rows into the job's cost "because they were sent":
|
|
181
|
+
they were paid inside the job's own calls' `pt` already; adding them again
|
|
182
|
+
is exactly the double count being removed.
|
|
183
|
+
- Cumulative keys (`pt_c` etc.) stay out of all sums; they are linear-chat
|
|
184
|
+
checkpoints (`Chat.meta`, meta.rb:140-160), not per-call evidence.
|
|
185
|
+
|
|
186
|
+
## Safety (dupes / cycles / degraded)
|
|
187
|
+
|
|
188
|
+
- **Duplicate inference copies**: already handled by identity grouping in
|
|
189
|
+
`provenance_token_events`; the attribution layer only adds an owner label
|
|
190
|
+
per event, so a row replayed into N logs is still one event with N
|
|
191
|
+
evidence records.
|
|
192
|
+
- **Conflicting token fields across copies**: existing conflict machinery
|
|
193
|
+
(immutable core `pt/ct/tt` + distinct `provider_response_id`, lines
|
|
194
|
+
632-662) stays authoritative; ownership never overrides a conflict, and
|
|
195
|
+
`strict:` keeps raising.
|
|
196
|
+
- **Cycles / self-edges**: traversal visited-set (`provenance_key`, lines
|
|
197
|
+
227-231) and the sidecar root-copy exclusion (`direct_chat_sidecar_files`,
|
|
198
|
+
provenance.rb:68-76) already prevent infinite recursion and self-duplication.
|
|
199
|
+
- **Mixed final segment**: attribution by identity/timestamp (not markers)
|
|
200
|
+
is immune; if both identity and timestamps were missing, fall back to
|
|
201
|
+
markers and flag the final segment as `ambiguous_tail` rather than
|
|
202
|
+
guessing.
|
|
203
|
+
- **Missing jobs**: unresolved receipt references produce warnings (measured
|
|
204
|
+
2 for `81e265cf`); attribution reports the job as `missing` with its
|
|
205
|
+
reference, and the residual accounting keeps global honesty (unknown, not
|
|
206
|
+
zero).
|
|
207
|
+
- **Log inspectability**: unchanged. Nothing is trimmed from logs; only the
|
|
208
|
+
accounting view filters.
|
|
209
|
+
|
|
210
|
+
## Future implementation surface (not implemented)
|
|
211
|
+
|
|
212
|
+
- `lib/scout/llm/chat/provenance.rb`
|
|
213
|
+
- new `Chat.provenance_token_attribution(root, warnings: nil, **opts)`:
|
|
214
|
+
reuse `provenance_token_events` (no new traversal), then resolve owners:
|
|
215
|
+
per job node from the same traversal, owned set = inference ids in
|
|
216
|
+
`job_result_chat_file` (fallback: log metas with `timestamp >= job
|
|
217
|
+
start`, from `Step#started`/info), inherited counts from `job=` markers.
|
|
218
|
+
- optional `scope: :owned` (or `attribution:` keyword) on
|
|
219
|
+
`provenance_token_totals` that filters events by owner == root job.
|
|
220
|
+
- Data shape: hash keyed by job path:
|
|
221
|
+
`{owned: {pt:, ct:, tt:, cct:, cwt:, rt:, uncached_prompt:, cache_read:},
|
|
222
|
+
inferences: n, inherited: {job_id => count}, missing: [refs],
|
|
223
|
+
ambiguous_tail: bool, warnings: [...]}`.
|
|
224
|
+
- CLI: `scout llm prov --tokens=owned` (and an `--attribution` table view),
|
|
225
|
+
reusing the existing totals plumbing.
|
|
226
|
+
- Flag/feature name suggestion: `token_attribution` (no env/config knob
|
|
227
|
+
needed; it is a view, not a change of evidence collection).
|
|
228
|
+
- Tests would mirror `test/scout/llm/chat/test_agent_meta_tokens.rb` with
|
|
229
|
+
fixture chats containing projected segments, a mixed final segment, a
|
|
230
|
+
missing referenced job, and a re-projected chat.
|
|
231
|
+
|
|
232
|
+
## Open questions
|
|
233
|
+
|
|
234
|
+
- **Unverified** - `pt` semantics vs `cct`: whether `pt` includes cached
|
|
235
|
+
tokens (provider convention) or excludes them. The uncached_prompt
|
|
236
|
+
derivation assumes inclusion; confirm against the API adapter code before
|
|
237
|
+
implementing.
|
|
238
|
+
- **Unverified** - `fin_fixes_recon` workspace copy: not found under
|
|
239
|
+
`/bulk` or `/home` (only `.scout/tmp` chat copies exist); the design rests
|
|
240
|
+
on job artifacts only. If the Cortex store reappears, confirm its
|
|
241
|
+
conversation file carries the same projected segments.
|
|
242
|
+
- **Unverified** - parentage claimed in the task brief (continuations as
|
|
243
|
+
children of `Planned/work 4bf865b6`): not measured; the measured structure
|
|
244
|
+
is a linear chain `c0b45189 -> 68b5b7fa -> 80070846` plus the separate
|
|
245
|
+
`de207a52 -> (missing) 81e265cf` case.
|
|
246
|
+
- **Open** - job start time source for the timestamp fallback: `Step` info
|
|
247
|
+
fields vs file mtime vs first owned inference timestamp; pick the most
|
|
248
|
+
robust and document it.
|
|
249
|
+
- **Open** - whether ownership should also attach to non-`chat` jobs whose
|
|
250
|
+
logs contain inference metas (e.g. `:json`/`:text` tasks that ran an
|
|
251
|
+
agent); the fallback rule covers them, but naming the owner for log-only
|
|
252
|
+
jobs needs a decision.
|
|
253
|
+
- **Open** - should `receipt_only` and owned views be combined in one CLI
|
|
254
|
+
table, or kept separate reports to avoid the documented non-additivity of
|
|
255
|
+
the coverage scopes?
|
data/scout-ai.gemspec
CHANGED
|
@@ -2,11 +2,11 @@
|
|
|
2
2
|
# DO NOT EDIT THIS FILE DIRECTLY
|
|
3
3
|
# Instead, edit Juwelier::Tasks in Rakefile, and run 'rake gemspec'
|
|
4
4
|
# -*- encoding: utf-8 -*-
|
|
5
|
-
# stub: scout-ai 2.
|
|
5
|
+
# stub: scout-ai 2.1.0 ruby lib
|
|
6
6
|
|
|
7
7
|
Gem::Specification.new do |s|
|
|
8
8
|
s.name = "scout-ai".freeze
|
|
9
|
-
s.version = "2.
|
|
9
|
+
s.version = "2.1.0".freeze
|
|
10
10
|
|
|
11
11
|
s.required_rubygems_version = Gem::Requirement.new(">= 0".freeze) if s.respond_to? :required_rubygems_version=
|
|
12
12
|
s.require_paths = ["lib".freeze]
|
|
@@ -52,6 +52,7 @@ Gem::Specification.new do |s|
|
|
|
52
52
|
"doc/user/WritingChats.md",
|
|
53
53
|
"lib/scout-ai.rb",
|
|
54
54
|
"lib/scout/llm/agent.rb",
|
|
55
|
+
"lib/scout/llm/agent/attach.rb",
|
|
55
56
|
"lib/scout/llm/agent/chat.rb",
|
|
56
57
|
"lib/scout/llm/agent/delegate.rb",
|
|
57
58
|
"lib/scout/llm/agent/iterate.rb",
|
|
@@ -152,6 +153,7 @@ Gem::Specification.new do |s|
|
|
|
152
153
|
"research/prompt-strategies-analysis.md",
|
|
153
154
|
"research/prov-verbosity-fix-notes.md",
|
|
154
155
|
"research/provenance-analysis.md",
|
|
156
|
+
"research/provenance-continuation-accounting.md",
|
|
155
157
|
"research/provenance-navigation-design.md",
|
|
156
158
|
"research/synthesis-report.md",
|
|
157
159
|
"research/tools-system-analysis.md",
|
data/scout_commands/llm/prov
CHANGED
|
@@ -17,12 +17,22 @@ Examine the provenance of a chat or workflow job
|
|
|
17
17
|
$ #{cmd} [<options>] <filename>
|
|
18
18
|
|
|
19
19
|
-h--help Print this help
|
|
20
|
-
-c--component Show per-component direct costs instead of
|
|
20
|
+
-c--component Show per-component direct costs (direct=) instead of subtree evidence closures (evidence=)
|
|
21
21
|
-f--flow Print a compact provenance flow
|
|
22
22
|
-l--long Show full filesystem paths instead of abbreviated names
|
|
23
23
|
-e--evidence List the deduplicated direct inference events and their evidence
|
|
24
24
|
--dot* Write the flow as Graphviz DOT
|
|
25
25
|
-p--plot* Render the flow as svg, png, or pdf
|
|
26
|
+
|
|
27
|
+
Default-tree numbers are evidence= subtree-deduplicated closures: they overlap
|
|
28
|
+
between nodes and are never per-part cost. Job nodes add delta= (direct tokens
|
|
29
|
+
of the persisted chat-typed result, the accounting delta). The root footer
|
|
30
|
+
always prints the authoritative deduplicated_total, in both modes; --component
|
|
31
|
+
relabels the node numbers direct=. cache= carries the cache-hit share of
|
|
32
|
+
prompt (cct/pt); the root footer adds fresh= (prompt not served from cache,
|
|
33
|
+
including cache_write) and cache_write= (Anthropic cache-write tokens).
|
|
34
|
+
Per-job delta= values sum to the root total on continuation chains but not on
|
|
35
|
+
general DAGs.
|
|
26
36
|
EOF
|
|
27
37
|
if options[:help]
|
|
28
38
|
defined?(scout_usage) ? scout_usage : puts(SOPT.doc)
|
|
@@ -38,6 +48,12 @@ evidence = options.delete(:evidence)
|
|
|
38
48
|
filename = ARGV.first
|
|
39
49
|
raise MissingParameterException, :filename if filename.nil?
|
|
40
50
|
|
|
51
|
+
# One parse-once cache for the whole CLI run (see
|
|
52
|
+
# Chat.provenance_chat_load): the initial traversal, the token-event
|
|
53
|
+
# collector and this script's own node computations then share a single
|
|
54
|
+
# Chat object per resolved path.
|
|
55
|
+
Chat.open_provenance_run_cache
|
|
56
|
+
|
|
41
57
|
# A persisted Step always has an info sidecar. A .files directory alone is NOT
|
|
42
58
|
# evidence: saved agent chats carry one too. Explicitly loading the Step after
|
|
43
59
|
# this evidence check avoids treating every readable chat as a job.
|
|
@@ -103,18 +119,36 @@ chat_cache = {}
|
|
|
103
119
|
direct_tokens = lambda do |key|
|
|
104
120
|
node = nodes[key]
|
|
105
121
|
if node[:kind] == :chat
|
|
106
|
-
chat = chat_cache[node[:path]] ||= Chat.
|
|
122
|
+
chat = chat_cache[node[:path]] ||= Chat.provenance_chat_load(node[:path])
|
|
107
123
|
Chat.token_totals([chat])
|
|
108
124
|
else
|
|
109
125
|
logs = edges.select { |e| e[:from] == key && e[:relation] == :log }
|
|
110
126
|
.collect { |e| nodes[e[:to]][:path] }.uniq
|
|
111
|
-
chats = logs.collect { |path| chat_cache[path] ||= Chat.
|
|
127
|
+
chats = logs.collect { |path| chat_cache[path] ||= Chat.provenance_chat_load(path) }
|
|
112
128
|
Chat.token_totals(chats)
|
|
113
129
|
end
|
|
114
130
|
rescue => error
|
|
115
131
|
warnings << { error: error, kind: node[:kind], object: node[:object], relation: :tokens }
|
|
116
132
|
Chat::TOKEN_KEYS.each_with_object({}) { |name, totals| totals[name.to_sym] = 0 }
|
|
117
133
|
end
|
|
134
|
+
# Per-job accounting delta: the direct tokens of the job's persisted
|
|
135
|
+
# chat-typed result (Chat.job_result_chat_file). This is the
|
|
136
|
+
# receipt-defined accounting object - the new chat the job returned - so on a
|
|
137
|
+
# continuation chain the per-job deltas sum to the root deduplicated_total
|
|
138
|
+
# while every ancestor's evidence= closure stays cumulative. Jobs whose
|
|
139
|
+
# result is not chat-typed omit the field (nil, never zero).
|
|
140
|
+
delta_cache = {}
|
|
141
|
+
delta_tokens = lambda do |key|
|
|
142
|
+
return delta_cache[key] if delta_cache.key?(key)
|
|
143
|
+
node = nodes[key]
|
|
144
|
+
delta_cache[key] = if node[:kind] == :job
|
|
145
|
+
file = Chat.job_result_chat_file(node[:object])
|
|
146
|
+
file ? Chat.token_totals([chat_cache[Chat.provenance_path(:chat, file)] ||= Chat.provenance_chat_load(file)]) : nil
|
|
147
|
+
end
|
|
148
|
+
rescue => error
|
|
149
|
+
warnings << { error: error, kind: node[:kind], object: node[:object], relation: :delta }
|
|
150
|
+
delta_cache[key] = nil
|
|
151
|
+
end
|
|
118
152
|
|
|
119
153
|
# ------------------------------------------------------------------
|
|
120
154
|
# Receipt (agent_meta) evidence: delegated inference events embedded in
|
|
@@ -204,21 +238,52 @@ end
|
|
|
204
238
|
# ------------------------------------------------------------------
|
|
205
239
|
# Formatting helpers
|
|
206
240
|
# ------------------------------------------------------------------
|
|
207
|
-
|
|
241
|
+
# `qualifier` names the total field after what it measures so a number can
|
|
242
|
+
# never be read as a different quantity: `evidence` (default tree; subtree
|
|
243
|
+
# deduplicated closure), `direct` (-c; per-component direct tokens),
|
|
244
|
+
# `deduplicated_total` (root footer; the authoritative cost). `events`
|
|
245
|
+
# appends the deduplicated event count where it is already computed (the
|
|
246
|
+
# footer), and `delta` inserts the job's result-chat delta right after it.
|
|
247
|
+
#
|
|
248
|
+
# Canonical field order, identical for node lines and the footer: qualifier
|
|
249
|
+
# total, [delta=], then the prompt axis `prompt= cache=<abs>@<rate> [fresh=]
|
|
250
|
+
# [cache_write=]`, then generation `cont= reason=`. `footer: true` adds the
|
|
251
|
+
# run-level economics (`fresh=`, `cache_write=`) that do not belong to a node
|
|
252
|
+
# label. The rate is always computed from RAW integer totals
|
|
253
|
+
# (`100.0 * cct / pt`): humanized values round at K/M boundaries and would
|
|
254
|
+
# corrupt boundary rates. `cache=` prints whenever pt is positive, including
|
|
255
|
+
# `cache=0@0.0%`; when pt is nil/0 the whole prompt axis is omitted (nothing
|
|
256
|
+
# to partition). A conflicting summed set can yield cct > pt; the raw ratio
|
|
257
|
+
# is then printed unclamped and the footer caveat line is the guard that
|
|
258
|
+
# labels such a figure best-effort.
|
|
259
|
+
def token_str(tokens, qualifier: nil, events: nil, delta: nil, footer: false)
|
|
208
260
|
parts = []
|
|
209
|
-
|
|
210
|
-
|
|
261
|
+
if tokens[:tt] && tokens[:tt].to_i > 0
|
|
262
|
+
value = Misc.human_number(tokens[:tt])
|
|
263
|
+
value += " (#{events} events)" if events
|
|
264
|
+
parts << "#{qualifier || 'total'}=#{value}"
|
|
265
|
+
parts << "delta=#{Misc.human_number(delta[:tt])}" if delta && delta[:tt].to_i > 0
|
|
266
|
+
end
|
|
267
|
+
pt = tokens[:pt].to_i
|
|
268
|
+
if pt > 0
|
|
269
|
+
cct = tokens[:cct].to_i
|
|
270
|
+
parts << "prompt=#{Misc.human_number(pt)}"
|
|
271
|
+
parts << "cache=#{Misc.human_number(cct)}@#{format('%.1f', 100.0 * cct / pt)}%"
|
|
272
|
+
if footer
|
|
273
|
+
parts << "fresh=#{Misc.human_number(pt - cct)}"
|
|
274
|
+
parts << "cache_write=#{Misc.human_number(tokens[:cwt].to_i)}" if tokens[:cwt] && tokens[:cwt].to_i > 0
|
|
275
|
+
end
|
|
276
|
+
end
|
|
211
277
|
parts << "cont=#{Misc.human_number(tokens[:ct])}" if tokens[:ct] && tokens[:ct].to_i > 0
|
|
212
|
-
parts << "cache=#{Misc.human_number(tokens[:cct])}" if tokens[:cct] && tokens[:cct].to_i > 0
|
|
213
278
|
parts << "reason=#{Misc.human_number(tokens[:rt])}" if tokens[:rt] && tokens[:rt].to_i > 0
|
|
214
279
|
parts * ' '
|
|
215
280
|
end
|
|
216
281
|
|
|
217
|
-
def report_line(kind, label, tokens, offset: 0)
|
|
282
|
+
def report_line(kind, label, tokens, offset: 0, qualifier: nil, delta: nil, events: nil)
|
|
218
283
|
color = kind == :job ? :yellow : :green
|
|
219
284
|
parts = [' ' * (offset * 2)]
|
|
220
285
|
parts << Log.color(color, kind.to_s)
|
|
221
|
-
tok = token_str(tokens)
|
|
286
|
+
tok = token_str(tokens, qualifier: qualifier, events: events, delta: delta)
|
|
222
287
|
parts << tok unless tok.empty?
|
|
223
288
|
parts << Log.color(:blue, label.to_s) unless label.to_s.empty?
|
|
224
289
|
parts * ' '
|
|
@@ -279,14 +344,34 @@ unless flow || dot_file || plot_file
|
|
|
279
344
|
node = nodes[key]
|
|
280
345
|
parent_node = parent_key ? nodes[parent_key] : nil
|
|
281
346
|
|
|
282
|
-
#
|
|
283
|
-
|
|
347
|
+
# Hidden nodes (result chats duplicating a job, top-level agent.chat
|
|
348
|
+
# logs) are not printed, but PROJECTION ONLY: the walk still descends
|
|
349
|
+
# through them so their children are attributed to the current display
|
|
350
|
+
# parent. This is what makes delegated children of tree-hidden carrier
|
|
351
|
+
# chats visible without adding, removing or rewriting any traversal
|
|
352
|
+
# edge: the :agent_job/:job edges keep pointing at the carrier chat.
|
|
353
|
+
# - no "printed" slot is consumed, so offsets stay as if the carrier
|
|
354
|
+
# were not there
|
|
355
|
+
# - children keep their own incoming relation (delegated-job label)
|
|
356
|
+
# - the global `printed` set still dedups nodes discovered from
|
|
357
|
+
# several carriers (first visit wins)
|
|
358
|
+
if hidden_node?(node, relation)
|
|
359
|
+
(adjacency[key] || []).each do |edge|
|
|
360
|
+
print_tree.call(edge[:to], offset, edge[:relation], parent_key)
|
|
361
|
+
end
|
|
362
|
+
return
|
|
363
|
+
end
|
|
284
364
|
|
|
285
365
|
printed << key
|
|
286
366
|
label = node_label(node, parent_node, root_key, key, long)
|
|
287
367
|
# A job reached through an agent_meta receipt is a delegated producer.
|
|
288
368
|
label = "delegated-job #{label}" if node[:kind] == :job && relation == :agent_job
|
|
289
|
-
|
|
369
|
+
# Both modes: job nodes carry their accounting delta (in -c the
|
|
370
|
+
# direct=/delta= contrast on one line is the question the mode answers).
|
|
371
|
+
delta = node[:kind] == :job ? delta_tokens.call(key) : nil
|
|
372
|
+
puts report_line(node[:kind], label, node_tokens.call(key), offset: offset,
|
|
373
|
+
qualifier: component ? 'direct' : 'evidence',
|
|
374
|
+
delta: delta)
|
|
290
375
|
|
|
291
376
|
# One compact annotation line for the delegated usage embedded in this
|
|
292
377
|
# chat's own outputs; never one line per receipt.
|
|
@@ -309,6 +394,25 @@ unless flow || dot_file || plot_file
|
|
|
309
394
|
end
|
|
310
395
|
print_tree.call(root_key, 0, nil, nil)
|
|
311
396
|
|
|
397
|
+
# Root footer (BOTH modes): the one authoritative cost figure, always
|
|
398
|
+
# printed, so -c no longer lacks a global bound. It carries the run-level
|
|
399
|
+
# economics (deduplicated event count, fresh=, cache_write=) that do not
|
|
400
|
+
# belong to a node label. It reuses the root aggregate computed for the
|
|
401
|
+
# root node line, which IS the root :deduplicated_total (every deduplicated
|
|
402
|
+
# event once). Identity conflicts demote it to best-effort; this footer
|
|
403
|
+
# caveat is the single home of the non-authoritative warning in both modes
|
|
404
|
+
# (the -c scope block no longer repeats it).
|
|
405
|
+
conflicts = {}
|
|
406
|
+
root_totals = Chat.provenance_token_totals(root, warnings: receipt_warnings, conflicts: conflicts)
|
|
407
|
+
puts Log.color(:cyan, 'root ') +
|
|
408
|
+
token_str(root_totals, qualifier: 'deduplicated_total',
|
|
409
|
+
events: token_events_for.call.length, footer: true) + ' ' +
|
|
410
|
+
Log.color(:cyan, '(authoritative cost; per-node evidence=/direct= values overlap)')
|
|
411
|
+
unless conflicts[:authoritative]
|
|
412
|
+
puts Log.color(:red, 'unresolved identity conflicts: ') +
|
|
413
|
+
"#{conflicts[:events]} conflicting event(s); totals above are best-effort (canonical evidence only), not exact"
|
|
414
|
+
end
|
|
415
|
+
|
|
312
416
|
# Component mode: make the evidence coverage behind the numbers explicit
|
|
313
417
|
# once receipts are present; chats without receipts keep the previous output
|
|
314
418
|
# exactly. These are COVERAGE figures, not additive cost categories:
|
|
@@ -316,8 +420,7 @@ unless flow || dot_file || plot_file
|
|
|
316
420
|
# both a saved child log and a receipt. Only deduplicated_total is the
|
|
317
421
|
# cost, and receipt_only is the disjoint delegated contribution.
|
|
318
422
|
if component && has_receipt_evidence
|
|
319
|
-
coverage = [[:
|
|
320
|
-
[:chat_evidence, 'events with saved chat/log evidence'],
|
|
423
|
+
coverage = [[:chat_evidence, 'events with saved chat/log evidence'],
|
|
321
424
|
[:receipt_evidence, 'events inside agent_meta receipts (overlaps chat_evidence)'],
|
|
322
425
|
[:receipt_only, 'receipt evidence with no saved chat/log']]
|
|
323
426
|
coverage.each do |scope, note|
|
|
@@ -326,13 +429,6 @@ unless flow || dot_file || plot_file
|
|
|
326
429
|
rendered = 'total=0' if rendered.empty?
|
|
327
430
|
puts Log.color(:cyan, "evidence #{scope}:") + ' ' + rendered + ' ' + Log.color(:cyan, "(#{note})")
|
|
328
431
|
end
|
|
329
|
-
|
|
330
|
-
conflict_info = {}
|
|
331
|
-
Chat.provenance_token_totals(root, warnings: [], conflicts: conflict_info)
|
|
332
|
-
unless conflict_info[:authoritative]
|
|
333
|
-
puts Log.color(:red, 'unresolved identity conflicts: ') +
|
|
334
|
-
"#{conflict_info[:events]} conflicting event(s); totals above are best-effort (canonical evidence only), not exact"
|
|
335
|
-
end
|
|
336
432
|
end
|
|
337
433
|
end
|
|
338
434
|
|
|
@@ -600,3 +696,7 @@ receipt_warnings.each do |warning|
|
|
|
600
696
|
Log.warn "agent_meta receipt: #{warning[:reason]} in #{warning[:source]} at #{location.inspect} call=#{warning[:call_id]}#{warning[:message] ? ' - ' + warning[:message] : ''}"
|
|
601
697
|
end
|
|
602
698
|
end
|
|
699
|
+
|
|
700
|
+
# End of the CLI run: drop the parse-once cache so a following run in the
|
|
701
|
+
# same process (tests, embedded use) re-parses freshly.
|
|
702
|
+
Chat.close_provenance_run_cache
|
|
@@ -32,8 +32,14 @@ module AgentMetaFixtures
|
|
|
32
32
|
end
|
|
33
33
|
|
|
34
34
|
# JSON payload of a function_call_output envelope carrying agent_meta.
|
|
35
|
-
|
|
36
|
-
|
|
35
|
+
# `envelope:` selects which serialized key carries the receipts:
|
|
36
|
+
# :agent_meta is the legacy envelope, :meta the current one written since
|
|
37
|
+
# the dual-envelope reader landed (lib/scout/llm/tools/call.rb).
|
|
38
|
+
# `name` only decorates the envelope; the receipt lifting ignores it.
|
|
39
|
+
def receipt_output(call_id, agent_meta, name: 'ask', content: 'child answer', envelope: :agent_meta)
|
|
40
|
+
payload = {name: name, content: content, id: call_id}
|
|
41
|
+
payload[envelope] = agent_meta
|
|
42
|
+
payload.to_json
|
|
37
43
|
end
|
|
38
44
|
|
|
39
45
|
# Persisted chat text with one paired ask call per receipt entry. Hash keys
|
|
@@ -43,11 +49,13 @@ module AgentMetaFixtures
|
|
|
43
49
|
# Message indexes produced by Chat.parse (single user turn, no leading
|
|
44
50
|
# empty user message since 49c0d20):
|
|
45
51
|
# 0 user, then per receipt: function_call, function_call_output.
|
|
46
|
-
|
|
52
|
+
# `tool` sets the function_call name; receipts are lifted from the output
|
|
53
|
+
# envelope regardless of it, so tests can pin that no tool name is special.
|
|
54
|
+
def receipt_chat_text(receipts, extra: nil, envelope: :agent_meta, tool: 'ask')
|
|
47
55
|
lines = ['user: Run the worker']
|
|
48
56
|
receipts.each do |call_id, agent_meta|
|
|
49
|
-
lines << 'function_call: ' + %({"name":"
|
|
50
|
-
lines << 'function_call_output: ' + receipt_output(call_id, agent_meta)
|
|
57
|
+
lines << 'function_call: ' + %({"name":"#{tool}","arguments":{},"id":"#{call_id}"})
|
|
58
|
+
lines << 'function_call_output: ' + receipt_output(call_id, agent_meta, envelope: envelope)
|
|
51
59
|
end
|
|
52
60
|
lines.concat(Array(extra)) if extra
|
|
53
61
|
lines << 'assistant: done'
|
|
@@ -128,4 +136,82 @@ module AgentMetaFixtures
|
|
|
128
136
|
"function_call_output: #{payload}\n" +
|
|
129
137
|
"assistant: done\n"
|
|
130
138
|
end
|
|
139
|
+
|
|
140
|
+
# Continuation-chain fixture (theme-2 A/B/C shape, in-repo). Three
|
|
141
|
+
# chat-typed jobs A -> B -> C whose agent logs are cumulative histories
|
|
142
|
+
# (each carries the projected parent metas before its own), so per-root
|
|
143
|
+
# closures are cumulative while each job's own delta is disjoint. Every
|
|
144
|
+
# meta line is a plain direct chat meta (not a receipt).
|
|
145
|
+
#
|
|
146
|
+
# Token plan (all figures arithmetically checkable):
|
|
147
|
+
# A own: 3 metas pt=40 tt=50 cct=30 cwt=5 -> delta tt=150, pt=120, cct=90
|
|
148
|
+
# B own: 2 metas pt=100 tt=120 (no cache) -> delta tt=240, pt=200
|
|
149
|
+
# C own: 1 meta pt=0 tt=600 (pt missing)-> delta tt=600, pt=0
|
|
150
|
+
# chain closure at C: 6 events, tt=990, pt=320, cct=90, cwt=15, ct=6
|
|
151
|
+
# sum of deltas 150+240+600 = 990 == root deduplicated_total
|
|
152
|
+
#
|
|
153
|
+
# Also builds a TSV-typed job reached from `parent.chat` through a
|
|
154
|
+
# receipt, to pin delta= omission for non-chat results in both modes.
|
|
155
|
+
# Returns [a_job, b_job, c_job, tsv_job, parent_chat].
|
|
156
|
+
def continuation_chain(dir)
|
|
157
|
+
meta = lambda do |id, pt, tt, cct: 0, cwt: 0|
|
|
158
|
+
parts = ["pt=#{pt}", 'ct=1', "tt=#{tt}"]
|
|
159
|
+
parts << "cct=#{cct}" if cct > 0
|
|
160
|
+
parts << "cwt=#{cwt}" if cwt > 0
|
|
161
|
+
parts << "inference_id=#{id}" << 'timestamp=2026-09-04T00:00:00Z'
|
|
162
|
+
"meta: " + parts * ' '
|
|
163
|
+
end
|
|
164
|
+
|
|
165
|
+
a_metas = (1..3).collect { |i| meta.call("a#{i}", 40, 50, cct: 30, cwt: 5) }
|
|
166
|
+
b_metas = (1..2).collect { |i| meta.call("b#{i}", 100, 120) }
|
|
167
|
+
c_meta = meta.call('c1', 0, 600)
|
|
168
|
+
|
|
169
|
+
a = make_job(dir, 'Cortex/continue/Default_a1', logs: {'agent.chat' => ''})
|
|
170
|
+
File.write(a + '.info', {dependencies: [], type: :chat}.to_json)
|
|
171
|
+
File.write(a, (['user: start'] + a_metas + ["meta: job=#{a}"]) * "
|
|
172
|
+
" + "
|
|
173
|
+
")
|
|
174
|
+
File.write(File.join(a + '.files', 'log', 'agent.chat'), a_metas * "
|
|
175
|
+
" + "
|
|
176
|
+
")
|
|
177
|
+
|
|
178
|
+
b = make_job(dir, 'Cortex/continue/Default_b2', logs: {'agent.chat' => ''})
|
|
179
|
+
File.write(b + '.info', {dependencies: [], type: :chat}.to_json)
|
|
180
|
+
File.write(b, (['user: continue B'] + b_metas + ["meta: job=#{b}"]) * "
|
|
181
|
+
" + "
|
|
182
|
+
")
|
|
183
|
+
File.write(File.join(b + '.files', 'log', 'agent.chat'),
|
|
184
|
+
(["meta: job=#{a}"] + a_metas + b_metas) * "
|
|
185
|
+
" + "
|
|
186
|
+
")
|
|
187
|
+
|
|
188
|
+
c = make_job(dir, 'Cortex/continue/Default_c3', logs: {'agent.chat' => ''})
|
|
189
|
+
File.write(c + '.info', {dependencies: [], type: :chat}.to_json)
|
|
190
|
+
File.write(c, (['user: continue C'] + [c_meta, "meta: job=#{c}"]) * "
|
|
191
|
+
" + "
|
|
192
|
+
")
|
|
193
|
+
File.write(File.join(c + '.files', 'log', 'agent.chat'),
|
|
194
|
+
(["meta: job=#{b}"] + a_metas + b_metas + [c_meta]) * "
|
|
195
|
+
" + "
|
|
196
|
+
")
|
|
197
|
+
|
|
198
|
+
# TSV-typed job (no chat result): must omit delta= while still showing
|
|
199
|
+
# its own log evidence.
|
|
200
|
+
tsv = make_job(dir, 'Other/step/Default_d4',
|
|
201
|
+
result: "a b
|
|
202
|
+
1 2
|
|
203
|
+
",
|
|
204
|
+
logs: {'agent.chat' => meta.call('d1', 7, 8) + "
|
|
205
|
+
"})
|
|
206
|
+
info = JSON.parse(File.read(tsv + '.info'))
|
|
207
|
+
File.write(tsv + '.info', {dependencies: info['dependencies'], type: :tsv}.to_json)
|
|
208
|
+
|
|
209
|
+
parent = write_chat(dir, 'parent.chat',
|
|
210
|
+
receipt_chat_text(
|
|
211
|
+
{'a1' => [meta_receipt("job=#{tsv}")]},
|
|
212
|
+
extra: ['meta: pt=10 ct=1 tt=11 inference_id=p1']
|
|
213
|
+
))
|
|
214
|
+
|
|
215
|
+
[a, b, c, tsv, parent]
|
|
216
|
+
end
|
|
131
217
|
end
|