scout-ai 2.0.0 → 2.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,255 @@
1
+ # Provenance token accounting for continued conversations (design, not implemented)
2
+
3
+ Step 3, Part 2 of the provenance-gap plan. Part 1 (provenance follows `job=`
4
+ receipt edges) is done and validated; this document designs token accounting
5
+ for continuation chains (Cortex `continue` jobs and any `chat_task` that
6
+ follows a previously projected chat). No library code was changed for this
7
+ design; probes live in `tmp/prov-cont/measure*.rb` and
8
+ `tmp/cortex-continuation-accounting-probes.md`.
9
+
10
+ All numbers below are measured on the real example chain unless marked
11
+ otherwise. Jobs live under `/home/mvazque2/.scout/var/jobs/Cortex/continue/`.
12
+
13
+ ## Evidence
14
+
15
+ ### The example chain
16
+
17
+ Three sequential continuations of one conversation, `fin_fixes_recon`
18
+ (workspace copy not found in this environment, see Open questions):
19
+
20
+ | job | result chat (`Default_<k>.chat`) | inference metas | pt sum |
21
+ |-----|----------------------------------|-----------------|--------|
22
+ | `c0b45189` | 445 messages, 1 `job=` marker, 147 metas | 147 | 6,297,210 |
23
+ | `68b5b7fa` | 711 messages, 1 `job=` marker, 237 metas | 237 | 10,707,400 |
24
+ | `80070846` | 60 messages, 1 `job=` marker, 20 metas | 20 | 1,124,684 |
25
+
26
+ Each result chat starts with its own `job=` meta marker followed by exactly
27
+ that job's own new inferences (this is `Chat.project(self.short_path, result)`
28
+ with `result = agent.current_chat - agent.start_chat`,
29
+ lib/scout/llm/agent/workflow.rb:127-139, projection at
30
+ lib/scout/llm/chat/process/meta.rb:410).
31
+
32
+ ### What the agent log actually contains (replay measured)
33
+
34
+ `Default_80070846...chat.files/agent.chat` (the third, latest job):
35
+
36
+ - 404 inference metas, 404 *distinct* `inference_id`s, 0 duplicated ids
37
+ (`tmp/prov-cont/measure41.rb`).
38
+ - Organized in projected segments delimited by `job=` meta markers:
39
+ - segment `job=...c0b45189...chat`: 147 metas, pt 6,297,210,
40
+ timestamps 2026-09-02T20:32:48Z .. 21:53:59Z;
41
+ - segment `job=...68b5b7fa...chat`: 257 metas, pt 11,832,084,
42
+ timestamps 21:54:59Z .. 22:48:55Z
43
+ (`measure43.rb`, `measure47.rb`).
44
+ - The last segment is MIXED: 257 = 237 inferences authored by job
45
+ `68b5b7fa` + 20 authored by `80070846` itself. There is no marker between
46
+ the second job's projected result and the third job's own new turns; the
47
+ third job appended its turns after `start_chat.follow chat`
48
+ (lib/scout/llm/agent/workflow.rb:85-87) of the projected second result.
49
+ - `Default_68b5b7fa...files/agent.chat`: 384 metas, all inside the
50
+ `job=c0b45189` segment (147 replayed + 237 own new, unmarked).
51
+ - `Default_c0b45189...files/agent.chat`: 147 unmarked metas, nothing replayed.
52
+
53
+ So: a continuation log carries the entire grown conversation; every replayed
54
+ inference retains its original per-call token fields (`pt`, `ct`, `tt`,
55
+ `cct`, `cwt`, `rt`, plus cumulative `*_c` and session `*_s` keys, all 404
56
+ rows), original `inference_id`, `provider_response_id` and `timestamp`.
57
+
58
+ ### Inference identity is stable across copies
59
+
60
+ Shared `inference_id`s between the second and third logs: 384; token-field
61
+ (`pt/ct/tt/cct`) mismatches among them: 0 (`measure39.rb`). The third log's
62
+ inherited rows sum to pt 17,004,610 / cct 16,288,896, byte-identical totals
63
+ to the second job's own log content. Identity, not position, is the reliable
64
+ key.
65
+
66
+ ### Traversal already reaches the whole chain
67
+
68
+ `Chat.provenance_chat_files(third result chat)` returns 6 files: the three
69
+ result chats plus the three `.files/agent.chat` logs (`measure44.rb`), via
70
+ the `job=` marker edges (Part 1). Per-call meta keys observed in the third
71
+ log, all 404 rows: `inference_id`, `timestamp`, `provider_response_id`,
72
+ `pt/ct/tt/cct/cwt/rt`, `pt_c...rt_c`, `pt_s...rt_s`, plus `reas` on 12 rows.
73
+
74
+ ## Current behavior (measured)
75
+
76
+ `Chat.provenance_token_totals(root)` (lib/scout/llm/chat/provenance.rb:729)
77
+ collects per-inference events (`provenance_token_events`, line 499),
78
+ deduplicates them by identity (`[:inference_id, id]`, else
79
+ `provider_response_id`, else lineage, else receipt address; lines 577-608),
80
+ then sums the canonical event's tokens (lines 706-708).
81
+
82
+ Measured per-root totals (`measure44.rb`):
83
+
84
+ | root | pt | interpretation |
85
+ |------|----|----------------|
86
+ | `c0b45189` | 6,297,210 | own only |
87
+ | `68b5b7fa` | 17,004,610 | cumulative: own 10,707,400 + first's 6,297,210 |
88
+ | `80070846` | 18,129,294 | cumulative: own 1,124,684 + 17,004,610 |
89
+
90
+ Verdicts:
91
+
92
+ 1. Global deduplicated total does NOT double count. 18,129,294 =
93
+ 6,297,210 + 10,707,400 + 1,124,684 exactly. Replayed copies of the same
94
+ inference in later logs collapse onto one event; the code's `next if
95
+ meta[:job]` filters (lines 534, 557) plus identity grouping do their job.
96
+ 2. Per-job root totals ARE cumulative. A continuation job's root total
97
+ includes every ancestor inference, so summing per-root totals across a
98
+ chain double counts (6,297,210 would be counted three times), and a
99
+ naive per-job cost read overstates job `68b5b7fa` by 6.3M pt. Today
100
+ there is no per-job "owned" breakdown: scope symbols
101
+ (`:deduplicated_total`, `:chat_evidence`, `:receipt_evidence`,
102
+ `:receipt_only`, lines 737-750) partition evidence coverage, not authorship.
103
+ 3. Re-send cost is present but unseparated. The third job's own 20 calls
104
+ cost pt 1,124,684 with cct 1,028,352 (91.5% of prompt tokens served from
105
+ cache; uncached prompt = pt - cct = 96,332). Per-call `pt` already
106
+ includes re-sent history for every call, which is the honest provider
107
+ view, but nothing at job level separates "what this job authored" from
108
+ "history this job re-paid for".
109
+ 4. Degraded case (missing child `81e265cf`): root
110
+ `Planned/work/Default_de207a52...chat` totals pt 37,969,140 with 2
111
+ warnings, both `unresolved_job_reference` for
112
+ `Cortex/continue/Default_81e265cf...chat` (file absent on disk). The
113
+ parent's log carries that job's projected segment with 0 inference metas
114
+ (its result was an exception receipt), so nothing is silently counted or
115
+ lost; the missing job's own cost is simply unknown, reported only as a
116
+ warning.
117
+
118
+ ## Attribution rules
119
+
120
+ Core rule: **an inference belongs to exactly one job - the job that issued
121
+ the model call - and its tokens are attributed to that job only; any other
122
+ copy of the same `inference_id` is evidence, never a second charge.**
123
+
124
+ 1. Ownership set of a job J = the `inference_id`s of the inference metas in
125
+ J's persisted result chat (for `chat` result-type jobs, `job_result_chat_file`,
126
+ provenance.rb:83). Measured: this reproduces the exact deltas above
127
+ (147 / 237 / 20).
128
+ 2. Fallback ownership (result chat absent, e.g. non-chat jobs or pruned
129
+ results): inference metas in J's own logs (`direct_job_chat_files`,
130
+ provenance.rb:44) whose `timestamp` >= J's start time. Measured: job
131
+ `68b5b7fa` issued 2026-09-02T21:54:53Z; the 237 rows with
132
+ `timestamp >= issued` equal its result-chat set exactly (set difference 0).
133
+ 3. Rows in J's logs belonging to other jobs are labeled inherited, with
134
+ provenance `{owning job id -> count}` derived from the `job=` segment
135
+ markers around them; they never enter J's owned sums.
136
+ 4. Global total = sum over distinct `inference_id`s (unchanged; already
137
+ correct today). Per-job owned totals must sum to the global total when
138
+ the chain is complete; the residual (global - sum of known owners) is
139
+ attributed to `unknown`, driven by missing jobs.
140
+ 5. A missing job referenced by a receipt keeps its `unresolved_job_reference`
141
+ warning; its contribution is reported as `missing`, never as zero-owned
142
+ and never dropped from the residual.
143
+
144
+ ## Continuation-point detection
145
+
146
+ Three independent signals, used in this priority:
147
+
148
+ 1. **Inference identity** (primary). `inference_id` is unique per call and
149
+ stable across copies (0 token-field mismatches over 384 shared ids).
150
+ Ownership from the result chat (Attribution rule 1) needs no markers at
151
+ all and survives arbitrary re-projection.
152
+ 2. **Timestamps** (secondary/corroborating). Every inference meta carries
153
+ `timestamp`; job issue time bounds the owned set. Use to cross-check rule
154
+ 1 and to recover ownership when a result chat is missing. Measured
155
+ boundary in the chain: first's rows end 21:53:59Z, second starts 21:54:59Z,
156
+ second job issued 21:54:53Z.
157
+ 3. **`job=` segment markers** (structural, not authoritative). Markers delimit
158
+ projected segments and are what the traversal follows, but the FINAL
159
+ segment of a continuation log mixes the parent's projected result with the
160
+ current job's own turns (measured: 257 = 237 + 20). Therefore markers
161
+ alone can only give an upper bound on inherited content; they must not be
162
+ the sole attribution mechanism. Markers are also where duplicate/cycle
163
+ safety and re-projection idempotence already live
164
+ (`Chat.project` `seen_inference` guard, meta.rb:410-431).
165
+
166
+ ## Cost vs content (honest cost model)
167
+
168
+ Two quantities must both stay visible and stay distinct:
169
+
170
+ - **Authored (content) cost**: tokens of inferences the job itself issued.
171
+ Report per job: owned `{pt, ct, tt, cct, cwt, rt}` over its owned set.
172
+ - **Re-send (transmission) cost**: history re-paid on each of the job's own
173
+ calls. It is already inside per-call `pt`; surface it as a derived view per
174
+ job: `uncached_prompt = sum(pt) - sum(cct)` over owned calls, plus
175
+ `cache_read = sum(cct)`, `cache_write = sum(cwt)`. Example, job `80070846`:
176
+ pt 1,124,684, cache_read 1,028,352, uncached 96,332 - the continuation was
177
+ cheap precisely because history was cached, and that is now legible.
178
+ - Do NOT hide re-send by subtracting history tokens from per-call `pt`: each
179
+ call really did pay its prompt; the cache fields carry the discount.
180
+ - Do NOT add inherited rows into the job's cost "because they were sent":
181
+ they were paid inside the job's own calls' `pt` already; adding them again
182
+ is exactly the double count being removed.
183
+ - Cumulative keys (`pt_c` etc.) stay out of all sums; they are linear-chat
184
+ checkpoints (`Chat.meta`, meta.rb:140-160), not per-call evidence.
185
+
186
+ ## Safety (dupes / cycles / degraded)
187
+
188
+ - **Duplicate inference copies**: already handled by identity grouping in
189
+ `provenance_token_events`; the attribution layer only adds an owner label
190
+ per event, so a row replayed into N logs is still one event with N
191
+ evidence records.
192
+ - **Conflicting token fields across copies**: existing conflict machinery
193
+ (immutable core `pt/ct/tt` + distinct `provider_response_id`, lines
194
+ 632-662) stays authoritative; ownership never overrides a conflict, and
195
+ `strict:` keeps raising.
196
+ - **Cycles / self-edges**: traversal visited-set (`provenance_key`, lines
197
+ 227-231) and the sidecar root-copy exclusion (`direct_chat_sidecar_files`,
198
+ provenance.rb:68-76) already prevent infinite recursion and self-duplication.
199
+ - **Mixed final segment**: attribution by identity/timestamp (not markers)
200
+ is immune; if both identity and timestamps were missing, fall back to
201
+ markers and flag the final segment as `ambiguous_tail` rather than
202
+ guessing.
203
+ - **Missing jobs**: unresolved receipt references produce warnings (measured
204
+ 2 for `81e265cf`); attribution reports the job as `missing` with its
205
+ reference, and the residual accounting keeps global honesty (unknown, not
206
+ zero).
207
+ - **Log inspectability**: unchanged. Nothing is trimmed from logs; only the
208
+ accounting view filters.
209
+
210
+ ## Future implementation surface (not implemented)
211
+
212
+ - `lib/scout/llm/chat/provenance.rb`
213
+ - new `Chat.provenance_token_attribution(root, warnings: nil, **opts)`:
214
+ reuse `provenance_token_events` (no new traversal), then resolve owners:
215
+ per job node from the same traversal, owned set = inference ids in
216
+ `job_result_chat_file` (fallback: log metas with `timestamp >= job
217
+ start`, from `Step#started`/info), inherited counts from `job=` markers.
218
+ - optional `scope: :owned` (or `attribution:` keyword) on
219
+ `provenance_token_totals` that filters events by owner == root job.
220
+ - Data shape: hash keyed by job path:
221
+ `{owned: {pt:, ct:, tt:, cct:, cwt:, rt:, uncached_prompt:, cache_read:},
222
+ inferences: n, inherited: {job_id => count}, missing: [refs],
223
+ ambiguous_tail: bool, warnings: [...]}`.
224
+ - CLI: `scout llm prov --tokens=owned` (and an `--attribution` table view),
225
+ reusing the existing totals plumbing.
226
+ - Flag/feature name suggestion: `token_attribution` (no env/config knob
227
+ needed; it is a view, not a change of evidence collection).
228
+ - Tests would mirror `test/scout/llm/chat/test_agent_meta_tokens.rb` with
229
+ fixture chats containing projected segments, a mixed final segment, a
230
+ missing referenced job, and a re-projected chat.
231
+
232
+ ## Open questions
233
+
234
+ - **Unverified** - `pt` semantics vs `cct`: whether `pt` includes cached
235
+ tokens (provider convention) or excludes them. The uncached_prompt
236
+ derivation assumes inclusion; confirm against the API adapter code before
237
+ implementing.
238
+ - **Unverified** - `fin_fixes_recon` workspace copy: not found under
239
+ `/bulk` or `/home` (only `.scout/tmp` chat copies exist); the design rests
240
+ on job artifacts only. If the Cortex store reappears, confirm its
241
+ conversation file carries the same projected segments.
242
+ - **Unverified** - parentage claimed in the task brief (continuations as
243
+ children of `Planned/work 4bf865b6`): not measured; the measured structure
244
+ is a linear chain `c0b45189 -> 68b5b7fa -> 80070846` plus the separate
245
+ `de207a52 -> (missing) 81e265cf` case.
246
+ - **Open** - job start time source for the timestamp fallback: `Step` info
247
+ fields vs file mtime vs first owned inference timestamp; pick the most
248
+ robust and document it.
249
+ - **Open** - whether ownership should also attach to non-`chat` jobs whose
250
+ logs contain inference metas (e.g. `:json`/`:text` tasks that ran an
251
+ agent); the fallback rule covers them, but naming the owner for log-only
252
+ jobs needs a decision.
253
+ - **Open** - should `receipt_only` and owned views be combined in one CLI
254
+ table, or kept separate reports to avoid the documented non-additivity of
255
+ the coverage scopes?
data/scout-ai.gemspec CHANGED
@@ -2,11 +2,11 @@
2
2
  # DO NOT EDIT THIS FILE DIRECTLY
3
3
  # Instead, edit Juwelier::Tasks in Rakefile, and run 'rake gemspec'
4
4
  # -*- encoding: utf-8 -*-
5
- # stub: scout-ai 2.0.0 ruby lib
5
+ # stub: scout-ai 2.1.0 ruby lib
6
6
 
7
7
  Gem::Specification.new do |s|
8
8
  s.name = "scout-ai".freeze
9
- s.version = "2.0.0".freeze
9
+ s.version = "2.1.0".freeze
10
10
 
11
11
  s.required_rubygems_version = Gem::Requirement.new(">= 0".freeze) if s.respond_to? :required_rubygems_version=
12
12
  s.require_paths = ["lib".freeze]
@@ -52,6 +52,7 @@ Gem::Specification.new do |s|
52
52
  "doc/user/WritingChats.md",
53
53
  "lib/scout-ai.rb",
54
54
  "lib/scout/llm/agent.rb",
55
+ "lib/scout/llm/agent/attach.rb",
55
56
  "lib/scout/llm/agent/chat.rb",
56
57
  "lib/scout/llm/agent/delegate.rb",
57
58
  "lib/scout/llm/agent/iterate.rb",
@@ -152,6 +153,7 @@ Gem::Specification.new do |s|
152
153
  "research/prompt-strategies-analysis.md",
153
154
  "research/prov-verbosity-fix-notes.md",
154
155
  "research/provenance-analysis.md",
156
+ "research/provenance-continuation-accounting.md",
155
157
  "research/provenance-navigation-design.md",
156
158
  "research/synthesis-report.md",
157
159
  "research/tools-system-analysis.md",
@@ -17,12 +17,22 @@ Examine the provenance of a chat or workflow job
17
17
  $ #{cmd} [<options>] <filename>
18
18
 
19
19
  -h--help Print this help
20
- -c--component Show per-component direct costs instead of aggregate totals
20
+ -c--component Show per-component direct costs (direct=) instead of subtree evidence closures (evidence=)
21
21
  -f--flow Print a compact provenance flow
22
22
  -l--long Show full filesystem paths instead of abbreviated names
23
23
  -e--evidence List the deduplicated direct inference events and their evidence
24
24
  --dot* Write the flow as Graphviz DOT
25
25
  -p--plot* Render the flow as svg, png, or pdf
26
+
27
+ Default-tree numbers are evidence= subtree-deduplicated closures: they overlap
28
+ between nodes and are never per-part cost. Job nodes add delta= (direct tokens
29
+ of the persisted chat-typed result, the accounting delta). The root footer
30
+ always prints the authoritative deduplicated_total, in both modes; --component
31
+ relabels the node numbers direct=. cache= carries the cache-hit share of
32
+ prompt (cct/pt); the root footer adds fresh= (prompt not served from cache,
33
+ including cache_write) and cache_write= (Anthropic cache-write tokens).
34
+ Per-job delta= values sum to the root total on continuation chains but not on
35
+ general DAGs.
26
36
  EOF
27
37
  if options[:help]
28
38
  defined?(scout_usage) ? scout_usage : puts(SOPT.doc)
@@ -38,6 +48,12 @@ evidence = options.delete(:evidence)
38
48
  filename = ARGV.first
39
49
  raise MissingParameterException, :filename if filename.nil?
40
50
 
51
+ # One parse-once cache for the whole CLI run (see
52
+ # Chat.provenance_chat_load): the initial traversal, the token-event
53
+ # collector and this script's own node computations then share a single
54
+ # Chat object per resolved path.
55
+ Chat.open_provenance_run_cache
56
+
41
57
  # A persisted Step always has an info sidecar. A .files directory alone is NOT
42
58
  # evidence: saved agent chats carry one too. Explicitly loading the Step after
43
59
  # this evidence check avoids treating every readable chat as a job.
@@ -103,18 +119,36 @@ chat_cache = {}
103
119
  direct_tokens = lambda do |key|
104
120
  node = nodes[key]
105
121
  if node[:kind] == :chat
106
- chat = chat_cache[node[:path]] ||= Chat.load(node[:path])
122
+ chat = chat_cache[node[:path]] ||= Chat.provenance_chat_load(node[:path])
107
123
  Chat.token_totals([chat])
108
124
  else
109
125
  logs = edges.select { |e| e[:from] == key && e[:relation] == :log }
110
126
  .collect { |e| nodes[e[:to]][:path] }.uniq
111
- chats = logs.collect { |path| chat_cache[path] ||= Chat.load(path) }
127
+ chats = logs.collect { |path| chat_cache[path] ||= Chat.provenance_chat_load(path) }
112
128
  Chat.token_totals(chats)
113
129
  end
114
130
  rescue => error
115
131
  warnings << { error: error, kind: node[:kind], object: node[:object], relation: :tokens }
116
132
  Chat::TOKEN_KEYS.each_with_object({}) { |name, totals| totals[name.to_sym] = 0 }
117
133
  end
134
+ # Per-job accounting delta: the direct tokens of the job's persisted
135
+ # chat-typed result (Chat.job_result_chat_file). This is the
136
+ # receipt-defined accounting object - the new chat the job returned - so on a
137
+ # continuation chain the per-job deltas sum to the root deduplicated_total
138
+ # while every ancestor's evidence= closure stays cumulative. Jobs whose
139
+ # result is not chat-typed omit the field (nil, never zero).
140
+ delta_cache = {}
141
+ delta_tokens = lambda do |key|
142
+ return delta_cache[key] if delta_cache.key?(key)
143
+ node = nodes[key]
144
+ delta_cache[key] = if node[:kind] == :job
145
+ file = Chat.job_result_chat_file(node[:object])
146
+ file ? Chat.token_totals([chat_cache[Chat.provenance_path(:chat, file)] ||= Chat.provenance_chat_load(file)]) : nil
147
+ end
148
+ rescue => error
149
+ warnings << { error: error, kind: node[:kind], object: node[:object], relation: :delta }
150
+ delta_cache[key] = nil
151
+ end
118
152
 
119
153
  # ------------------------------------------------------------------
120
154
  # Receipt (agent_meta) evidence: delegated inference events embedded in
@@ -204,21 +238,52 @@ end
204
238
  # ------------------------------------------------------------------
205
239
  # Formatting helpers
206
240
  # ------------------------------------------------------------------
207
- def token_str(tokens)
241
+ # `qualifier` names the total field after what it measures so a number can
242
+ # never be read as a different quantity: `evidence` (default tree; subtree
243
+ # deduplicated closure), `direct` (-c; per-component direct tokens),
244
+ # `deduplicated_total` (root footer; the authoritative cost). `events`
245
+ # appends the deduplicated event count where it is already computed (the
246
+ # footer), and `delta` inserts the job's result-chat delta right after it.
247
+ #
248
+ # Canonical field order, identical for node lines and the footer: qualifier
249
+ # total, [delta=], then the prompt axis `prompt= cache=<abs>@<rate> [fresh=]
250
+ # [cache_write=]`, then generation `cont= reason=`. `footer: true` adds the
251
+ # run-level economics (`fresh=`, `cache_write=`) that do not belong to a node
252
+ # label. The rate is always computed from RAW integer totals
253
+ # (`100.0 * cct / pt`): humanized values round at K/M boundaries and would
254
+ # corrupt boundary rates. `cache=` prints whenever pt is positive, including
255
+ # `cache=0@0.0%`; when pt is nil/0 the whole prompt axis is omitted (nothing
256
+ # to partition). A conflicting summed set can yield cct > pt; the raw ratio
257
+ # is then printed unclamped and the footer caveat line is the guard that
258
+ # labels such a figure best-effort.
259
+ def token_str(tokens, qualifier: nil, events: nil, delta: nil, footer: false)
208
260
  parts = []
209
- parts << "total=#{Misc.human_number(tokens[:tt])}" if tokens[:tt] && tokens[:tt].to_i > 0
210
- parts << "prompt=#{Misc.human_number(tokens[:pt])}" if tokens[:pt] && tokens[:pt].to_i > 0
261
+ if tokens[:tt] && tokens[:tt].to_i > 0
262
+ value = Misc.human_number(tokens[:tt])
263
+ value += " (#{events} events)" if events
264
+ parts << "#{qualifier || 'total'}=#{value}"
265
+ parts << "delta=#{Misc.human_number(delta[:tt])}" if delta && delta[:tt].to_i > 0
266
+ end
267
+ pt = tokens[:pt].to_i
268
+ if pt > 0
269
+ cct = tokens[:cct].to_i
270
+ parts << "prompt=#{Misc.human_number(pt)}"
271
+ parts << "cache=#{Misc.human_number(cct)}@#{format('%.1f', 100.0 * cct / pt)}%"
272
+ if footer
273
+ parts << "fresh=#{Misc.human_number(pt - cct)}"
274
+ parts << "cache_write=#{Misc.human_number(tokens[:cwt].to_i)}" if tokens[:cwt] && tokens[:cwt].to_i > 0
275
+ end
276
+ end
211
277
  parts << "cont=#{Misc.human_number(tokens[:ct])}" if tokens[:ct] && tokens[:ct].to_i > 0
212
- parts << "cache=#{Misc.human_number(tokens[:cct])}" if tokens[:cct] && tokens[:cct].to_i > 0
213
278
  parts << "reason=#{Misc.human_number(tokens[:rt])}" if tokens[:rt] && tokens[:rt].to_i > 0
214
279
  parts * ' '
215
280
  end
216
281
 
217
- def report_line(kind, label, tokens, offset: 0)
282
+ def report_line(kind, label, tokens, offset: 0, qualifier: nil, delta: nil, events: nil)
218
283
  color = kind == :job ? :yellow : :green
219
284
  parts = [' ' * (offset * 2)]
220
285
  parts << Log.color(color, kind.to_s)
221
- tok = token_str(tokens)
286
+ tok = token_str(tokens, qualifier: qualifier, events: events, delta: delta)
222
287
  parts << tok unless tok.empty?
223
288
  parts << Log.color(:blue, label.to_s) unless label.to_s.empty?
224
289
  parts * ' '
@@ -279,14 +344,34 @@ unless flow || dot_file || plot_file
279
344
  node = nodes[key]
280
345
  parent_node = parent_key ? nodes[parent_key] : nil
281
346
 
282
- # Skip hidden nodes entirely — don't print, don't traverse children
283
- return if hidden_node?(node, relation)
347
+ # Hidden nodes (result chats duplicating a job, top-level agent.chat
348
+ # logs) are not printed, but PROJECTION ONLY: the walk still descends
349
+ # through them so their children are attributed to the current display
350
+ # parent. This is what makes delegated children of tree-hidden carrier
351
+ # chats visible without adding, removing or rewriting any traversal
352
+ # edge: the :agent_job/:job edges keep pointing at the carrier chat.
353
+ # - no "printed" slot is consumed, so offsets stay as if the carrier
354
+ # were not there
355
+ # - children keep their own incoming relation (delegated-job label)
356
+ # - the global `printed` set still dedups nodes discovered from
357
+ # several carriers (first visit wins)
358
+ if hidden_node?(node, relation)
359
+ (adjacency[key] || []).each do |edge|
360
+ print_tree.call(edge[:to], offset, edge[:relation], parent_key)
361
+ end
362
+ return
363
+ end
284
364
 
285
365
  printed << key
286
366
  label = node_label(node, parent_node, root_key, key, long)
287
367
  # A job reached through an agent_meta receipt is a delegated producer.
288
368
  label = "delegated-job #{label}" if node[:kind] == :job && relation == :agent_job
289
- puts report_line(node[:kind], label, node_tokens.call(key), offset: offset)
369
+ # Both modes: job nodes carry their accounting delta (in -c the
370
+ # direct=/delta= contrast on one line is the question the mode answers).
371
+ delta = node[:kind] == :job ? delta_tokens.call(key) : nil
372
+ puts report_line(node[:kind], label, node_tokens.call(key), offset: offset,
373
+ qualifier: component ? 'direct' : 'evidence',
374
+ delta: delta)
290
375
 
291
376
  # One compact annotation line for the delegated usage embedded in this
292
377
  # chat's own outputs; never one line per receipt.
@@ -309,6 +394,25 @@ unless flow || dot_file || plot_file
309
394
  end
310
395
  print_tree.call(root_key, 0, nil, nil)
311
396
 
397
+ # Root footer (BOTH modes): the one authoritative cost figure, always
398
+ # printed, so -c no longer lacks a global bound. It carries the run-level
399
+ # economics (deduplicated event count, fresh=, cache_write=) that do not
400
+ # belong to a node label. It reuses the root aggregate computed for the
401
+ # root node line, which IS the root :deduplicated_total (every deduplicated
402
+ # event once). Identity conflicts demote it to best-effort; this footer
403
+ # caveat is the single home of the non-authoritative warning in both modes
404
+ # (the -c scope block no longer repeats it).
405
+ conflicts = {}
406
+ root_totals = Chat.provenance_token_totals(root, warnings: receipt_warnings, conflicts: conflicts)
407
+ puts Log.color(:cyan, 'root ') +
408
+ token_str(root_totals, qualifier: 'deduplicated_total',
409
+ events: token_events_for.call.length, footer: true) + ' ' +
410
+ Log.color(:cyan, '(authoritative cost; per-node evidence=/direct= values overlap)')
411
+ unless conflicts[:authoritative]
412
+ puts Log.color(:red, 'unresolved identity conflicts: ') +
413
+ "#{conflicts[:events]} conflicting event(s); totals above are best-effort (canonical evidence only), not exact"
414
+ end
415
+
312
416
  # Component mode: make the evidence coverage behind the numbers explicit
313
417
  # once receipts are present; chats without receipts keep the previous output
314
418
  # exactly. These are COVERAGE figures, not additive cost categories:
@@ -316,8 +420,7 @@ unless flow || dot_file || plot_file
316
420
  # both a saved child log and a receipt. Only deduplicated_total is the
317
421
  # cost, and receipt_only is the disjoint delegated contribution.
318
422
  if component && has_receipt_evidence
319
- coverage = [[:deduplicated_total, 'total'],
320
- [:chat_evidence, 'events with saved chat/log evidence'],
423
+ coverage = [[:chat_evidence, 'events with saved chat/log evidence'],
321
424
  [:receipt_evidence, 'events inside agent_meta receipts (overlaps chat_evidence)'],
322
425
  [:receipt_only, 'receipt evidence with no saved chat/log']]
323
426
  coverage.each do |scope, note|
@@ -326,13 +429,6 @@ unless flow || dot_file || plot_file
326
429
  rendered = 'total=0' if rendered.empty?
327
430
  puts Log.color(:cyan, "evidence #{scope}:") + ' ' + rendered + ' ' + Log.color(:cyan, "(#{note})")
328
431
  end
329
-
330
- conflict_info = {}
331
- Chat.provenance_token_totals(root, warnings: [], conflicts: conflict_info)
332
- unless conflict_info[:authoritative]
333
- puts Log.color(:red, 'unresolved identity conflicts: ') +
334
- "#{conflict_info[:events]} conflicting event(s); totals above are best-effort (canonical evidence only), not exact"
335
- end
336
432
  end
337
433
  end
338
434
 
@@ -600,3 +696,7 @@ receipt_warnings.each do |warning|
600
696
  Log.warn "agent_meta receipt: #{warning[:reason]} in #{warning[:source]} at #{location.inspect} call=#{warning[:call_id]}#{warning[:message] ? ' - ' + warning[:message] : ''}"
601
697
  end
602
698
  end
699
+
700
+ # End of the CLI run: drop the parse-once cache so a following run in the
701
+ # same process (tests, embedded use) re-parses freshly.
702
+ Chat.close_provenance_run_cache
@@ -32,8 +32,14 @@ module AgentMetaFixtures
32
32
  end
33
33
 
34
34
  # JSON payload of a function_call_output envelope carrying agent_meta.
35
- def receipt_output(call_id, agent_meta, name: 'ask', content: 'child answer')
36
- {name: name, content: content, id: call_id, agent_meta: agent_meta}.to_json
35
+ # `envelope:` selects which serialized key carries the receipts:
36
+ # :agent_meta is the legacy envelope, :meta the current one written since
37
+ # the dual-envelope reader landed (lib/scout/llm/tools/call.rb).
38
+ # `name` only decorates the envelope; the receipt lifting ignores it.
39
+ def receipt_output(call_id, agent_meta, name: 'ask', content: 'child answer', envelope: :agent_meta)
40
+ payload = {name: name, content: content, id: call_id}
41
+ payload[envelope] = agent_meta
42
+ payload.to_json
37
43
  end
38
44
 
39
45
  # Persisted chat text with one paired ask call per receipt entry. Hash keys
@@ -43,11 +49,13 @@ module AgentMetaFixtures
43
49
  # Message indexes produced by Chat.parse (single user turn, no leading
44
50
  # empty user message since 49c0d20):
45
51
  # 0 user, then per receipt: function_call, function_call_output.
46
- def receipt_chat_text(receipts, extra: nil)
52
+ # `tool` sets the function_call name; receipts are lifted from the output
53
+ # envelope regardless of it, so tests can pin that no tool name is special.
54
+ def receipt_chat_text(receipts, extra: nil, envelope: :agent_meta, tool: 'ask')
47
55
  lines = ['user: Run the worker']
48
56
  receipts.each do |call_id, agent_meta|
49
- lines << 'function_call: ' + %({"name":"ask","arguments":{},"id":"#{call_id}"})
50
- lines << 'function_call_output: ' + receipt_output(call_id, agent_meta)
57
+ lines << 'function_call: ' + %({"name":"#{tool}","arguments":{},"id":"#{call_id}"})
58
+ lines << 'function_call_output: ' + receipt_output(call_id, agent_meta, envelope: envelope)
51
59
  end
52
60
  lines.concat(Array(extra)) if extra
53
61
  lines << 'assistant: done'
@@ -128,4 +136,82 @@ module AgentMetaFixtures
128
136
  "function_call_output: #{payload}\n" +
129
137
  "assistant: done\n"
130
138
  end
139
+
140
+ # Continuation-chain fixture (theme-2 A/B/C shape, in-repo). Three
141
+ # chat-typed jobs A -> B -> C whose agent logs are cumulative histories
142
+ # (each carries the projected parent metas before its own), so per-root
143
+ # closures are cumulative while each job's own delta is disjoint. Every
144
+ # meta line is a plain direct chat meta (not a receipt).
145
+ #
146
+ # Token plan (all figures arithmetically checkable):
147
+ # A own: 3 metas pt=40 tt=50 cct=30 cwt=5 -> delta tt=150, pt=120, cct=90
148
+ # B own: 2 metas pt=100 tt=120 (no cache) -> delta tt=240, pt=200
149
+ # C own: 1 meta pt=0 tt=600 (pt missing)-> delta tt=600, pt=0
150
+ # chain closure at C: 6 events, tt=990, pt=320, cct=90, cwt=15, ct=6
151
+ # sum of deltas 150+240+600 = 990 == root deduplicated_total
152
+ #
153
+ # Also builds a TSV-typed job reached from `parent.chat` through a
154
+ # receipt, to pin delta= omission for non-chat results in both modes.
155
+ # Returns [a_job, b_job, c_job, tsv_job, parent_chat].
156
+ def continuation_chain(dir)
157
+ meta = lambda do |id, pt, tt, cct: 0, cwt: 0|
158
+ parts = ["pt=#{pt}", 'ct=1', "tt=#{tt}"]
159
+ parts << "cct=#{cct}" if cct > 0
160
+ parts << "cwt=#{cwt}" if cwt > 0
161
+ parts << "inference_id=#{id}" << 'timestamp=2026-09-04T00:00:00Z'
162
+ "meta: " + parts * ' '
163
+ end
164
+
165
+ a_metas = (1..3).collect { |i| meta.call("a#{i}", 40, 50, cct: 30, cwt: 5) }
166
+ b_metas = (1..2).collect { |i| meta.call("b#{i}", 100, 120) }
167
+ c_meta = meta.call('c1', 0, 600)
168
+
169
+ a = make_job(dir, 'Cortex/continue/Default_a1', logs: {'agent.chat' => ''})
170
+ File.write(a + '.info', {dependencies: [], type: :chat}.to_json)
171
+ File.write(a, (['user: start'] + a_metas + ["meta: job=#{a}"]) * "
172
+ " + "
173
+ ")
174
+ File.write(File.join(a + '.files', 'log', 'agent.chat'), a_metas * "
175
+ " + "
176
+ ")
177
+
178
+ b = make_job(dir, 'Cortex/continue/Default_b2', logs: {'agent.chat' => ''})
179
+ File.write(b + '.info', {dependencies: [], type: :chat}.to_json)
180
+ File.write(b, (['user: continue B'] + b_metas + ["meta: job=#{b}"]) * "
181
+ " + "
182
+ ")
183
+ File.write(File.join(b + '.files', 'log', 'agent.chat'),
184
+ (["meta: job=#{a}"] + a_metas + b_metas) * "
185
+ " + "
186
+ ")
187
+
188
+ c = make_job(dir, 'Cortex/continue/Default_c3', logs: {'agent.chat' => ''})
189
+ File.write(c + '.info', {dependencies: [], type: :chat}.to_json)
190
+ File.write(c, (['user: continue C'] + [c_meta, "meta: job=#{c}"]) * "
191
+ " + "
192
+ ")
193
+ File.write(File.join(c + '.files', 'log', 'agent.chat'),
194
+ (["meta: job=#{b}"] + a_metas + b_metas + [c_meta]) * "
195
+ " + "
196
+ ")
197
+
198
+ # TSV-typed job (no chat result): must omit delta= while still showing
199
+ # its own log evidence.
200
+ tsv = make_job(dir, 'Other/step/Default_d4',
201
+ result: "a b
202
+ 1 2
203
+ ",
204
+ logs: {'agent.chat' => meta.call('d1', 7, 8) + "
205
+ "})
206
+ info = JSON.parse(File.read(tsv + '.info'))
207
+ File.write(tsv + '.info', {dependencies: info['dependencies'], type: :tsv}.to_json)
208
+
209
+ parent = write_chat(dir, 'parent.chat',
210
+ receipt_chat_text(
211
+ {'a1' => [meta_receipt("job=#{tsv}")]},
212
+ extra: ['meta: pt=10 ct=1 tt=11 inference_id=p1']
213
+ ))
214
+
215
+ [a, b, c, tsv, parent]
216
+ end
131
217
  end