log-foundry 0.10.2.dev67__tar.gz → 0.10.2.dev68__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (59) hide show
  1. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/PKG-INFO +41 -10
  2. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/README.md +40 -9
  3. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/pyproject.toml +1 -1
  4. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/__init__.py +9 -1
  5. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/api.py +28 -2
  6. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/decorator.py +103 -4
  7. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/model.py +7 -0
  8. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/sinks/multi.py +36 -8
  9. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/worker.py +19 -0
  10. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/LICENSE +0 -0
  11. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/_diag.py +0 -0
  12. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/_fork.py +0 -0
  13. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/_lifecycle.py +0 -0
  14. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/config.py +0 -0
  15. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/console.py +0 -0
  16. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/context.py +0 -0
  17. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/ids.py +0 -0
  18. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/py.typed +0 -0
  19. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/results.py +0 -0
  20. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/sanitize.py +0 -0
  21. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/sinks/__init__.py +0 -0
  22. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/sinks/_batch.py +0 -0
  23. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/sinks/_chunk.py +0 -0
  24. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/sinks/_retry.py +0 -0
  25. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/sinks/_socket.py +0 -0
  26. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/sinks/_time.py +0 -0
  27. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/sinks/base.py +0 -0
  28. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/sinks/callback.py +0 -0
  29. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/sinks/clickhouse.py +0 -0
  30. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/sinks/datadog.py +0 -0
  31. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/sinks/elasticsearch.py +0 -0
  32. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/sinks/eventhubs.py +0 -0
  33. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/sinks/file.py +0 -0
  34. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/sinks/filtering.py +0 -0
  35. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/sinks/firehose.py +0 -0
  36. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/sinks/honeycomb.py +0 -0
  37. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/sinks/http.py +0 -0
  38. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/sinks/kafka.py +0 -0
  39. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/sinks/kinesis.py +0 -0
  40. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/sinks/logging_sink.py +0 -0
  41. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/sinks/logstash.py +0 -0
  42. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/sinks/loki.py +0 -0
  43. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/sinks/memory.py +0 -0
  44. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/sinks/mongodb.py +0 -0
  45. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/sinks/nats.py +0 -0
  46. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/sinks/newrelic.py +0 -0
  47. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/sinks/null.py +0 -0
  48. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/sinks/postgres.py +0 -0
  49. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/sinks/pubsub.py +0 -0
  50. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/sinks/rabbitmq.py +0 -0
  51. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/sinks/redis.py +0 -0
  52. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/sinks/sentry.py +0 -0
  53. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/sinks/sns.py +0 -0
  54. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/sinks/splunk.py +0 -0
  55. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/sinks/sqlite.py +0 -0
  56. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/sinks/sqs.py +0 -0
  57. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/sinks/stdout.py +0 -0
  58. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/sinks/syslog.py +0 -0
  59. {log_foundry-0.10.2.dev67 → log_foundry-0.10.2.dev68}/src/log_foundry/sinks/transform.py +0 -0
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: log-foundry
3
- Version: 0.10.2.dev67
3
+ Version: 0.10.2.dev68
4
4
  Summary: Generate logs for your console and JSON events for downstream consumption.
5
5
  License-Expression: MIT
6
6
  License-File: LICENSE
@@ -413,15 +413,23 @@ def enqueue_check(location: str) -> None:
413
413
 
414
414
  ```python
415
415
  @lf.trace
416
- def handler(event, context):
416
+ def _handler(event, context):
417
417
  lf.continue_trace(event.get("traceparent"), baggage=event.get("baggage"))
418
418
  lf.info("inspecting") # same trace_id as the producer; parent is its span
419
+ return inspect(event)
420
+
421
+
422
+ def handler(event, context): # the entry point, deliberately not decorated
419
423
  try:
420
- return inspect(event)
424
+ return _handler(event, context)
421
425
  finally:
422
- lf.flush()
426
+ lf.flush() # the span has closed, so its events are drained
423
427
  ```
424
428
 
429
+ `flush()` goes **outside** the traced function, not in its `finally`. An in-span event lives on the
430
+ span until the span *closes*, and `flush()` drains the queue — so a `flush()` inside the span has
431
+ nothing to drain yet.
432
+
425
433
  | Call | Does |
426
434
  |---|---|
427
435
  | `continue_trace(traceparent=None, *, trace_id=None, parent_span_id=None, baggage=None)` | Adopt an inbound context. Returns a `ContinueResult`: truthy if adopted, else falsy with `reason` of `"nothing-supplied"` or `"rejected"`. The verdict is about the **trace context** — `baggage=` is merged independently and does not make it truthy. Never raises. |
@@ -953,6 +961,7 @@ if (
953
961
  h.dropped or h.failed_batches or h.stopped_reason or h.incomplete_swaps
954
962
  or (h.sink and (h.sink.dropped or h.sink.failed))
955
963
  or (h.retired and h.submitted_after_shutdown)
964
+ or h.orphan_lost or h.in_span_lost
956
965
  ):
957
966
  ... # logs were silently lost — worth an alert
958
967
  ```
@@ -974,8 +983,14 @@ They tell you different things, and they want different responses:
974
983
  | `retired` + `submitted_after_shutdown` | `shutdown()` was called and the process **kept logging**. Those events are queued where nothing will drain them — total loss, for as long as the process runs. | Use `flush()`, not `shutdown()`, in a process that logs again. This is the serverless mistake below. |
975
984
  | `incomplete_swaps` | A late `configure(sink=...)` could not confirm the previous sink was drained. The swap took effect; that sink was left open and some queued events may have gone to the new one. | Investigate the previous sink — it was hung or failing. Configure the sink before the first log where you can. |
976
985
  | `inherited_sink` | This process is delivering to a sink it **inherited across a `fork`** and may not release, so it will not be closed here. Not a loss and not an alert term. | Nothing, usually. It explains a handle still open after `shutdown()`, and tells you a deployment shares one sink across a fork at all. `True` for a shared `StdoutSink` too, whose `close()` only flushes — so a `True` is not by itself evidence that anything is held. If you want the child to own its transport, build the sink in the worker process (see Forking). |
986
+ | `orphan_lost` | An event logged **with no active span** never reached the sink. That call emits on your own thread with no worker behind it, so no other field here can carry it — it is not a batch, there was no retry, and there may be no worker at all. Covers a sink that failed to *construct* as well as one that raised. | Fix the destination, or the data. The stderr line names the exception type. If a process logs this way at all, this is the field to alert on: nothing else describes that path. |
987
+ | `in_span_lost` | An event logged **inside a span** could not be built — a value that could not be turned into an event. Always the data, never the destination: the in-span path cannot fail at delivery, which is `failed_batches`. | Fix the call site. Passing a non-string message (an exception object, say) is the common cause. |
977
988
  | `closing_sinks` | Swapped-out sinks inside `close()` **right now** — a live gauge, not a counter, and the only field that falls as well as rises. Non-zero on a single read is normal during a swap. | Nothing, unless it stays non-zero. That means a destination is stuck in `close()` and will not release its resources. |
978
989
 
990
+ `orphan_lost` and `in_span_lost` are deliberately two fields and their sum is deliberately not
991
+ reported. They aggregate different failure populations — one can mean the destination *or* the
992
+ data, the other can only mean the data — so a single number would hide which fix applies.
993
+
979
994
  `h.sink` is a `SinkLosses(dropped, failed)` or `None` — `None` when no worker exists yet, or when
980
995
  the configured sink reports nothing (`losses()` is optional). Note the two `dropped` fields count
981
996
  different things: the worker's is backpressure at *its* queue, the sink's is an event that never
@@ -996,14 +1011,21 @@ logged, and after a clean `shutdown()`, so a plain truthiness check is safe. Wit
996
1011
  thread showed up only indirectly, as `dropped` climbing once the queue filled — the wrong signal,
997
1012
  pointing at the wrong fix.
998
1013
 
1014
+ `orphan_lost` climbing is on its own a reason to look: unlike `dropped`, it is never
1015
+ backpressure and never transient. Each increment is one event that reached no destination, and on
1016
+ a process that logs only outside a span it is the *only* field that can say so.
1017
+
999
1018
  `retired` is deliberately **not** alerted on by itself. A process that shuts down and then stops
1000
1019
  logging is doing the right thing; it is the *pair* — retired, and still being handed events — that
1001
1020
  means every log line since the shutdown has gone nowhere. That state used to read as perfectly
1002
1021
  healthy: `stopped_reason` is `None` after a clean shutdown, and the queue simply grows.
1003
1022
 
1004
- `retired` is also the one field reported for a process that has **no worker at all**. A program
1005
- that only ever calls `info()`/`error()` outside a span emits synchronously and builds no background
1006
- worker, so every other field describes something that does not exist and reads zero. Its
1023
+ `retired`, `orphan_lost` and `in_span_lost` are the fields reported for a process that has **no
1024
+ worker at all**. A program that only ever calls `info()`/`error()` outside a span emits
1025
+ synchronously and builds no background worker, so the rest describe something that does not exist
1026
+ and read zero — which is why that path needs counters of its own. Until it had them, such a process
1027
+ reported `queued=0 dropped=0 failed_batches=0 stopped_reason=None` over total, permanent loss, and
1028
+ the only thing that said otherwise was a line on stderr. Its
1007
1029
  `shutdown()` still closes the sink, exactly once and without starting a thread, and `retired` reads
1008
1030
  `True` afterwards rather than staying vacuously `False`. `submitted_after_shutdown` stays `0` there
1009
1031
  by design: a later level call is *refused* at the closed sink and announced on stderr — if the sink
@@ -1124,19 +1146,28 @@ from log_foundry.sinks.sqs import SQSSink
1124
1146
  lf.configure(service="billing-api", env="prod", sink=SQSSink(queue_url=QUEUE_URL))
1125
1147
 
1126
1148
  @lf.trace
1127
- def handler(event, context):
1149
+ def _handler(event, context):
1128
1150
  lf.info("received", records=len(event["Records"]))
1151
+ return do_work(event)
1152
+
1153
+
1154
+ def handler(event, context):
1155
+ # NOT decorated, so the span closes when `_handler` returns and its events reach the queue
1156
+ # before `flush()` runs. A `flush()` *inside* the traced function has nothing to drain yet.
1129
1157
  try:
1130
- return do_work(event)
1158
+ return _handler(event, context)
1131
1159
  finally:
1132
1160
  # In `finally`: the failed invocation is the one worth logging. NEVER shutdown() here —
1133
1161
  # the worker does not come back, and every later invocation on this warm container
1134
1162
  # would log nothing.
1135
1163
  drained = lf.flush()
1136
1164
  h = lf.health()
1137
- if not drained or h.failed_batches or h.dropped or h.stopped_reason or h.retired:
1165
+ if (not drained or h.failed_batches or h.dropped or h.stopped_reason or h.retired
1166
+ or h.orphan_lost or h.in_span_lost):
1138
1167
  # `drained` covers this invocation's tail; the counters cover anything the worker
1139
1168
  # lost earlier — a batch its own interval trigger already gave up on, for instance.
1169
+ # `orphan_lost`/`in_span_lost` cover the two paths no worker field can describe: a
1170
+ # log made outside any span, and an event that could not be built.
1140
1171
  # `h.retired` catches the mistake above: inside a handler it can only mean something
1141
1172
  # called shutdown(), and from here on this container logs nothing.
1142
1173
  # Emitting this through your platform's own logger keeps it outside the pipeline
@@ -377,15 +377,23 @@ def enqueue_check(location: str) -> None:
377
377
 
378
378
  ```python
379
379
  @lf.trace
380
- def handler(event, context):
380
+ def _handler(event, context):
381
381
  lf.continue_trace(event.get("traceparent"), baggage=event.get("baggage"))
382
382
  lf.info("inspecting") # same trace_id as the producer; parent is its span
383
+ return inspect(event)
384
+
385
+
386
+ def handler(event, context): # the entry point, deliberately not decorated
383
387
  try:
384
- return inspect(event)
388
+ return _handler(event, context)
385
389
  finally:
386
- lf.flush()
390
+ lf.flush() # the span has closed, so its events are drained
387
391
  ```
388
392
 
393
+ `flush()` goes **outside** the traced function, not in its `finally`. An in-span event lives on the
394
+ span until the span *closes*, and `flush()` drains the queue — so a `flush()` inside the span has
395
+ nothing to drain yet.
396
+
389
397
  | Call | Does |
390
398
  |---|---|
391
399
  | `continue_trace(traceparent=None, *, trace_id=None, parent_span_id=None, baggage=None)` | Adopt an inbound context. Returns a `ContinueResult`: truthy if adopted, else falsy with `reason` of `"nothing-supplied"` or `"rejected"`. The verdict is about the **trace context** — `baggage=` is merged independently and does not make it truthy. Never raises. |
@@ -917,6 +925,7 @@ if (
917
925
  h.dropped or h.failed_batches or h.stopped_reason or h.incomplete_swaps
918
926
  or (h.sink and (h.sink.dropped or h.sink.failed))
919
927
  or (h.retired and h.submitted_after_shutdown)
928
+ or h.orphan_lost or h.in_span_lost
920
929
  ):
921
930
  ... # logs were silently lost — worth an alert
922
931
  ```
@@ -938,8 +947,14 @@ They tell you different things, and they want different responses:
938
947
  | `retired` + `submitted_after_shutdown` | `shutdown()` was called and the process **kept logging**. Those events are queued where nothing will drain them — total loss, for as long as the process runs. | Use `flush()`, not `shutdown()`, in a process that logs again. This is the serverless mistake below. |
939
948
  | `incomplete_swaps` | A late `configure(sink=...)` could not confirm the previous sink was drained. The swap took effect; that sink was left open and some queued events may have gone to the new one. | Investigate the previous sink — it was hung or failing. Configure the sink before the first log where you can. |
940
949
  | `inherited_sink` | This process is delivering to a sink it **inherited across a `fork`** and may not release, so it will not be closed here. Not a loss and not an alert term. | Nothing, usually. It explains a handle still open after `shutdown()`, and tells you a deployment shares one sink across a fork at all. `True` for a shared `StdoutSink` too, whose `close()` only flushes — so a `True` is not by itself evidence that anything is held. If you want the child to own its transport, build the sink in the worker process (see Forking). |
950
+ | `orphan_lost` | An event logged **with no active span** never reached the sink. That call emits on your own thread with no worker behind it, so no other field here can carry it — it is not a batch, there was no retry, and there may be no worker at all. Covers a sink that failed to *construct* as well as one that raised. | Fix the destination, or the data. The stderr line names the exception type. If a process logs this way at all, this is the field to alert on: nothing else describes that path. |
951
+ | `in_span_lost` | An event logged **inside a span** could not be built — a value that could not be turned into an event. Always the data, never the destination: the in-span path cannot fail at delivery, which is `failed_batches`. | Fix the call site. Passing a non-string message (an exception object, say) is the common cause. |
941
952
  | `closing_sinks` | Swapped-out sinks inside `close()` **right now** — a live gauge, not a counter, and the only field that falls as well as rises. Non-zero on a single read is normal during a swap. | Nothing, unless it stays non-zero. That means a destination is stuck in `close()` and will not release its resources. |
942
953
 
954
+ `orphan_lost` and `in_span_lost` are deliberately two fields and their sum is deliberately not
955
+ reported. They aggregate different failure populations — one can mean the destination *or* the
956
+ data, the other can only mean the data — so a single number would hide which fix applies.
957
+
943
958
  `h.sink` is a `SinkLosses(dropped, failed)` or `None` — `None` when no worker exists yet, or when
944
959
  the configured sink reports nothing (`losses()` is optional). Note the two `dropped` fields count
945
960
  different things: the worker's is backpressure at *its* queue, the sink's is an event that never
@@ -960,14 +975,21 @@ logged, and after a clean `shutdown()`, so a plain truthiness check is safe. Wit
960
975
  thread showed up only indirectly, as `dropped` climbing once the queue filled — the wrong signal,
961
976
  pointing at the wrong fix.
962
977
 
978
+ `orphan_lost` climbing is on its own a reason to look: unlike `dropped`, it is never
979
+ backpressure and never transient. Each increment is one event that reached no destination, and on
980
+ a process that logs only outside a span it is the *only* field that can say so.
981
+
963
982
  `retired` is deliberately **not** alerted on by itself. A process that shuts down and then stops
964
983
  logging is doing the right thing; it is the *pair* — retired, and still being handed events — that
965
984
  means every log line since the shutdown has gone nowhere. That state used to read as perfectly
966
985
  healthy: `stopped_reason` is `None` after a clean shutdown, and the queue simply grows.
967
986
 
968
- `retired` is also the one field reported for a process that has **no worker at all**. A program
969
- that only ever calls `info()`/`error()` outside a span emits synchronously and builds no background
970
- worker, so every other field describes something that does not exist and reads zero. Its
987
+ `retired`, `orphan_lost` and `in_span_lost` are the fields reported for a process that has **no
988
+ worker at all**. A program that only ever calls `info()`/`error()` outside a span emits
989
+ synchronously and builds no background worker, so the rest describe something that does not exist
990
+ and read zero — which is why that path needs counters of its own. Until it had them, such a process
991
+ reported `queued=0 dropped=0 failed_batches=0 stopped_reason=None` over total, permanent loss, and
992
+ the only thing that said otherwise was a line on stderr. Its
971
993
  `shutdown()` still closes the sink, exactly once and without starting a thread, and `retired` reads
972
994
  `True` afterwards rather than staying vacuously `False`. `submitted_after_shutdown` stays `0` there
973
995
  by design: a later level call is *refused* at the closed sink and announced on stderr — if the sink
@@ -1088,19 +1110,28 @@ from log_foundry.sinks.sqs import SQSSink
1088
1110
  lf.configure(service="billing-api", env="prod", sink=SQSSink(queue_url=QUEUE_URL))
1089
1111
 
1090
1112
  @lf.trace
1091
- def handler(event, context):
1113
+ def _handler(event, context):
1092
1114
  lf.info("received", records=len(event["Records"]))
1115
+ return do_work(event)
1116
+
1117
+
1118
+ def handler(event, context):
1119
+ # NOT decorated, so the span closes when `_handler` returns and its events reach the queue
1120
+ # before `flush()` runs. A `flush()` *inside* the traced function has nothing to drain yet.
1093
1121
  try:
1094
- return do_work(event)
1122
+ return _handler(event, context)
1095
1123
  finally:
1096
1124
  # In `finally`: the failed invocation is the one worth logging. NEVER shutdown() here —
1097
1125
  # the worker does not come back, and every later invocation on this warm container
1098
1126
  # would log nothing.
1099
1127
  drained = lf.flush()
1100
1128
  h = lf.health()
1101
- if not drained or h.failed_batches or h.dropped or h.stopped_reason or h.retired:
1129
+ if (not drained or h.failed_batches or h.dropped or h.stopped_reason or h.retired
1130
+ or h.orphan_lost or h.in_span_lost):
1102
1131
  # `drained` covers this invocation's tail; the counters cover anything the worker
1103
1132
  # lost earlier — a batch its own interval trigger already gave up on, for instance.
1133
+ # `orphan_lost`/`in_span_lost` cover the two paths no worker field can describe: a
1134
+ # log made outside any span, and an event that could not be built.
1104
1135
  # `h.retired` catches the mistake above: inside a handler it can only mean something
1105
1136
  # called shutdown(), and from here on this container logs nothing.
1106
1137
  # Emitting this through your platform's own logger keeps it outside the pipeline
@@ -20,7 +20,7 @@ dependencies = [
20
20
  ]
21
21
 
22
22
  # Optional features. Install with: pip install log-foundry[aws]
23
- version = "0.10.2.dev67"
23
+ version = "0.10.2.dev68"
24
24
 
25
25
  [project.optional-dependencies]
26
26
  aws = ["boto3>=1.43.61"] # SQSSink, SNSSink, KinesisSink, FirehoseSink
@@ -72,9 +72,17 @@ def health() -> Health:
72
72
  h = log_foundry.health()
73
73
  if h.dropped or h.failed_batches or h.stopped_reason or h.incomplete_swaps or (
74
74
  h.sink and (h.sink.dropped or h.sink.failed)
75
- ) or (h.retired and h.submitted_after_shutdown):
75
+ ) or (h.retired and h.submitted_after_shutdown) or h.orphan_lost or h.in_span_lost:
76
76
  ... # raise an alert; logs were silently lost
77
77
 
78
+ ``orphan_lost`` and ``in_span_lost`` are the two terms that do **not** describe the worker
79
+ (SPEC-036 FR-003). A level call made with no active span emits on the caller's own thread, and
80
+ an event that cannot be *built* never reaches a queue at all — so every other field here
81
+ describes machinery those two losses never touched, and a process that only logs outside a
82
+ span read all zeros over total loss until they existed. They stay separate because one can
83
+ mean the destination or the data and the other can only mean the data; their sum is a number
84
+ nobody can act on.
85
+
78
86
  ``retired`` alone is not a fault — a process that shuts down and then stops logging is
79
87
  doing the right thing, which is why it is paired with the count rather than alerted on.
80
88
  ``closing_sinks`` is deliberately absent for the same kind of reason: it is briefly non-zero
@@ -8,7 +8,7 @@ from log_foundry import _diag, context
8
8
  from log_foundry.config import _ensure_sink
9
9
  from log_foundry.console import ConsoleWriter
10
10
  from log_foundry.context import set_baggage
11
- from log_foundry.decorator import _note_orphan_emit
11
+ from log_foundry.decorator import _note_in_span_loss, _note_orphan_emit, _note_orphan_loss
12
12
  from log_foundry.ids import new_span_id, new_trace_id
13
13
  from log_foundry.model import Span, build_event
14
14
 
@@ -71,6 +71,30 @@ def _log(level: str, message: str, echo: bool, fields: dict[str, object]) -> Non
71
71
  A guard that reads "this cannot fail" is a guard that stops being true when the code under
72
72
  it changes, and this one had already stopped.
73
73
 
74
+ **An append to a span that has already closed takes the orphan route** (SPEC-036 FR-004).
75
+ ``contextvars`` copies the same ``Span`` object into every task created inside a span, so a
76
+ fire-and-forget ``create_task`` can log after its parent returned — and since ``_flush``
77
+ detaches at submit, that append lands in a buffer nothing will ever emit. This is the only
78
+ place that can notice, because nothing in the library reads a span again after it closes; a
79
+ post-hoc check of the buffer has no observer. A span that has closed is not a span, so the
80
+ event becomes a fresh one-event span and inherits the orphan path's accounting entire —
81
+ delivered on success, ``orphan_lost`` on failure. The flag is read before the event is built
82
+ and the append happens after, so a thread *sharing* the span can still slip between the two and
83
+ strand its event in the post-swap list; that window is recorded in ``architecture.md`` §13
84
+ rather than closed, since closing it costs a per-span lock on the hottest path. That is why FR-004 adds no ``Health`` field
85
+ of its own, and why the destination is named here rather than left implied.
86
+
87
+ The cost is stated rather than hidden: the event gets a **fresh** ``trace_id``, so it leaves
88
+ its trace, not merely its span. ``contextvars`` cannot deliver the alternative once the span
89
+ is gone, and the choice is against losing the event outright.
90
+
91
+ **Both losses are counted, and counted apart** (SPEC-036 FR-003). Each guard records before it
92
+ announces, so the stderr line SPEC-025 already writes is unchanged and this adds a counter
93
+ rather than a second announcement. The orphan increment sits in the ``except`` rather than
94
+ after ``sink.emit``, which is what makes a sink that fails to *construct* count too. They stay
95
+ two fields because the populations differ: this branch can fail at the destination, the
96
+ in-span branch can only fail at the data.
97
+
74
98
  The echo runs after the emit, so a closed or redirected stream never costs the event itself.
75
99
 
76
100
  Args:
@@ -89,11 +113,12 @@ def _log(level: str, message: str, echo: bool, fields: dict[str, object]) -> Non
89
113
  baggage = context._live_baggage()
90
114
  span = context.current_span()
91
115
  event: dict[str, object] | None = None
92
- if span is not None:
116
+ if span is not None and not span.closed:
93
117
  try:
94
118
  event = build_event(span, level, message, fields=fields, baggage=baggage)
95
119
  span.events.append(event)
96
120
  except Exception as exc:
121
+ _note_in_span_loss()
97
122
  _diag.absorbed("building an in-span log", exc, "the event was lost")
98
123
  else:
99
124
  try:
@@ -109,6 +134,7 @@ def _log(level: str, message: str, echo: bool, fields: dict[str, object]) -> Non
109
134
  _note_orphan_emit(sink)
110
135
  sink.emit([event])
111
136
  except Exception as exc:
137
+ _note_orphan_loss()
112
138
  _diag.absorbed("emitting an orphan log", exc, "the event was lost")
113
139
  if echo and event is not None:
114
140
  try:
@@ -60,6 +60,19 @@ shutdown's event has every backoff collapsed to zero.
60
60
  """
61
61
  _orphan_retired = False
62
62
 
63
+ _loss_lock = threading.Lock()
64
+ """Guards the two loss counters, and deliberately not ``_worker_lock`` (SPEC-036 FR-003 AC-5).
65
+
66
+ SPEC-028's ordering rule: a counter takes its own lock, because the orphan path runs on arbitrary
67
+ application threads and ``_worker_lock`` is held across ``Worker(_ensure_sink())`` in
68
+ :func:`_get_worker` — a blocking build a counter increment must never queue behind. It cannot
69
+ deadlock here either: the increment sits in ``api._log``'s ``except``, where
70
+ :func:`_note_orphan_emit` has already released ``_worker_lock`` and the propagating exception has
71
+ released any sink lock.
72
+ """
73
+ _orphan_lost = 0
74
+ _in_span_lost = 0
75
+
63
76
  F = TypeVar("F", bound=Callable[..., Any])
64
77
 
65
78
 
@@ -728,6 +741,64 @@ def _flush_worker(timeout: float | None = 5.0) -> FlushResult:
728
741
  return FlushResult(ok=False, reason="thread-died")
729
742
 
730
743
 
744
+ def _note_orphan_loss() -> None:
745
+ """Counts one event lost on the synchronous path (SPEC-036 FR-003).
746
+
747
+ Called from ``api._log``'s orphan ``except``, which wraps the span construction, the event
748
+ build, ``_ensure_sink`` and the emit — so a sink that failed to *construct* is counted here
749
+ too, which an increment placed after ``sink.emit`` would miss.
750
+
751
+ Args:
752
+ None.
753
+
754
+ Returns:
755
+ None.
756
+
757
+ Raises:
758
+ None.
759
+ """
760
+ global _orphan_lost
761
+ with _loss_lock:
762
+ _orphan_lost += 1
763
+
764
+
765
+ def _note_in_span_loss() -> None:
766
+ """Counts one event lost while being built inside a span (SPEC-036 FR-003).
767
+
768
+ Separate from :func:`_note_orphan_loss` because the two aggregate different failure
769
+ populations: this path cannot fail at ``emit``, so a non-zero count here always means the
770
+ data, never the destination.
771
+
772
+ Args:
773
+ None.
774
+
775
+ Returns:
776
+ None.
777
+
778
+ Raises:
779
+ None.
780
+ """
781
+ global _in_span_lost
782
+ with _loss_lock:
783
+ _in_span_lost += 1
784
+
785
+
786
+ def _read_losses() -> tuple[int, int]:
787
+ """Reads both loss counters under the lock they are written under.
788
+
789
+ Args:
790
+ None.
791
+
792
+ Returns:
793
+ The orphan and in-span loss counts, in that order.
794
+
795
+ Raises:
796
+ None.
797
+ """
798
+ with _loss_lock:
799
+ return _orphan_lost, _in_span_lost
800
+
801
+
731
802
  def _delivering_to_an_inherited_sink() -> bool:
732
803
  """Whether the sink this process last installed for delivery is one it may not release.
733
804
 
@@ -786,6 +857,12 @@ def _worker_health() -> Health:
786
857
  defines that count as submissions queued where nothing will drain them, and a later orphan
787
858
  log is refused at the closed sink and announced instead. The two are not the same claim.
788
859
 
860
+ The two loss counters are synthesized on **both** branches, for the reason ``retired`` is on
861
+ one: they describe the caller's own path, not the worker's, and ``Worker`` cannot report them
862
+ because it does not know they exist — ``worker.py`` imports nothing from this module, and the
863
+ reverse read would be a cycle. A process that only ever logged outside a span has no worker
864
+ and is exactly the process whose loss they exist to show (SPEC-036 FR-003 AC-7).
865
+
789
866
  The synthesis also survives a worker built *after* that shutdown, which is why it is an
790
867
  ``or`` rather than a fallback. An orphan-only ``shutdown()`` leaves ``_worker`` unset, so a
791
868
  later ``@trace`` constructs a fresh worker whose own ``retired`` is ``False`` — and reading
@@ -806,6 +883,7 @@ def _worker_health() -> Health:
806
883
  None.
807
884
  """
808
885
  worker = _worker
886
+ orphan_lost, in_span_lost = _read_losses()
809
887
  if worker is None:
810
888
  return Health(
811
889
  queued=0,
@@ -814,11 +892,14 @@ def _worker_health() -> Health:
814
892
  retired=_orphan_retired,
815
893
  closing_sinks=_lifecycle.closing_count(),
816
894
  inherited_sink=_delivering_to_an_inherited_sink(),
895
+ orphan_lost=orphan_lost,
896
+ in_span_lost=in_span_lost,
817
897
  )
818
898
  health = worker.health()
819
- if _orphan_retired and not health.retired:
820
- return replace(health, retired=True)
821
- return health
899
+ retired = _orphan_retired or health.retired
900
+ return replace(
901
+ health, retired=retired, orphan_lost=orphan_lost, in_span_lost=in_span_lost
902
+ )
822
903
 
823
904
 
824
905
  def _flush(span: Span) -> None:
@@ -828,6 +909,16 @@ def _flush(span: Span) -> None:
828
909
  ``_ensure_sink``, so a zero-config ``@trace`` still falls back to ``StdoutSink`` rather
829
910
  than crashing.
830
911
 
912
+ **The buffer is detached, not handed over** (SPEC-036 FR-004). ``submit`` used to take the
913
+ live list, so a task outliving its parent span appended to a list the worker already owned:
914
+ under the flush interval the event was delivered but ordered after ``span.end``, and over it
915
+ silently lost — a pure race on a timer. The swap gives the worker the old list and leaves the
916
+ span a fresh one, so the outcome stops depending on timing. A copy would do the same and
917
+ allocates a second list per span on the hottest path in the library; the swap is free.
918
+
919
+ The late append is now landing in a buffer nothing will emit, which is *also* loss — that
920
+ half is ``api._log``'s, keyed on :attr:`Span.closed`.
921
+
831
922
  Args:
832
923
  span: The finished span whose buffered events are submitted.
833
924
 
@@ -837,7 +928,8 @@ def _flush(span: Span) -> None:
837
928
  Raises:
838
929
  Exception: Whatever creating the worker or submitting raises; :func:`_end` is the guard.
839
930
  """
840
- _get_worker().submit(span.events)
931
+ events, span.events = span.events, []
932
+ _get_worker().submit(events)
841
933
 
842
934
 
843
935
  type _SpanScope = tuple[
@@ -947,6 +1039,12 @@ def _close_span(span: Span, status: str, exc: BaseException | None) -> None:
947
1039
  baggage known at their construction, empty for ``span.start``, so SPEC-015 completes them
948
1040
  from the baggage live now, while the batch is still ours.
949
1041
 
1042
+ ``closed`` is set **before** the flush, not after (SPEC-036 FR-004). ``_flush`` can raise —
1043
+ building the worker, or the queue path — and :func:`_end`'s guard absorbs it, so setting the
1044
+ flag afterwards would leave it ``False`` on exactly the runs where a later append most needs
1045
+ the orphan route. ``sinks/base.py`` settles the general form of this: set it before releasing
1046
+ anything.
1047
+
950
1048
  Args:
951
1049
  span: The span being closed.
952
1050
  status: The span outcome, ``"ok"`` or ``"error"``.
@@ -961,6 +1059,7 @@ def _close_span(span: Span, status: str, exc: BaseException | None) -> None:
961
1059
  """
962
1060
  span.events.append(end_event(span, status, exc))
963
1061
  backfill_baggage(span, context._live_baggage())
1062
+ span.closed = True
964
1063
  _flush(span)
965
1064
 
966
1065
 
@@ -30,6 +30,12 @@ class Span:
30
30
  ``start_ts`` is a ``time.monotonic()`` reading rather than wall-clock, so
31
31
  ``duration_ms`` can never go negative across a clock change. ``defaults`` are per-
32
32
  decorator field overrides, and ``events`` is the queue flushed together at span end.
33
+
34
+ ``closed`` marks a span whose events have already been handed to the worker (SPEC-036
35
+ FR-004). ``contextvars`` copies the *same* ``Span`` object into every task created inside a
36
+ span, so a fire-and-forget ``create_task`` can outlive its parent and append to a buffer
37
+ nothing will emit again. It is read by ``api._log`` at append time, which is the only place
38
+ that can notice: nothing in the library looks at a span after ``_close_span`` returns.
33
39
  """
34
40
 
35
41
  trace_id: str
@@ -39,6 +45,7 @@ class Span:
39
45
  start_ts: float
40
46
  defaults: dict[str, object] = field(default_factory=dict)
41
47
  events: list[dict[str, object]] = field(default_factory=list)
48
+ closed: bool = False
42
49
 
43
50
 
44
51
  def _iso_now() -> str:
@@ -52,12 +52,21 @@ class MultiSink:
52
52
  """
53
53
  self._sinks = sinks
54
54
  self.failed = 0
55
+ self._silent_failed = 0
55
56
  self._counter_lock = threading.Lock()
56
57
  self._stop_signal: threading.Event | None = None
57
58
 
58
59
  def emit(self, batch: list[dict[str, object]]) -> None:
59
60
  """Forwards the batch to every child in construction order, isolating failures (FR-002).
60
61
 
62
+ A failing child that reports **nothing** of its own is counted here, in events, because
63
+ otherwise it is invisible: ``losses()`` sums only children with a ``losses()``, so a
64
+ destination that has delivered nothing since the process started reported zero loss
65
+ forever (SPEC-036 FR-005). Which children are silent is decided by ``read_losses`` per
66
+ call, the same probe the aggregate uses, so a child that gains a ``losses()`` later moves
67
+ categories on its own — at the cost that the aggregate can then *fall*, which is why the
68
+ counter is per child rather than a single total.
69
+
61
70
  Partial success stays isolated: a retry there would re-deliver the batch to the children
62
71
  that already took it, and duplicates are worse than the one failure already counted on
63
72
  ``failed`` and written to stderr. Total failure delivered nothing, so it has no
@@ -84,8 +93,11 @@ class MultiSink:
84
93
  try:
85
94
  sink.emit(batch)
86
95
  except Exception as err:
96
+ silent = read_losses(sink) is None
87
97
  with self._counter_lock:
88
98
  self.failed += 1
99
+ if silent:
100
+ self._silent_failed += len(batch)
89
101
  if first_error is None:
90
102
  first_error = err
91
103
  _diag.absorbed(
@@ -147,18 +159,32 @@ class MultiSink:
147
159
  """Sums the children's losses so a fan-out reports the whole tree (SPEC-026 FR-002).
148
160
 
149
161
  Nesting is handled for free, since a child ``MultiSink`` is just another sink with a
150
- ``losses()``. ``MultiSink.failed`` is deliberately absent from the total: it counts
162
+ ``losses()``. ~~``MultiSink.failed`` is deliberately absent from the total: it counts
151
163
  child calls that raised, not events, so adding it to a per-event figure would produce a
152
- number with no unit, and the children already report their own loss in events.
164
+ number with no unit, and the children already report their own loss in events.~~ —
165
+ superseded in part by SPEC-036 FR-005. ``failed`` is still absent, and still for that
166
+ reason. What was wrong was concluding that the fan-out therefore reports nothing: a child
167
+ with no ``losses()`` contributed nothing to the total, so a permanently dead destination
168
+ was invisible with ``health()`` reading all zeros. ``_silent_failed`` is the same loss in
169
+ the **right unit** — events from failing children that report nothing — so a reporting
170
+ child is still left to report itself and nothing is counted twice.
171
+
172
+ The figure is a **total**, not a breakdown: it cannot say which child is failing. The
173
+ per-batch stderr line names the class (:meth:`emit`), which is where that lives.
174
+
175
+ ``close()`` deliberately does not move ``_silent_failed``. Its failure path has no batch,
176
+ so there is no event count to add, and a ``+1`` there would mix units in exactly the way
177
+ that keeps ``failed`` out of this sum. A client buffer lost to a failed close is of
178
+ unknown size, and an invented number is worse than an absent one.
153
179
 
154
180
  Args:
155
181
  None.
156
182
 
157
183
  Returns:
158
- The summed losses, or ``None`` when no child reported anything, including an empty
159
- fan-out. FR-003 separates "reports nothing" from "reports no loss", and a tree of
160
- silent children has not given a clean bill of health; one reporting child is enough to
161
- make the total meaningful, since the silent ones contribute zero.
184
+ The summed losses, or ``None`` when no child reported anything **and** no silent child
185
+ has lost a batch. FR-003 separates "reports nothing" from "reports no loss", and a tree
186
+ of silent children that has lost nothing has still not given a clean bill of health —
187
+ but one that *has* lost something now says so, which is the whole of FR-005.
162
188
 
163
189
  Raises:
164
190
  None. A child without ``losses()`` contributes zero, and a child whose accessor raises
@@ -167,11 +193,13 @@ class MultiSink:
167
193
  """
168
194
  children = [read_losses(sink) for sink in self._sinks]
169
195
  reported = [child for child in children if child is not None]
170
- if not reported:
196
+ with self._counter_lock:
197
+ silent_failed = self._silent_failed
198
+ if not reported and not silent_failed:
171
199
  return None
172
200
  return SinkLosses(
173
201
  dropped=sum(child.dropped for child in reported),
174
- failed=sum(child.failed for child in reported),
202
+ failed=sum(child.failed for child in reported) + silent_failed,
175
203
  )
176
204
 
177
205
  def close(self) -> None:
@@ -144,6 +144,23 @@ class Health:
144
144
  ``shutdown()``, and it is the signal that a deployment shares a sink across a fork at
145
145
  all. ``True`` for a shared ``StdoutSink`` too, whose ``close()`` only flushes — so a
146
146
  ``True`` is not by itself evidence that anything is held.
147
+ orphan_lost: Events lost on the **synchronous** path — a level call made with no active
148
+ span, which emits on the caller's own thread with no worker between it and the sink
149
+ (SPEC-036 FR-003). Until this field existed that loss was counted nowhere: ``health()``
150
+ describes a worker, and this path has none, so a process logging only this way read all
151
+ zeros over total loss. Not ``failed_batches``, which means batches a worker abandoned
152
+ after spending a retry budget and kept running past; there is no batch, no retry and no
153
+ worker here. Not ``SinkLosses`` either — the sink did not absorb anything, it raised,
154
+ which is what SPEC-026 requires of it. It covers everything inside the orphan guard, a
155
+ sink that failed to *construct* included, so it climbing means **the destination or the
156
+ data**.
157
+ in_span_lost: Events lost while being built *inside* a span (SPEC-037 AC-5c, deferred to
158
+ SPEC-036 FR-003 so the pair was designed together). The in-span path cannot lose an
159
+ event at ``emit`` — that is ``failed_batches`` — so this climbing means **the data**,
160
+ always: a value that could not be built into an event. Two fields rather than one
161
+ because they aggregate different failure populations and so fail SPEC-026's test,
162
+ *would one number hide which fix applies*. Their sum is deliberately not reported: with
163
+ different populations it is a number nobody can act on.
147
164
  """
148
165
 
149
166
  queued: int
@@ -156,6 +173,8 @@ class Health:
156
173
  incomplete_swaps: int = 0
157
174
  closing_sinks: int = 0
158
175
  inherited_sink: bool = False
176
+ orphan_lost: int = 0
177
+ in_span_lost: int = 0
159
178
 
160
179
 
161
180
  class _FlushMarker: