margin-meter 0.6.2__tar.gz → 0.6.4__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (27) hide show
  1. {margin_meter-0.6.2 → margin_meter-0.6.4}/PKG-INFO +45 -1
  2. {margin_meter-0.6.2 → margin_meter-0.6.4}/README.md +44 -0
  3. {margin_meter-0.6.2 → margin_meter-0.6.4}/margin_meter/__init__.py +1 -1
  4. {margin_meter-0.6.2 → margin_meter-0.6.4}/margin_meter/client.py +60 -0
  5. {margin_meter-0.6.2 → margin_meter-0.6.4}/margin_meter/pytest_plugin.py +10 -0
  6. {margin_meter-0.6.2 → margin_meter-0.6.4}/margin_meter.egg-info/PKG-INFO +45 -1
  7. {margin_meter-0.6.2 → margin_meter-0.6.4}/margin_meter/__main__.py +0 -0
  8. {margin_meter-0.6.2 → margin_meter-0.6.4}/margin_meter/batching.py +0 -0
  9. {margin_meter-0.6.2 → margin_meter-0.6.4}/margin_meter/braintrust.py +0 -0
  10. {margin_meter-0.6.2 → margin_meter-0.6.4}/margin_meter/deepeval.py +0 -0
  11. {margin_meter-0.6.2 → margin_meter-0.6.4}/margin_meter/doctor.py +0 -0
  12. {margin_meter-0.6.2 → margin_meter-0.6.4}/margin_meter/evals.py +0 -0
  13. {margin_meter-0.6.2 → margin_meter-0.6.4}/margin_meter/instrument.py +0 -0
  14. {margin_meter-0.6.2 → margin_meter-0.6.4}/margin_meter/langfuse.py +0 -0
  15. {margin_meter-0.6.2 → margin_meter-0.6.4}/margin_meter/langsmith.py +0 -0
  16. {margin_meter-0.6.2 → margin_meter-0.6.4}/margin_meter/model_override.py +0 -0
  17. {margin_meter-0.6.2 → margin_meter-0.6.4}/margin_meter/promptfoo.py +0 -0
  18. {margin_meter-0.6.2 → margin_meter-0.6.4}/margin_meter/providers.py +0 -0
  19. {margin_meter-0.6.2 → margin_meter-0.6.4}/margin_meter/ragas.py +0 -0
  20. {margin_meter-0.6.2 → margin_meter-0.6.4}/margin_meter/substrate.py +0 -0
  21. {margin_meter-0.6.2 → margin_meter-0.6.4}/margin_meter/values.py +0 -0
  22. {margin_meter-0.6.2 → margin_meter-0.6.4}/margin_meter.egg-info/SOURCES.txt +0 -0
  23. {margin_meter-0.6.2 → margin_meter-0.6.4}/margin_meter.egg-info/dependency_links.txt +0 -0
  24. {margin_meter-0.6.2 → margin_meter-0.6.4}/margin_meter.egg-info/entry_points.txt +0 -0
  25. {margin_meter-0.6.2 → margin_meter-0.6.4}/margin_meter.egg-info/top_level.txt +0 -0
  26. {margin_meter-0.6.2 → margin_meter-0.6.4}/pyproject.toml +0 -0
  27. {margin_meter-0.6.2 → margin_meter-0.6.4}/setup.cfg +0 -0
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: margin-meter
3
- Version: 0.6.2
3
+ Version: 0.6.4
4
4
  Summary: Tiny stdlib-only client that auto-instruments OpenAI/Anthropic/Gemini and emits LLM call + outcome economics to a Margin ingest API.
5
5
  Author: Margin
6
6
  License: MIT
@@ -452,6 +452,50 @@ current_task.set(ticket.topic) # "billing" | "password-reset" | …
452
452
  Leave it off and the row's `task_key` is simply unset (`unknown`) — never a
453
453
  fabricated bucket.
454
454
 
455
+ ## Say where each call RAN (`environment`)
456
+
457
+ `environment` splits the work you sell from the work you did to decide what to
458
+ sell. A spike, an eval pass and your production traffic are three different kinds
459
+ of spend, and only one of them belongs in the number you report upward.
460
+
461
+ ```python
462
+ meter.record_call(
463
+ workflow_id="support-triage",
464
+ provider="openai",
465
+ model="gpt-4o-mini",
466
+ input_tokens=812,
467
+ output_tokens=96,
468
+ environment="prod", # "prod" | "dev" | "eval"
469
+ )
470
+ ```
471
+
472
+ Setting it once per process is usually easier. A deployment already knows whether
473
+ it is production; the call site does not:
474
+
475
+ ```bash
476
+ export MARGIN_ENVIRONMENT=prod
477
+ ```
478
+
479
+ An explicit argument wins over the variable, so one process can still label a
480
+ single call differently.
481
+
482
+ | value | send it for |
483
+ | --- | --- |
484
+ | `prod` | traffic your customers cause |
485
+ | `dev` | a spike, a local run, a branch deploy |
486
+ | `eval` | a graded pass, in CI or from an eval harness |
487
+
488
+ Why it earns a field. `cost_per_outcome` is total call spend divided by passed
489
+ outcomes. With `environment` unset, your R&D and your operations land in that one
490
+ number, so "your AI workforce costs $X per outcome" describes neither. Margin
491
+ already keeps its own face-off spend out of your number; this is how you get the
492
+ same split on your side.
493
+
494
+ Two things it will not do. It does not default: leave it off and the row is
495
+ `unknown`, which is honest and costs you nothing, where guessing `prod` would book
496
+ your spike as production. And it is never priced and never rolled into cost. It is
497
+ a label, like `team`.
498
+
455
499
  ## API
456
500
 
457
501
  | Method | Emits to | Notes |
@@ -440,6 +440,50 @@ current_task.set(ticket.topic) # "billing" | "password-reset" | …
440
440
  Leave it off and the row's `task_key` is simply unset (`unknown`) — never a
441
441
  fabricated bucket.
442
442
 
443
+ ## Say where each call RAN (`environment`)
444
+
445
+ `environment` splits the work you sell from the work you did to decide what to
446
+ sell. A spike, an eval pass and your production traffic are three different kinds
447
+ of spend, and only one of them belongs in the number you report upward.
448
+
449
+ ```python
450
+ meter.record_call(
451
+ workflow_id="support-triage",
452
+ provider="openai",
453
+ model="gpt-4o-mini",
454
+ input_tokens=812,
455
+ output_tokens=96,
456
+ environment="prod", # "prod" | "dev" | "eval"
457
+ )
458
+ ```
459
+
460
+ Setting it once per process is usually easier. A deployment already knows whether
461
+ it is production; the call site does not:
462
+
463
+ ```bash
464
+ export MARGIN_ENVIRONMENT=prod
465
+ ```
466
+
467
+ An explicit argument wins over the variable, so one process can still label a
468
+ single call differently.
469
+
470
+ | value | send it for |
471
+ | --- | --- |
472
+ | `prod` | traffic your customers cause |
473
+ | `dev` | a spike, a local run, a branch deploy |
474
+ | `eval` | a graded pass, in CI or from an eval harness |
475
+
476
+ Why it earns a field. `cost_per_outcome` is total call spend divided by passed
477
+ outcomes. With `environment` unset, your R&D and your operations land in that one
478
+ number, so "your AI workforce costs $X per outcome" describes neither. Margin
479
+ already keeps its own face-off spend out of your number; this is how you get the
480
+ same split on your side.
481
+
482
+ Two things it will not do. It does not default: leave it off and the row is
483
+ `unknown`, which is honest and costs you nothing, where guessing `prod` would book
484
+ your spike as production. And it is never priced and never rolled into cost. It is
485
+ a label, like `team`.
486
+
443
487
  ## API
444
488
 
445
489
  | Method | Emits to | Notes |
@@ -106,4 +106,4 @@ __all__ = [
106
106
  # closing the invoice-reconcile coverage gap is "ship the fix, wait for the seed's
107
107
  # dependency refresh, re-run the reconcile", and a runtime that misreports which
108
108
  # SDK metered a call makes that unanswerable.
109
- __version__ = "0.6.2"
109
+ __version__ = "0.6.4"
@@ -1479,6 +1479,8 @@ class MarginMeter:
1479
1479
  harness: str | None = None,
1480
1480
  finish_reason: str | None = None,
1481
1481
  provider_call_id: str | None = None,
1482
+ case_id: str | None = None,
1483
+ prompt_hash: str | None = None,
1482
1484
  ) -> IngestResult:
1483
1485
  """Emit one measured LLM call to ``POST /api/ingest/calls``.
1484
1486
 
@@ -1630,6 +1632,18 @@ class MarginMeter:
1630
1632
  harness = harness[:64]
1631
1633
  if harness is not None:
1632
1634
  payload["harness"] = harness
1635
+ # WHICH EVAL CASE this call ran under (MAR-1782). Read from MARGIN_CASE_ID AT CALL
1636
+ # TIME, because the case changes per test inside one process; the pytest plugin sets
1637
+ # it around each test. Blank reads as absent. ⛔ It is NOT ``task_key``: that is the
1638
+ # task CLASS the parity gate buckets on, and a unique case id there would give every
1639
+ # call its own class. Only repeat detection reads this. A dimension, never money.
1640
+ case_id = (case_id or os.environ.get("MARGIN_CASE_ID") or "").strip()[:128] or None
1641
+ if case_id is not None:
1642
+ payload["case_id"] = case_id
1643
+ # sha256 hex of the request, never the text. Lets repeat detection tell an identical
1644
+ # re-send from a different request of similar size.
1645
+ if prompt_hash:
1646
+ payload["prompt_hash"] = prompt_hash
1633
1647
  if cost_usd is not None:
1634
1648
  payload["cost_usd"] = cost_usd
1635
1649
  # Only send a non-default tier: a synchronous call omits it (the server
@@ -1719,6 +1733,7 @@ class MarginMeter:
1719
1733
  is_simulated: bool = False,
1720
1734
  event_id: str | None = None,
1721
1735
  span_id: str | None = None,
1736
+ environment: str | None = None,
1722
1737
  value_usd: float | None = None,
1723
1738
  ) -> IngestResult:
1724
1739
  """Emit one outcome (a unit of productivity) to ``POST /api/ingest/outcomes``.
@@ -1748,11 +1763,25 @@ class MarginMeter:
1748
1763
  value (revenue booked, a ticket's handling cost avoided); omit it and the
1749
1764
  outcome is measured as cost-per-outcome, the honest default. It is optional
1750
1765
  by construction — absent is the normal case, never defaulted to 0.
1766
+
1767
+ ``environment`` is WHERE this outcome was produced — the DENOMINATOR twin of
1768
+ :meth:`record_call`'s (MAR-1562): ``prod`` for the work the customer sells,
1769
+ ``dev``/``eval`` for the work they did to decide what to sell. Same resolution
1770
+ rules as the call path — explicit wins, then ``MARGIN_ENVIRONMENT``, then
1771
+ absent. ⛔ Omit it and the server stores NULL (UNKNOWN); it is never assumed to
1772
+ be ``prod``.
1751
1773
  """
1752
1774
  if span_id is None:
1753
1775
  _span = _SPAN.get()
1754
1776
  if _span is not None:
1755
1777
  span_id = _span.span_id
1778
+ # Same env-var fallback and same server-width clamp as `record_call`, for the same
1779
+ # reason: `ingest._require_str` RAISES (422) on an over-long value, so an outcome
1780
+ # label longer than the column would drop the metered row rather than mislabel it —
1781
+ # and a dropped outcome is a silently smaller denominator, the flattering direction.
1782
+ environment = environment or (os.environ.get("MARGIN_ENVIRONMENT") or "").strip() or None
1783
+ if environment:
1784
+ environment = environment[:32]
1756
1785
  payload: dict[str, Any] = {
1757
1786
  "workflow_id": workflow_id,
1758
1787
  "passed": passed,
@@ -1781,6 +1810,10 @@ class MarginMeter:
1781
1810
  # send. An older server that predates the column ignores an unknown key.
1782
1811
  if value_usd is not None:
1783
1812
  payload["value_usd"] = value_usd
1813
+ # Rides only when one was declared (explicitly or via the process env). An
1814
+ # unlabelled outcome sends nothing and the server stores NULL, never `prod`.
1815
+ if environment is not None:
1816
+ payload["environment"] = environment
1784
1817
  return self._emit(OUTCOMES_PATH, payload)
1785
1818
 
1786
1819
  def measure(
@@ -1837,6 +1870,7 @@ class MarginMeter:
1837
1870
  substrate: str | None = None,
1838
1871
  task_key: str | None = None,
1839
1872
  provider_call_id: str | None = None,
1873
+ prompt_hash: str | None = None,
1840
1874
  ) -> IngestResult:
1841
1875
  """Meter one call by reading tokens straight off its provider ``response``.
1842
1876
 
@@ -1882,6 +1916,7 @@ class MarginMeter:
1882
1916
  # forwarded as finish_reason (MAR-209). A signal, never a billing verdict.
1883
1917
  finish_reason=usage.stop_reason,
1884
1918
  is_simulated=is_simulated,
1919
+ prompt_hash=prompt_hash,
1885
1920
  )
1886
1921
 
1887
1922
  def wrap(
@@ -2001,6 +2036,7 @@ class MarginMeter:
2001
2036
  is_retry=is_retry,
2002
2037
  session_id=session_id,
2003
2038
  prompt_id=prompt_id_from(kwargs),
2039
+ prompt_hash=prompt_hash_from(kwargs),
2004
2040
  operation=operation,
2005
2041
  substrate=call_substrate,
2006
2042
  task_key=call_task_key,
@@ -2028,6 +2064,7 @@ class MarginMeter:
2028
2064
  is_retry=is_retry,
2029
2065
  session_id=session_id,
2030
2066
  prompt_id=prompt_id_from(kwargs),
2067
+ prompt_hash=prompt_hash_from(kwargs),
2031
2068
  operation=operation,
2032
2069
  substrate=call_substrate,
2033
2070
  task_key=call_task_key,
@@ -2084,6 +2121,7 @@ class MarginMeter:
2084
2121
  is_retry=is_retry,
2085
2122
  session_id=session_id,
2086
2123
  prompt_id=prompt_id_from(kwargs),
2124
+ prompt_hash=prompt_hash_from(kwargs),
2087
2125
  operation=operation,
2088
2126
  substrate=call_substrate,
2089
2127
  task_key=call_task_key,
@@ -2223,6 +2261,28 @@ class MarginMeter:
2223
2261
  return _wrapped
2224
2262
 
2225
2263
 
2264
+ _PROMPT_FIELDS = ("messages", "input", "contents", "prompt", "system", "tools")
2265
+
2266
+
2267
+ def prompt_hash_from(kwargs: dict) -> str | None:
2268
+ """sha256 hex of the request's prompt-bearing kwargs, or ``None``. Never the text.
2269
+
2270
+ Two calls that send the same messages to the same model hash alike, which is all
2271
+ repeat detection needs. Best-effort: an unserialisable request yields ``None``.
2272
+ """
2273
+ import hashlib
2274
+ import json
2275
+
2276
+ part = {k: kwargs[k] for k in _PROMPT_FIELDS if kwargs.get(k) is not None}
2277
+ if not part:
2278
+ return None
2279
+ try:
2280
+ blob = json.dumps(part, sort_keys=True, default=repr).encode()
2281
+ except Exception:
2282
+ return None
2283
+ return hashlib.sha256(blob).hexdigest()
2284
+
2285
+
2226
2286
  def prompt_id_from(kwargs: dict) -> str | None:
2227
2287
  """Pull a caller-supplied ``margin_prompt_id`` out of call kwargs, if present.
2228
2288
 
@@ -291,6 +291,12 @@ def pytest_runtest_setup(item):
291
291
  nothing and is NOT warned about here. A warning raised in the setup phase becomes an ERROR
292
292
  under ``-W error``, which would flip a passing test to a fail in a suite that only has the
293
293
  SDK installed (the adapter, when enabled, reports a bad marker itself)."""
294
+ # Name the eval case for every call this test meters (MAR-1782), unconditionally like
295
+ # the marker below. Cleared in teardown so a call outside a test carries no stale case.
296
+ try:
297
+ os.environ["MARGIN_CASE_ID"] = str(item.nodeid)[:128]
298
+ except Exception:
299
+ pass
294
300
  try:
295
301
  value = value_from_marker(item)
296
302
  except ValueError:
@@ -299,6 +305,10 @@ def pytest_runtest_setup(item):
299
305
  item.user_properties.append((VALUE_PROPERTY, value))
300
306
 
301
307
 
308
+ def pytest_runtest_teardown(item):
309
+ os.environ.pop("MARGIN_CASE_ID", None)
310
+
311
+
302
312
  def pytest_configure(config):
303
313
  # Register the marker unconditionally so a suite that uses it without the
304
314
  # flag does not trip PytestUnknownMarkWarning.
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: margin-meter
3
- Version: 0.6.2
3
+ Version: 0.6.4
4
4
  Summary: Tiny stdlib-only client that auto-instruments OpenAI/Anthropic/Gemini and emits LLM call + outcome economics to a Margin ingest API.
5
5
  Author: Margin
6
6
  License: MIT
@@ -452,6 +452,50 @@ current_task.set(ticket.topic) # "billing" | "password-reset" | …
452
452
  Leave it off and the row's `task_key` is simply unset (`unknown`) — never a
453
453
  fabricated bucket.
454
454
 
455
+ ## Say where each call RAN (`environment`)
456
+
457
+ `environment` splits the work you sell from the work you did to decide what to
458
+ sell. A spike, an eval pass and your production traffic are three different kinds
459
+ of spend, and only one of them belongs in the number you report upward.
460
+
461
+ ```python
462
+ meter.record_call(
463
+ workflow_id="support-triage",
464
+ provider="openai",
465
+ model="gpt-4o-mini",
466
+ input_tokens=812,
467
+ output_tokens=96,
468
+ environment="prod", # "prod" | "dev" | "eval"
469
+ )
470
+ ```
471
+
472
+ Setting it once per process is usually easier. A deployment already knows whether
473
+ it is production; the call site does not:
474
+
475
+ ```bash
476
+ export MARGIN_ENVIRONMENT=prod
477
+ ```
478
+
479
+ An explicit argument wins over the variable, so one process can still label a
480
+ single call differently.
481
+
482
+ | value | send it for |
483
+ | --- | --- |
484
+ | `prod` | traffic your customers cause |
485
+ | `dev` | a spike, a local run, a branch deploy |
486
+ | `eval` | a graded pass, in CI or from an eval harness |
487
+
488
+ Why it earns a field. `cost_per_outcome` is total call spend divided by passed
489
+ outcomes. With `environment` unset, your R&D and your operations land in that one
490
+ number, so "your AI workforce costs $X per outcome" describes neither. Margin
491
+ already keeps its own face-off spend out of your number; this is how you get the
492
+ same split on your side.
493
+
494
+ Two things it will not do. It does not default: leave it off and the row is
495
+ `unknown`, which is honest and costs you nothing, where guessing `prod` would book
496
+ your spike as production. And it is never priced and never rolled into cost. It is
497
+ a label, like `team`.
498
+
455
499
  ## API
456
500
 
457
501
  | Method | Emits to | Notes |
File without changes