margin-meter 0.6.2__tar.gz → 0.6.4__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {margin_meter-0.6.2 → margin_meter-0.6.4}/PKG-INFO +45 -1
- {margin_meter-0.6.2 → margin_meter-0.6.4}/README.md +44 -0
- {margin_meter-0.6.2 → margin_meter-0.6.4}/margin_meter/__init__.py +1 -1
- {margin_meter-0.6.2 → margin_meter-0.6.4}/margin_meter/client.py +60 -0
- {margin_meter-0.6.2 → margin_meter-0.6.4}/margin_meter/pytest_plugin.py +10 -0
- {margin_meter-0.6.2 → margin_meter-0.6.4}/margin_meter.egg-info/PKG-INFO +45 -1
- {margin_meter-0.6.2 → margin_meter-0.6.4}/margin_meter/__main__.py +0 -0
- {margin_meter-0.6.2 → margin_meter-0.6.4}/margin_meter/batching.py +0 -0
- {margin_meter-0.6.2 → margin_meter-0.6.4}/margin_meter/braintrust.py +0 -0
- {margin_meter-0.6.2 → margin_meter-0.6.4}/margin_meter/deepeval.py +0 -0
- {margin_meter-0.6.2 → margin_meter-0.6.4}/margin_meter/doctor.py +0 -0
- {margin_meter-0.6.2 → margin_meter-0.6.4}/margin_meter/evals.py +0 -0
- {margin_meter-0.6.2 → margin_meter-0.6.4}/margin_meter/instrument.py +0 -0
- {margin_meter-0.6.2 → margin_meter-0.6.4}/margin_meter/langfuse.py +0 -0
- {margin_meter-0.6.2 → margin_meter-0.6.4}/margin_meter/langsmith.py +0 -0
- {margin_meter-0.6.2 → margin_meter-0.6.4}/margin_meter/model_override.py +0 -0
- {margin_meter-0.6.2 → margin_meter-0.6.4}/margin_meter/promptfoo.py +0 -0
- {margin_meter-0.6.2 → margin_meter-0.6.4}/margin_meter/providers.py +0 -0
- {margin_meter-0.6.2 → margin_meter-0.6.4}/margin_meter/ragas.py +0 -0
- {margin_meter-0.6.2 → margin_meter-0.6.4}/margin_meter/substrate.py +0 -0
- {margin_meter-0.6.2 → margin_meter-0.6.4}/margin_meter/values.py +0 -0
- {margin_meter-0.6.2 → margin_meter-0.6.4}/margin_meter.egg-info/SOURCES.txt +0 -0
- {margin_meter-0.6.2 → margin_meter-0.6.4}/margin_meter.egg-info/dependency_links.txt +0 -0
- {margin_meter-0.6.2 → margin_meter-0.6.4}/margin_meter.egg-info/entry_points.txt +0 -0
- {margin_meter-0.6.2 → margin_meter-0.6.4}/margin_meter.egg-info/top_level.txt +0 -0
- {margin_meter-0.6.2 → margin_meter-0.6.4}/pyproject.toml +0 -0
- {margin_meter-0.6.2 → margin_meter-0.6.4}/setup.cfg +0 -0
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.4
|
|
2
2
|
Name: margin-meter
|
|
3
|
-
Version: 0.6.
|
|
3
|
+
Version: 0.6.4
|
|
4
4
|
Summary: Tiny stdlib-only client that auto-instruments OpenAI/Anthropic/Gemini and emits LLM call + outcome economics to a Margin ingest API.
|
|
5
5
|
Author: Margin
|
|
6
6
|
License: MIT
|
|
@@ -452,6 +452,50 @@ current_task.set(ticket.topic) # "billing" | "password-reset" | …
|
|
|
452
452
|
Leave it off and the row's `task_key` is simply unset (`unknown`) — never a
|
|
453
453
|
fabricated bucket.
|
|
454
454
|
|
|
455
|
+
## Say where each call RAN (`environment`)
|
|
456
|
+
|
|
457
|
+
`environment` splits the work you sell from the work you did to decide what to
|
|
458
|
+
sell. A spike, an eval pass and your production traffic are three different kinds
|
|
459
|
+
of spend, and only one of them belongs in the number you report upward.
|
|
460
|
+
|
|
461
|
+
```python
|
|
462
|
+
meter.record_call(
|
|
463
|
+
workflow_id="support-triage",
|
|
464
|
+
provider="openai",
|
|
465
|
+
model="gpt-4o-mini",
|
|
466
|
+
input_tokens=812,
|
|
467
|
+
output_tokens=96,
|
|
468
|
+
environment="prod", # "prod" | "dev" | "eval"
|
|
469
|
+
)
|
|
470
|
+
```
|
|
471
|
+
|
|
472
|
+
Setting it once per process is usually easier. A deployment already knows whether
|
|
473
|
+
it is production; the call site does not:
|
|
474
|
+
|
|
475
|
+
```bash
|
|
476
|
+
export MARGIN_ENVIRONMENT=prod
|
|
477
|
+
```
|
|
478
|
+
|
|
479
|
+
An explicit argument wins over the variable, so one process can still label a
|
|
480
|
+
single call differently.
|
|
481
|
+
|
|
482
|
+
| value | send it for |
|
|
483
|
+
| --- | --- |
|
|
484
|
+
| `prod` | traffic your customers cause |
|
|
485
|
+
| `dev` | a spike, a local run, a branch deploy |
|
|
486
|
+
| `eval` | a graded pass, in CI or from an eval harness |
|
|
487
|
+
|
|
488
|
+
Why it earns a field. `cost_per_outcome` is total call spend divided by passed
|
|
489
|
+
outcomes. With `environment` unset, your R&D and your operations land in that one
|
|
490
|
+
number, so "your AI workforce costs $X per outcome" describes neither. Margin
|
|
491
|
+
already keeps its own face-off spend out of your number; this is how you get the
|
|
492
|
+
same split on your side.
|
|
493
|
+
|
|
494
|
+
Two things it will not do. It does not default: leave it off and the row is
|
|
495
|
+
`unknown`, which is honest and costs you nothing, where guessing `prod` would book
|
|
496
|
+
your spike as production. And it is never priced and never rolled into cost. It is
|
|
497
|
+
a label, like `team`.
|
|
498
|
+
|
|
455
499
|
## API
|
|
456
500
|
|
|
457
501
|
| Method | Emits to | Notes |
|
|
@@ -440,6 +440,50 @@ current_task.set(ticket.topic) # "billing" | "password-reset" | …
|
|
|
440
440
|
Leave it off and the row's `task_key` is simply unset (`unknown`) — never a
|
|
441
441
|
fabricated bucket.
|
|
442
442
|
|
|
443
|
+
## Say where each call RAN (`environment`)
|
|
444
|
+
|
|
445
|
+
`environment` splits the work you sell from the work you did to decide what to
|
|
446
|
+
sell. A spike, an eval pass and your production traffic are three different kinds
|
|
447
|
+
of spend, and only one of them belongs in the number you report upward.
|
|
448
|
+
|
|
449
|
+
```python
|
|
450
|
+
meter.record_call(
|
|
451
|
+
workflow_id="support-triage",
|
|
452
|
+
provider="openai",
|
|
453
|
+
model="gpt-4o-mini",
|
|
454
|
+
input_tokens=812,
|
|
455
|
+
output_tokens=96,
|
|
456
|
+
environment="prod", # "prod" | "dev" | "eval"
|
|
457
|
+
)
|
|
458
|
+
```
|
|
459
|
+
|
|
460
|
+
Setting it once per process is usually easier. A deployment already knows whether
|
|
461
|
+
it is production; the call site does not:
|
|
462
|
+
|
|
463
|
+
```bash
|
|
464
|
+
export MARGIN_ENVIRONMENT=prod
|
|
465
|
+
```
|
|
466
|
+
|
|
467
|
+
An explicit argument wins over the variable, so one process can still label a
|
|
468
|
+
single call differently.
|
|
469
|
+
|
|
470
|
+
| value | send it for |
|
|
471
|
+
| --- | --- |
|
|
472
|
+
| `prod` | traffic your customers cause |
|
|
473
|
+
| `dev` | a spike, a local run, a branch deploy |
|
|
474
|
+
| `eval` | a graded pass, in CI or from an eval harness |
|
|
475
|
+
|
|
476
|
+
Why it earns a field. `cost_per_outcome` is total call spend divided by passed
|
|
477
|
+
outcomes. With `environment` unset, your R&D and your operations land in that one
|
|
478
|
+
number, so "your AI workforce costs $X per outcome" describes neither. Margin
|
|
479
|
+
already keeps its own face-off spend out of your number; this is how you get the
|
|
480
|
+
same split on your side.
|
|
481
|
+
|
|
482
|
+
Two things it will not do. It does not default: leave it off and the row is
|
|
483
|
+
`unknown`, which is honest and costs you nothing, where guessing `prod` would book
|
|
484
|
+
your spike as production. And it is never priced and never rolled into cost. It is
|
|
485
|
+
a label, like `team`.
|
|
486
|
+
|
|
443
487
|
## API
|
|
444
488
|
|
|
445
489
|
| Method | Emits to | Notes |
|
|
@@ -106,4 +106,4 @@ __all__ = [
|
|
|
106
106
|
# closing the invoice-reconcile coverage gap is "ship the fix, wait for the seed's
|
|
107
107
|
# dependency refresh, re-run the reconcile", and a runtime that misreports which
|
|
108
108
|
# SDK metered a call makes that unanswerable.
|
|
109
|
-
__version__ = "0.6.
|
|
109
|
+
__version__ = "0.6.4"
|
|
@@ -1479,6 +1479,8 @@ class MarginMeter:
|
|
|
1479
1479
|
harness: str | None = None,
|
|
1480
1480
|
finish_reason: str | None = None,
|
|
1481
1481
|
provider_call_id: str | None = None,
|
|
1482
|
+
case_id: str | None = None,
|
|
1483
|
+
prompt_hash: str | None = None,
|
|
1482
1484
|
) -> IngestResult:
|
|
1483
1485
|
"""Emit one measured LLM call to ``POST /api/ingest/calls``.
|
|
1484
1486
|
|
|
@@ -1630,6 +1632,18 @@ class MarginMeter:
|
|
|
1630
1632
|
harness = harness[:64]
|
|
1631
1633
|
if harness is not None:
|
|
1632
1634
|
payload["harness"] = harness
|
|
1635
|
+
# WHICH EVAL CASE this call ran under (MAR-1782). Read from MARGIN_CASE_ID AT CALL
|
|
1636
|
+
# TIME, because the case changes per test inside one process; the pytest plugin sets
|
|
1637
|
+
# it around each test. Blank reads as absent. ⛔ It is NOT ``task_key``: that is the
|
|
1638
|
+
# task CLASS the parity gate buckets on, and a unique case id there would give every
|
|
1639
|
+
# call its own class. Only repeat detection reads this. A dimension, never money.
|
|
1640
|
+
case_id = (case_id or os.environ.get("MARGIN_CASE_ID") or "").strip()[:128] or None
|
|
1641
|
+
if case_id is not None:
|
|
1642
|
+
payload["case_id"] = case_id
|
|
1643
|
+
# sha256 hex of the request, never the text. Lets repeat detection tell an identical
|
|
1644
|
+
# re-send from a different request of similar size.
|
|
1645
|
+
if prompt_hash:
|
|
1646
|
+
payload["prompt_hash"] = prompt_hash
|
|
1633
1647
|
if cost_usd is not None:
|
|
1634
1648
|
payload["cost_usd"] = cost_usd
|
|
1635
1649
|
# Only send a non-default tier: a synchronous call omits it (the server
|
|
@@ -1719,6 +1733,7 @@ class MarginMeter:
|
|
|
1719
1733
|
is_simulated: bool = False,
|
|
1720
1734
|
event_id: str | None = None,
|
|
1721
1735
|
span_id: str | None = None,
|
|
1736
|
+
environment: str | None = None,
|
|
1722
1737
|
value_usd: float | None = None,
|
|
1723
1738
|
) -> IngestResult:
|
|
1724
1739
|
"""Emit one outcome (a unit of productivity) to ``POST /api/ingest/outcomes``.
|
|
@@ -1748,11 +1763,25 @@ class MarginMeter:
|
|
|
1748
1763
|
value (revenue booked, a ticket's handling cost avoided); omit it and the
|
|
1749
1764
|
outcome is measured as cost-per-outcome, the honest default. It is optional
|
|
1750
1765
|
by construction — absent is the normal case, never defaulted to 0.
|
|
1766
|
+
|
|
1767
|
+
``environment`` is WHERE this outcome was produced — the DENOMINATOR twin of
|
|
1768
|
+
:meth:`record_call`'s (MAR-1562): ``prod`` for the work the customer sells,
|
|
1769
|
+
``dev``/``eval`` for the work they did to decide what to sell. Same resolution
|
|
1770
|
+
rules as the call path — explicit wins, then ``MARGIN_ENVIRONMENT``, then
|
|
1771
|
+
absent. ⛔ Omit it and the server stores NULL (UNKNOWN); it is never assumed to
|
|
1772
|
+
be ``prod``.
|
|
1751
1773
|
"""
|
|
1752
1774
|
if span_id is None:
|
|
1753
1775
|
_span = _SPAN.get()
|
|
1754
1776
|
if _span is not None:
|
|
1755
1777
|
span_id = _span.span_id
|
|
1778
|
+
# Same env-var fallback and same server-width clamp as `record_call`, for the same
|
|
1779
|
+
# reason: `ingest._require_str` RAISES (422) on an over-long value, so an outcome
|
|
1780
|
+
# label longer than the column would drop the metered row rather than mislabel it —
|
|
1781
|
+
# and a dropped outcome is a silently smaller denominator, the flattering direction.
|
|
1782
|
+
environment = environment or (os.environ.get("MARGIN_ENVIRONMENT") or "").strip() or None
|
|
1783
|
+
if environment:
|
|
1784
|
+
environment = environment[:32]
|
|
1756
1785
|
payload: dict[str, Any] = {
|
|
1757
1786
|
"workflow_id": workflow_id,
|
|
1758
1787
|
"passed": passed,
|
|
@@ -1781,6 +1810,10 @@ class MarginMeter:
|
|
|
1781
1810
|
# send. An older server that predates the column ignores an unknown key.
|
|
1782
1811
|
if value_usd is not None:
|
|
1783
1812
|
payload["value_usd"] = value_usd
|
|
1813
|
+
# Rides only when one was declared (explicitly or via the process env). An
|
|
1814
|
+
# unlabelled outcome sends nothing and the server stores NULL, never `prod`.
|
|
1815
|
+
if environment is not None:
|
|
1816
|
+
payload["environment"] = environment
|
|
1784
1817
|
return self._emit(OUTCOMES_PATH, payload)
|
|
1785
1818
|
|
|
1786
1819
|
def measure(
|
|
@@ -1837,6 +1870,7 @@ class MarginMeter:
|
|
|
1837
1870
|
substrate: str | None = None,
|
|
1838
1871
|
task_key: str | None = None,
|
|
1839
1872
|
provider_call_id: str | None = None,
|
|
1873
|
+
prompt_hash: str | None = None,
|
|
1840
1874
|
) -> IngestResult:
|
|
1841
1875
|
"""Meter one call by reading tokens straight off its provider ``response``.
|
|
1842
1876
|
|
|
@@ -1882,6 +1916,7 @@ class MarginMeter:
|
|
|
1882
1916
|
# forwarded as finish_reason (MAR-209). A signal, never a billing verdict.
|
|
1883
1917
|
finish_reason=usage.stop_reason,
|
|
1884
1918
|
is_simulated=is_simulated,
|
|
1919
|
+
prompt_hash=prompt_hash,
|
|
1885
1920
|
)
|
|
1886
1921
|
|
|
1887
1922
|
def wrap(
|
|
@@ -2001,6 +2036,7 @@ class MarginMeter:
|
|
|
2001
2036
|
is_retry=is_retry,
|
|
2002
2037
|
session_id=session_id,
|
|
2003
2038
|
prompt_id=prompt_id_from(kwargs),
|
|
2039
|
+
prompt_hash=prompt_hash_from(kwargs),
|
|
2004
2040
|
operation=operation,
|
|
2005
2041
|
substrate=call_substrate,
|
|
2006
2042
|
task_key=call_task_key,
|
|
@@ -2028,6 +2064,7 @@ class MarginMeter:
|
|
|
2028
2064
|
is_retry=is_retry,
|
|
2029
2065
|
session_id=session_id,
|
|
2030
2066
|
prompt_id=prompt_id_from(kwargs),
|
|
2067
|
+
prompt_hash=prompt_hash_from(kwargs),
|
|
2031
2068
|
operation=operation,
|
|
2032
2069
|
substrate=call_substrate,
|
|
2033
2070
|
task_key=call_task_key,
|
|
@@ -2084,6 +2121,7 @@ class MarginMeter:
|
|
|
2084
2121
|
is_retry=is_retry,
|
|
2085
2122
|
session_id=session_id,
|
|
2086
2123
|
prompt_id=prompt_id_from(kwargs),
|
|
2124
|
+
prompt_hash=prompt_hash_from(kwargs),
|
|
2087
2125
|
operation=operation,
|
|
2088
2126
|
substrate=call_substrate,
|
|
2089
2127
|
task_key=call_task_key,
|
|
@@ -2223,6 +2261,28 @@ class MarginMeter:
|
|
|
2223
2261
|
return _wrapped
|
|
2224
2262
|
|
|
2225
2263
|
|
|
2264
|
+
_PROMPT_FIELDS = ("messages", "input", "contents", "prompt", "system", "tools")
|
|
2265
|
+
|
|
2266
|
+
|
|
2267
|
+
def prompt_hash_from(kwargs: dict) -> str | None:
|
|
2268
|
+
"""sha256 hex of the request's prompt-bearing kwargs, or ``None``. Never the text.
|
|
2269
|
+
|
|
2270
|
+
Two calls that send the same messages to the same model hash alike, which is all
|
|
2271
|
+
repeat detection needs. Best-effort: an unserialisable request yields ``None``.
|
|
2272
|
+
"""
|
|
2273
|
+
import hashlib
|
|
2274
|
+
import json
|
|
2275
|
+
|
|
2276
|
+
part = {k: kwargs[k] for k in _PROMPT_FIELDS if kwargs.get(k) is not None}
|
|
2277
|
+
if not part:
|
|
2278
|
+
return None
|
|
2279
|
+
try:
|
|
2280
|
+
blob = json.dumps(part, sort_keys=True, default=repr).encode()
|
|
2281
|
+
except Exception:
|
|
2282
|
+
return None
|
|
2283
|
+
return hashlib.sha256(blob).hexdigest()
|
|
2284
|
+
|
|
2285
|
+
|
|
2226
2286
|
def prompt_id_from(kwargs: dict) -> str | None:
|
|
2227
2287
|
"""Pull a caller-supplied ``margin_prompt_id`` out of call kwargs, if present.
|
|
2228
2288
|
|
|
@@ -291,6 +291,12 @@ def pytest_runtest_setup(item):
|
|
|
291
291
|
nothing and is NOT warned about here. A warning raised in the setup phase becomes an ERROR
|
|
292
292
|
under ``-W error``, which would flip a passing test to a fail in a suite that only has the
|
|
293
293
|
SDK installed (the adapter, when enabled, reports a bad marker itself)."""
|
|
294
|
+
# Name the eval case for every call this test meters (MAR-1782), unconditionally like
|
|
295
|
+
# the marker below. Cleared in teardown so a call outside a test carries no stale case.
|
|
296
|
+
try:
|
|
297
|
+
os.environ["MARGIN_CASE_ID"] = str(item.nodeid)[:128]
|
|
298
|
+
except Exception:
|
|
299
|
+
pass
|
|
294
300
|
try:
|
|
295
301
|
value = value_from_marker(item)
|
|
296
302
|
except ValueError:
|
|
@@ -299,6 +305,10 @@ def pytest_runtest_setup(item):
|
|
|
299
305
|
item.user_properties.append((VALUE_PROPERTY, value))
|
|
300
306
|
|
|
301
307
|
|
|
308
|
+
def pytest_runtest_teardown(item):
|
|
309
|
+
os.environ.pop("MARGIN_CASE_ID", None)
|
|
310
|
+
|
|
311
|
+
|
|
302
312
|
def pytest_configure(config):
|
|
303
313
|
# Register the marker unconditionally so a suite that uses it without the
|
|
304
314
|
# flag does not trip PytestUnknownMarkWarning.
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.4
|
|
2
2
|
Name: margin-meter
|
|
3
|
-
Version: 0.6.
|
|
3
|
+
Version: 0.6.4
|
|
4
4
|
Summary: Tiny stdlib-only client that auto-instruments OpenAI/Anthropic/Gemini and emits LLM call + outcome economics to a Margin ingest API.
|
|
5
5
|
Author: Margin
|
|
6
6
|
License: MIT
|
|
@@ -452,6 +452,50 @@ current_task.set(ticket.topic) # "billing" | "password-reset" | …
|
|
|
452
452
|
Leave it off and the row's `task_key` is simply unset (`unknown`) — never a
|
|
453
453
|
fabricated bucket.
|
|
454
454
|
|
|
455
|
+
## Say where each call RAN (`environment`)
|
|
456
|
+
|
|
457
|
+
`environment` splits the work you sell from the work you did to decide what to
|
|
458
|
+
sell. A spike, an eval pass and your production traffic are three different kinds
|
|
459
|
+
of spend, and only one of them belongs in the number you report upward.
|
|
460
|
+
|
|
461
|
+
```python
|
|
462
|
+
meter.record_call(
|
|
463
|
+
workflow_id="support-triage",
|
|
464
|
+
provider="openai",
|
|
465
|
+
model="gpt-4o-mini",
|
|
466
|
+
input_tokens=812,
|
|
467
|
+
output_tokens=96,
|
|
468
|
+
environment="prod", # "prod" | "dev" | "eval"
|
|
469
|
+
)
|
|
470
|
+
```
|
|
471
|
+
|
|
472
|
+
Setting it once per process is usually easier. A deployment already knows whether
|
|
473
|
+
it is production; the call site does not:
|
|
474
|
+
|
|
475
|
+
```bash
|
|
476
|
+
export MARGIN_ENVIRONMENT=prod
|
|
477
|
+
```
|
|
478
|
+
|
|
479
|
+
An explicit argument wins over the variable, so one process can still label a
|
|
480
|
+
single call differently.
|
|
481
|
+
|
|
482
|
+
| value | send it for |
|
|
483
|
+
| --- | --- |
|
|
484
|
+
| `prod` | traffic your customers cause |
|
|
485
|
+
| `dev` | a spike, a local run, a branch deploy |
|
|
486
|
+
| `eval` | a graded pass, in CI or from an eval harness |
|
|
487
|
+
|
|
488
|
+
Why it earns a field. `cost_per_outcome` is total call spend divided by passed
|
|
489
|
+
outcomes. With `environment` unset, your R&D and your operations land in that one
|
|
490
|
+
number, so "your AI workforce costs $X per outcome" describes neither. Margin
|
|
491
|
+
already keeps its own face-off spend out of your number; this is how you get the
|
|
492
|
+
same split on your side.
|
|
493
|
+
|
|
494
|
+
Two things it will not do. It does not default: leave it off and the row is
|
|
495
|
+
`unknown`, which is honest and costs you nothing, where guessing `prod` would book
|
|
496
|
+
your spike as production. And it is never priced and never rolled into cost. It is
|
|
497
|
+
a label, like `team`.
|
|
498
|
+
|
|
455
499
|
## API
|
|
456
500
|
|
|
457
501
|
| Method | Emits to | Notes |
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|