revenium-python-sdk 0.1.3__tar.gz → 0.1.4__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (88) hide show
  1. {revenium_python_sdk-0.1.3/revenium_python_sdk.egg-info → revenium_python_sdk-0.1.4}/PKG-INFO +107 -1
  2. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/README.md +106 -0
  3. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/pyproject.toml +1 -1
  4. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/_core/__init__.py +7 -0
  5. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/_core/config.py +9 -0
  6. revenium_python_sdk-0.1.4/revenium_middleware/_core/enforcement.py +344 -0
  7. revenium_python_sdk-0.1.4/revenium_middleware/_core/exceptions.py +35 -0
  8. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/openai/__init__.py +3 -1
  9. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/openai/exceptions.py +5 -0
  10. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/openai/middleware.py +20 -0
  11. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4/revenium_python_sdk.egg-info}/PKG-INFO +107 -1
  12. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_python_sdk.egg-info/SOURCES.txt +2 -0
  13. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/LICENSE +0 -0
  14. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/__init__.py +0 -0
  15. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/_core/context.py +0 -0
  16. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/_core/decorators.py +0 -0
  17. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/_core/fields.py +0 -0
  18. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/_core/metering.py +0 -0
  19. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/_core/patch_registry.py +0 -0
  20. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/_core/prompt_extraction.py +0 -0
  21. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/_core/subscriber.py +0 -0
  22. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/_core/trace_fields.py +0 -0
  23. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/anthropic/__init__.py +0 -0
  24. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/anthropic/bedrock_adapter.py +0 -0
  25. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/anthropic/config.py +0 -0
  26. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/anthropic/middleware.py +0 -0
  27. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/anthropic/prompt_extractor.py +0 -0
  28. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/anthropic/provider.py +0 -0
  29. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/anthropic/summary_printer.py +0 -0
  30. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/anthropic/trace_fields.py +0 -0
  31. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/fal/__init__.py +0 -0
  32. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/fal/_metering.py +0 -0
  33. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/fal/config.py +0 -0
  34. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/fal/middleware.py +0 -0
  35. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/fal/trace_fields.py +0 -0
  36. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/google/__init__.py +0 -0
  37. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/google/common/__init__.py +0 -0
  38. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/google/common/exceptions.py +0 -0
  39. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/google/common/protocols.py +0 -0
  40. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/google/common/summary_printer.py +0 -0
  41. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/google/common/trace_fields.py +0 -0
  42. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/google/common/types.py +0 -0
  43. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/google/common/utils.py +0 -0
  44. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/google/config.py +0 -0
  45. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/google/google_ai/__init__.py +0 -0
  46. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/google/google_ai/middleware.py +0 -0
  47. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/google/google_ai/provider.py +0 -0
  48. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/google/prompt_extractor.py +0 -0
  49. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/google/vertex_ai/__init__.py +0 -0
  50. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/google/vertex_ai/middleware.py +0 -0
  51. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/google/vertex_ai/provider.py +0 -0
  52. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/litellm/__init__.py +0 -0
  53. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/litellm/client/__init__.py +0 -0
  54. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/litellm/client/config.py +0 -0
  55. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/litellm/client/context.py +0 -0
  56. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/litellm/client/decorators.py +0 -0
  57. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/litellm/client/hooks.py +0 -0
  58. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/litellm/client/integrations/__init__.py +0 -0
  59. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/litellm/client/integrations/crewai.py +0 -0
  60. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/litellm/client/middleware.py +0 -0
  61. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/litellm/client/summary_printer.py +0 -0
  62. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/litellm/client/trace_fields.py +0 -0
  63. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/litellm/client/validation.py +0 -0
  64. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/litellm/proxy/__init__.py +0 -0
  65. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/litellm/proxy/middleware.py +0 -0
  66. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/ollama/__init__.py +0 -0
  67. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/ollama/middleware.py +0 -0
  68. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/ollama/trace_fields.py +0 -0
  69. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/openai/azure_config.py +0 -0
  70. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/openai/azure_model_resolver.py +0 -0
  71. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/openai/config.py +0 -0
  72. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/openai/langchain/__init__.py +0 -0
  73. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/openai/langchain/_utils.py +0 -0
  74. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/openai/langchain/unified_handler.py +0 -0
  75. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/openai/prompt_extractor.py +0 -0
  76. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/openai/provider.py +0 -0
  77. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/openai/summary_printer.py +0 -0
  78. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/openai/trace_fields.py +0 -0
  79. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/perplexity/__init__.py +0 -0
  80. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/perplexity/middleware.py +0 -0
  81. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/perplexity/perplexity_sdk.py +0 -0
  82. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/perplexity/provider.py +0 -0
  83. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_middleware/perplexity/trace_fields.py +0 -0
  84. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_python_sdk.egg-info/dependency_links.txt +0 -0
  85. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_python_sdk.egg-info/requires.txt +0 -0
  86. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/revenium_python_sdk.egg-info/top_level.txt +0 -0
  87. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/setup.cfg +0 -0
  88. {revenium_python_sdk-0.1.3 → revenium_python_sdk-0.1.4}/tests/test_metering.py +0 -0
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: revenium-python-sdk
3
- Version: 0.1.3
3
+ Version: 0.1.4
4
4
  Summary: The official Revenium Python SDK — unified AI metering middleware for OpenAI, Anthropic, Google, Ollama, LiteLLM, Perplexity, and fal.ai.
5
5
  Author-email: Revenium <support@revenium.io>
6
6
  License: MIT
@@ -1136,6 +1136,106 @@ Trace ID: abc-123
1136
1136
 
1137
1137
  ---
1138
1138
 
1139
+ ## Cost Controls / Enforcement
1140
+
1141
+ Block outbound provider requests client-side when a Revenium cost-limit rule trips. When the circuit breaker is enabled, the middleware polls enforcement rules from the Revenium API in a background daemon thread and raises `ReveniumCostLimitExceeded` **before** the upstream call, preventing spend beyond the configured limit.
1142
+
1143
+ Currently wired for the OpenAI provider (other providers land via per-provider follow-on tickets).
1144
+
1145
+ ### Enable
1146
+
1147
+ ```bash
1148
+ pip install 'revenium-python-sdk[openai]'
1149
+ ```
1150
+
1151
+ ```env
1152
+ REVENIUM_CIRCUIT_BREAKER_ENABLED=true
1153
+ REVENIUM_METERING_API_KEY=hak_your_key_here
1154
+ REVENIUM_TEAM_ID=your_hashed_team_id
1155
+ REVENIUM_ENFORCEMENT_BASE_URL=https://api.revenium.ai/profitstream # optional
1156
+ ```
1157
+
1158
+ ### Environment Variables
1159
+
1160
+ | Variable | Default | Description |
1161
+ |----------|---------|-------------|
1162
+ | `REVENIUM_CIRCUIT_BREAKER_ENABLED` | `false` | Master switch. `true` / `1` / `yes` / `on` to enable. |
1163
+ | `REVENIUM_BYPASS` | `false` | When `true`, every `check_enforcement` call short-circuits to a no-op. Useful for incident response. |
1164
+ | `REVENIUM_TEAM_ID` | — | Hashed team ID. Path component on rule fetches; required when the breaker is enabled. |
1165
+ | `REVENIUM_ENFORCEMENT_BASE_URL` | origin of `REVENIUM_METERING_BASE_URL` | Base URL for the enforcement API. Set when the enforcement API lives behind a context-path. |
1166
+ | `REVENIUM_CB_POLL_INTERVAL_SECONDS` | `60` | Background poll interval for rule refreshes. |
1167
+ | `REVENIUM_CB_FAIL_MODE` | `open` | `open` (default) lets calls through when no cache exists; `closed` raises `ReveniumCostLimitExceeded` until rules are loaded. |
1168
+ | `REVENIUM_CACHE_DIR` | — | When set, the rule cache is mirrored to `<dir>/revenium_enforcement_rules.json` so a restarted process doesn't fail-closed on the very first call. |
1169
+
1170
+ ### Public API
1171
+
1172
+ Enforcement auto-initializes when the OpenAI middleware loads:
1173
+
1174
+ ```python
1175
+ import revenium_middleware.openai # auto-instruments openai
1176
+ import openai
1177
+
1178
+ client = openai.OpenAI()
1179
+ ```
1180
+
1181
+ The pre-call check fires before every chat / embeddings / responses call. When the circuit breaker is disabled, it is a no-op. When enabled:
1182
+
1183
+ 1. A daemon thread (`revenium-enforcement-poll`) starts on first use.
1184
+ 2. It polls `GET {REVENIUM_ENFORCEMENT_BASE_URL}/v2/api/ai/enforcement-rules/{REVENIUM_TEAM_ID}` every `REVENIUM_CB_POLL_INTERVAL_SECONDS` with the `x-api-key` header.
1185
+ 3. Rules are cached in-process (120 s TTL, refresh-on-stale with thundering-herd guard).
1186
+ 4. `204 No Content` is treated as "no rules configured" — the cache is cleared.
1187
+
1188
+ ### Exception Contract
1189
+
1190
+ ```python
1191
+ from revenium_middleware.openai import ReveniumCostLimitExceeded
1192
+ ```
1193
+
1194
+ When a tripped rule matches the current request, the middleware raises before the OpenAI call is made. All structured fields are populated when the server provides them:
1195
+
1196
+ | Attribute | Type | Description |
1197
+ |-----------|------|-------------|
1198
+ | `message` | `str` | Human-readable reason, e.g. `"Request blocked by Revenium enforcement rule: monthly-gpt4-cap"` |
1199
+ | `rule_name` | `str \| None` | Server-side rule name |
1200
+ | `current_value` | `float \| None` | Current metric value at the time of the block |
1201
+ | `threshold` | `float \| None` | Configured limit |
1202
+ | `resets_at` | `str \| None` | ISO-8601 timestamp the rule next resets |
1203
+ | `rule_id` | `str \| int \| None` | Server-side rule identifier |
1204
+
1205
+ `ReveniumCostLimitExceeded` does **not** inherit from `ReveniumMiddlewareError`, so the OpenAI middleware's `handle_exception_safely` decorator never swallows it — it always reaches your `except` block.
1206
+
1207
+ ```python
1208
+ from revenium_middleware.openai import ReveniumCostLimitExceeded
1209
+ import openai
1210
+
1211
+ client = openai.OpenAI()
1212
+
1213
+ try:
1214
+ response = client.chat.completions.create(
1215
+ model="gpt-4o-mini",
1216
+ messages=[{"role": "user", "content": "Summarize the meeting notes"}],
1217
+ )
1218
+ except ReveniumCostLimitExceeded as exc:
1219
+ print(f"Cost limit reached: {exc.message}")
1220
+ print(f"Rule {exc.rule_name}: {exc.current_value} / {exc.threshold}; resets {exc.resets_at}")
1221
+ ```
1222
+
1223
+ ### Fail-Open vs Fail-Closed
1224
+
1225
+ By default (`REVENIUM_CB_FAIL_MODE=open`) enforcement failures never propagate to user code. If the rule fetch errors (network, 5xx, auth), the previous in-memory cache is preserved and a debug log line is emitted. If there is no cache yet, enforcement behaves as if no rules are configured and the request continues.
1226
+
1227
+ Set `REVENIUM_CB_FAIL_MODE=closed` to refuse calls until at least one rule fetch (or `REVENIUM_CACHE_DIR` snapshot) succeeds. Pair with `REVENIUM_CACHE_DIR` so a process restart loads the last-known rules rather than blocking every call until the first poll completes.
1228
+
1229
+ ### Shadow Mode
1230
+
1231
+ Rules with `shadowMode: true` are observe-and-log: they are skipped by `check_enforcement`. Use shadow mode on the server side to audit a rule before flipping it to enforce.
1232
+
1233
+ ### End-to-End Example
1234
+
1235
+ See [`examples/openai/openai_blocking_demo.py`](examples/openai/openai_blocking_demo.py) for a runnable end-to-end demo using a seeded budget rule.
1236
+
1237
+ ---
1238
+
1139
1239
  ## Configuration Reference
1140
1240
 
1141
1241
  ### Required Environment Variables
@@ -1238,6 +1338,12 @@ Available log levels:
1238
1338
 
1239
1339
  For detailed documentation, visit [docs.revenium.io](https://docs.revenium.io)
1240
1340
 
1341
+ ### Server-Side Cost Controls
1342
+
1343
+ Cost controls (spend limits, throttling, alerts) are managed server-side in Revenium, not in this SDK. The SDK reports usage; Revenium evaluates it against your configured cost controls.
1344
+
1345
+ The cost-controls API endpoint is `/v2/api/ai/cost-controls`. This Python SDK does not call the endpoint directly — no SDK changes are required to use cost controls. If you manage cost controls via the Revenium API, HTTP client, or `curl`, see [docs.revenium.io](https://docs.revenium.io) for the current API reference.
1346
+
1241
1347
  ## Contributing
1242
1348
 
1243
1349
  See [CONTRIBUTING.md](./CONTRIBUTING.md)
@@ -1048,6 +1048,106 @@ Trace ID: abc-123
1048
1048
 
1049
1049
  ---
1050
1050
 
1051
+ ## Cost Controls / Enforcement
1052
+
1053
+ Block outbound provider requests client-side when a Revenium cost-limit rule trips. When the circuit breaker is enabled, the middleware polls enforcement rules from the Revenium API in a background daemon thread and raises `ReveniumCostLimitExceeded` **before** the upstream call, preventing spend beyond the configured limit.
1054
+
1055
+ Currently wired for the OpenAI provider (other providers land via per-provider follow-on tickets).
1056
+
1057
+ ### Enable
1058
+
1059
+ ```bash
1060
+ pip install 'revenium-python-sdk[openai]'
1061
+ ```
1062
+
1063
+ ```env
1064
+ REVENIUM_CIRCUIT_BREAKER_ENABLED=true
1065
+ REVENIUM_METERING_API_KEY=hak_your_key_here
1066
+ REVENIUM_TEAM_ID=your_hashed_team_id
1067
+ REVENIUM_ENFORCEMENT_BASE_URL=https://api.revenium.ai/profitstream # optional
1068
+ ```
1069
+
1070
+ ### Environment Variables
1071
+
1072
+ | Variable | Default | Description |
1073
+ |----------|---------|-------------|
1074
+ | `REVENIUM_CIRCUIT_BREAKER_ENABLED` | `false` | Master switch. `true` / `1` / `yes` / `on` to enable. |
1075
+ | `REVENIUM_BYPASS` | `false` | When `true`, every `check_enforcement` call short-circuits to a no-op. Useful for incident response. |
1076
+ | `REVENIUM_TEAM_ID` | — | Hashed team ID. Path component on rule fetches; required when the breaker is enabled. |
1077
+ | `REVENIUM_ENFORCEMENT_BASE_URL` | origin of `REVENIUM_METERING_BASE_URL` | Base URL for the enforcement API. Set when the enforcement API lives behind a context-path. |
1078
+ | `REVENIUM_CB_POLL_INTERVAL_SECONDS` | `60` | Background poll interval for rule refreshes. |
1079
+ | `REVENIUM_CB_FAIL_MODE` | `open` | `open` (default) lets calls through when no cache exists; `closed` raises `ReveniumCostLimitExceeded` until rules are loaded. |
1080
+ | `REVENIUM_CACHE_DIR` | — | When set, the rule cache is mirrored to `<dir>/revenium_enforcement_rules.json` so a restarted process doesn't fail-closed on the very first call. |
1081
+
1082
+ ### Public API
1083
+
1084
+ Enforcement auto-initializes when the OpenAI middleware loads:
1085
+
1086
+ ```python
1087
+ import revenium_middleware.openai # auto-instruments openai
1088
+ import openai
1089
+
1090
+ client = openai.OpenAI()
1091
+ ```
1092
+
1093
+ The pre-call check fires before every chat / embeddings / responses call. When the circuit breaker is disabled, it is a no-op. When enabled:
1094
+
1095
+ 1. A daemon thread (`revenium-enforcement-poll`) starts on first use.
1096
+ 2. It polls `GET {REVENIUM_ENFORCEMENT_BASE_URL}/v2/api/ai/enforcement-rules/{REVENIUM_TEAM_ID}` every `REVENIUM_CB_POLL_INTERVAL_SECONDS` with the `x-api-key` header.
1097
+ 3. Rules are cached in-process (120 s TTL, refresh-on-stale with thundering-herd guard).
1098
+ 4. `204 No Content` is treated as "no rules configured" — the cache is cleared.
1099
+
1100
+ ### Exception Contract
1101
+
1102
+ ```python
1103
+ from revenium_middleware.openai import ReveniumCostLimitExceeded
1104
+ ```
1105
+
1106
+ When a tripped rule matches the current request, the middleware raises before the OpenAI call is made. All structured fields are populated when the server provides them:
1107
+
1108
+ | Attribute | Type | Description |
1109
+ |-----------|------|-------------|
1110
+ | `message` | `str` | Human-readable reason, e.g. `"Request blocked by Revenium enforcement rule: monthly-gpt4-cap"` |
1111
+ | `rule_name` | `str \| None` | Server-side rule name |
1112
+ | `current_value` | `float \| None` | Current metric value at the time of the block |
1113
+ | `threshold` | `float \| None` | Configured limit |
1114
+ | `resets_at` | `str \| None` | ISO-8601 timestamp the rule next resets |
1115
+ | `rule_id` | `str \| int \| None` | Server-side rule identifier |
1116
+
1117
+ `ReveniumCostLimitExceeded` does **not** inherit from `ReveniumMiddlewareError`, so the OpenAI middleware's `handle_exception_safely` decorator never swallows it — it always reaches your `except` block.
1118
+
1119
+ ```python
1120
+ from revenium_middleware.openai import ReveniumCostLimitExceeded
1121
+ import openai
1122
+
1123
+ client = openai.OpenAI()
1124
+
1125
+ try:
1126
+ response = client.chat.completions.create(
1127
+ model="gpt-4o-mini",
1128
+ messages=[{"role": "user", "content": "Summarize the meeting notes"}],
1129
+ )
1130
+ except ReveniumCostLimitExceeded as exc:
1131
+ print(f"Cost limit reached: {exc.message}")
1132
+ print(f"Rule {exc.rule_name}: {exc.current_value} / {exc.threshold}; resets {exc.resets_at}")
1133
+ ```
1134
+
1135
+ ### Fail-Open vs Fail-Closed
1136
+
1137
+ By default (`REVENIUM_CB_FAIL_MODE=open`) enforcement failures never propagate to user code. If the rule fetch errors (network, 5xx, auth), the previous in-memory cache is preserved and a debug log line is emitted. If there is no cache yet, enforcement behaves as if no rules are configured and the request continues.
1138
+
1139
+ Set `REVENIUM_CB_FAIL_MODE=closed` to refuse calls until at least one rule fetch (or `REVENIUM_CACHE_DIR` snapshot) succeeds. Pair with `REVENIUM_CACHE_DIR` so a process restart loads the last-known rules rather than blocking every call until the first poll completes.
1140
+
1141
+ ### Shadow Mode
1142
+
1143
+ Rules with `shadowMode: true` are observe-and-log: they are skipped by `check_enforcement`. Use shadow mode on the server side to audit a rule before flipping it to enforce.
1144
+
1145
+ ### End-to-End Example
1146
+
1147
+ See [`examples/openai/openai_blocking_demo.py`](examples/openai/openai_blocking_demo.py) for a runnable end-to-end demo using a seeded budget rule.
1148
+
1149
+ ---
1150
+
1051
1151
  ## Configuration Reference
1052
1152
 
1053
1153
  ### Required Environment Variables
@@ -1150,6 +1250,12 @@ Available log levels:
1150
1250
 
1151
1251
  For detailed documentation, visit [docs.revenium.io](https://docs.revenium.io)
1152
1252
 
1253
+ ### Server-Side Cost Controls
1254
+
1255
+ Cost controls (spend limits, throttling, alerts) are managed server-side in Revenium, not in this SDK. The SDK reports usage; Revenium evaluates it against your configured cost controls.
1256
+
1257
+ The cost-controls API endpoint is `/v2/api/ai/cost-controls`. This Python SDK does not call the endpoint directly — no SDK changes are required to use cost controls. If you manage cost controls via the Revenium API, HTTP client, or `curl`, see [docs.revenium.io](https://docs.revenium.io) for the current API reference.
1258
+
1153
1259
  ## Contributing
1154
1260
 
1155
1261
  See [CONTRIBUTING.md](./CONTRIBUTING.md)
@@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta"
4
4
 
5
5
  [project]
6
6
  name = "revenium-python-sdk"
7
- version = "0.1.3"
7
+ version = "0.1.4"
8
8
  description = "The official Revenium Python SDK — unified AI metering middleware for OpenAI, Anthropic, Google, Ollama, LiteLLM, Perplexity, and fal.ai."
9
9
  readme = "README.md"
10
10
  requires-python = ">=3.8"
@@ -6,6 +6,8 @@ provider-specific middleware implementations.
6
6
  """
7
7
 
8
8
  from .metering import run_async_in_thread, shutdown_event, client
9
+ from .exceptions import ReveniumCostLimitExceeded
10
+ from .enforcement import check_enforcement, is_circuit_breaker_enabled, stop_polling
9
11
  from .context import (
10
12
  is_inside_decorated_function,
11
13
  get_function_metadata,
@@ -46,6 +48,11 @@ __all__ = [
46
48
  "client",
47
49
  "run_async_in_thread",
48
50
  "shutdown_event",
51
+ # Enforcement / circuit breaker
52
+ "ReveniumCostLimitExceeded",
53
+ "check_enforcement",
54
+ "is_circuit_breaker_enabled",
55
+ "stop_polling",
49
56
  # Decorators
50
57
  "revenium_meter",
51
58
  "revenium_metadata",
@@ -74,6 +74,15 @@ class Config:
74
74
  SUMMARY_API_TIMEOUT: float = 5.0
75
75
  DEFAULT_BASE_URL: str = "https://api.revenium.ai"
76
76
 
77
+ # Enforcement / circuit breaker — see _core/enforcement.py.
78
+ # Taxonomy locked in BACK-1154 docs; keep in sync across SDKs.
79
+ ENV_REVENIUM_BYPASS: str = "REVENIUM_BYPASS"
80
+ ENV_CIRCUIT_BREAKER_ENABLED: str = "REVENIUM_CIRCUIT_BREAKER_ENABLED"
81
+ ENV_REVENIUM_ENFORCEMENT_BASE_URL: str = "REVENIUM_ENFORCEMENT_BASE_URL"
82
+ ENV_REVENIUM_CB_POLL_INTERVAL_SECONDS: str = "REVENIUM_CB_POLL_INTERVAL_SECONDS"
83
+ ENV_REVENIUM_CB_FAIL_MODE: str = "REVENIUM_CB_FAIL_MODE"
84
+ ENV_REVENIUM_CACHE_DIR: str = "REVENIUM_CACHE_DIR"
85
+
77
86
 
78
87
  class SecurityConfig:
79
88
  """Security-related configuration shared across all providers."""
@@ -0,0 +1,344 @@
1
+ """
2
+ Enforcement engine for the Revenium circuit breaker.
3
+
4
+ Polls cost-limit rules from the Revenium API in a daemon thread and caches
5
+ them in memory. ``check_enforcement(...)`` is a pre-call hook that raises
6
+ ``ReveniumCostLimitExceeded`` when a tripped rule matches the current
7
+ request, blocking the outbound provider call before any spend occurs.
8
+
9
+ Opt-in via ``REVENIUM_CIRCUIT_BREAKER_ENABLED``. Disabled by default so the
10
+ SDK stays no-op for callers who haven't enrolled in cost controls.
11
+ """
12
+
13
+ import json
14
+ import logging
15
+ import os
16
+ import threading
17
+ import time
18
+ from typing import List, Optional, Tuple
19
+ from urllib.parse import quote, urlparse
20
+
21
+ import httpx
22
+
23
+ from .config import Config
24
+ from .exceptions import ReveniumCostLimitExceeded
25
+
26
+ logger = logging.getLogger("revenium_middleware.extension")
27
+
28
+ _DEFAULT_POLL_INTERVAL = 60 # seconds between background refreshes
29
+ _CACHE_TTL = 120 # seconds before a cached rule is considered stale
30
+ _RULES_CACHE_FILENAME = "revenium_enforcement_rules.json"
31
+
32
+ _cached_rules: List[dict] = []
33
+ _cache_lock = threading.Lock()
34
+ _cache_timestamp = 0.0
35
+ # True once any successful fetch (even an empty list / HTTP 204) or a disk
36
+ # snapshot load has populated the cache. Distinguishes "server says no rules
37
+ # apply" from "we have never heard back from the server" — fail-closed must
38
+ # only block in the latter case.
39
+ _cache_initialized = False
40
+
41
+ _poll_thread: Optional[threading.Thread] = None
42
+ _poll_lock = threading.Lock()
43
+ _stop_event = threading.Event()
44
+
45
+ # Serializes synchronous stale-cache refreshes to prevent thundering herd
46
+ _refresh_lock = threading.Lock()
47
+
48
+ # Single-shot warnings so misconfigured environments don't spam logs
49
+ _team_id_warned = False
50
+ _disk_load_attempted = False
51
+
52
+
53
+ def _env_truthy(name: str) -> bool:
54
+ return os.environ.get(name, "").lower() in ("1", "true", "yes", "on")
55
+
56
+
57
+ def is_circuit_breaker_enabled() -> bool:
58
+ """Return True when the operator has opted in to enforcement."""
59
+ return _env_truthy(Config.ENV_CIRCUIT_BREAKER_ENABLED)
60
+
61
+
62
+ def is_bypass_enabled() -> bool:
63
+ """``REVENIUM_BYPASS=true`` short-circuits enforcement at every callsite."""
64
+ return _env_truthy(Config.ENV_REVENIUM_BYPASS)
65
+
66
+
67
+ def _poll_interval_seconds() -> int:
68
+ raw = os.environ.get(Config.ENV_REVENIUM_CB_POLL_INTERVAL_SECONDS, "")
69
+ if not raw:
70
+ return _DEFAULT_POLL_INTERVAL
71
+ try:
72
+ value = int(raw)
73
+ return value if value > 0 else _DEFAULT_POLL_INTERVAL
74
+ except ValueError:
75
+ logger.debug("Invalid %s=%r, using default", Config.ENV_REVENIUM_CB_POLL_INTERVAL_SECONDS, raw)
76
+ return _DEFAULT_POLL_INTERVAL
77
+
78
+
79
+ def _fail_mode_is_closed() -> bool:
80
+ """``REVENIUM_CB_FAIL_MODE=closed`` raises when no usable cache exists."""
81
+ return os.environ.get(Config.ENV_REVENIUM_CB_FAIL_MODE, "open").lower() == "closed"
82
+
83
+
84
+ def _cache_file_path() -> Optional[str]:
85
+ cache_dir = os.environ.get(Config.ENV_REVENIUM_CACHE_DIR, "")
86
+ if not cache_dir:
87
+ return None
88
+ return os.path.join(cache_dir, _RULES_CACHE_FILENAME)
89
+
90
+
91
+ def _get_enforcement_base_url() -> str:
92
+ """Base URL for enforcement API calls.
93
+
94
+ Prefers ``REVENIUM_ENFORCEMENT_BASE_URL`` so a context-path
95
+ (``http://localhost:8080/profitstream``) survives intact. Falls back to
96
+ the origin of the metering URL when unset.
97
+ """
98
+ explicit = os.environ.get(Config.ENV_REVENIUM_ENFORCEMENT_BASE_URL, "")
99
+ if explicit:
100
+ return explicit.rstrip("/")
101
+ metering_url = os.environ.get(Config.ENV_REVENIUM_BASE_URL, "https://api.revenium.ai/meter/")
102
+ parsed = urlparse(metering_url)
103
+ return f"{parsed.scheme}://{parsed.netloc}"
104
+
105
+
106
+ def _load_cache_from_disk() -> None:
107
+ global _cached_rules, _cache_timestamp, _disk_load_attempted, _cache_initialized
108
+ # Check-and-set under _cache_lock so two cold-start callers can't both
109
+ # observe _disk_load_attempted=False and race on the snapshot read.
110
+ with _cache_lock:
111
+ if _disk_load_attempted:
112
+ return
113
+ _disk_load_attempted = True
114
+ path = _cache_file_path()
115
+ if not path or not os.path.exists(path):
116
+ return
117
+ try:
118
+ with open(path, "r", encoding="utf-8") as handle:
119
+ data = json.load(handle)
120
+ if isinstance(data, list):
121
+ with _cache_lock:
122
+ _cached_rules = data
123
+ # Treat as stale so the next call still triggers a refresh,
124
+ # but the disk snapshot prevents fail-closed from raising on
125
+ # the very first request after a process restart.
126
+ _cache_timestamp = 0.0
127
+ _cache_initialized = True
128
+ logger.debug("Loaded %d enforcement rule(s) from %s", len(data), path)
129
+ except Exception:
130
+ logger.debug("Failed to read enforcement cache from %s", path, exc_info=True)
131
+
132
+
133
+ def _persist_cache_to_disk(rules: list) -> None:
134
+ path = _cache_file_path()
135
+ if not path:
136
+ return
137
+ try:
138
+ os.makedirs(os.path.dirname(path), exist_ok=True)
139
+ with open(path, "w", encoding="utf-8") as handle:
140
+ json.dump(rules, handle)
141
+ except Exception:
142
+ logger.debug("Failed to write enforcement cache to %s", path, exc_info=True)
143
+
144
+
145
+ def _fetch_rules() -> Optional[list]:
146
+ """Fetch current enforcement rules from the Revenium API.
147
+
148
+ Returns a list (possibly empty) on success, or ``None`` on failure so
149
+ the caller can preserve the previous cache.
150
+ """
151
+ global _team_id_warned
152
+
153
+ api_key = os.environ.get(Config.ENV_REVENIUM_API_KEY, "")
154
+ if not api_key:
155
+ logger.debug("No API key configured, skipping enforcement rule fetch")
156
+ return None
157
+
158
+ team_id = os.environ.get(Config.ENV_REVENIUM_TEAM_ID, "")
159
+ if not team_id:
160
+ if not _team_id_warned:
161
+ logger.warning(
162
+ "REVENIUM_TEAM_ID is not set — enforcement rule polling disabled. "
163
+ "Set this to your hashed team ID to enable cost-limit enforcement."
164
+ )
165
+ _team_id_warned = True
166
+ return None
167
+
168
+ base_url = _get_enforcement_base_url()
169
+ # Percent-encode the team_id path segment so a misconfigured value
170
+ # containing '/', '..', or '?' cannot retarget the request to a
171
+ # different endpoint on the same origin.
172
+ safe_team_id = quote(team_id, safe="")
173
+ try:
174
+ response = httpx.get(
175
+ f"{base_url}/v2/api/ai/enforcement-rules/{safe_team_id}",
176
+ headers={"x-api-key": api_key},
177
+ timeout=10,
178
+ )
179
+ # 204 No Content == no rules configured for this team; cache empty list
180
+ if response.status_code == 204:
181
+ return []
182
+ response.raise_for_status()
183
+ data = response.json()
184
+ # Server currently returns ``{"rules": [...], "compiledAt": ...}`` but
185
+ # accept a bare list too so a future schema change does not silently
186
+ # AttributeError its way into a stale cache.
187
+ if isinstance(data, list):
188
+ return data
189
+ if isinstance(data, dict):
190
+ return data.get("rules", [])
191
+ logger.warning("Unexpected enforcement response shape: %r", type(data).__name__)
192
+ return None
193
+ except Exception:
194
+ logger.debug("Failed to fetch enforcement rules, falling open", exc_info=True)
195
+ return None
196
+
197
+
198
+ def _refresh_cache() -> None:
199
+ """Refresh the in-memory rule cache.
200
+
201
+ Only advances ``_cache_timestamp`` on a successful fetch — a transient
202
+ network error must not poison the stale-cache trigger and silently
203
+ suppress retries for the next ``_CACHE_TTL`` window. Disk persistence
204
+ happens outside ``_cache_lock`` so a slow filesystem write can't block
205
+ concurrent ``check_enforcement`` callers on the pre-call path.
206
+ """
207
+ global _cached_rules, _cache_timestamp, _cache_initialized
208
+ rules = _fetch_rules()
209
+ if rules is None:
210
+ return
211
+ with _cache_lock:
212
+ _cached_rules = rules
213
+ _cache_timestamp = time.monotonic()
214
+ _cache_initialized = True
215
+ _persist_cache_to_disk(rules)
216
+
217
+
218
+ def _poll_loop() -> None:
219
+ interval = _poll_interval_seconds()
220
+ while not _stop_event.is_set():
221
+ _refresh_cache()
222
+ _stop_event.wait(interval)
223
+
224
+
225
+ def _ensure_poller_running() -> None:
226
+ global _poll_thread
227
+ with _poll_lock:
228
+ if _poll_thread is not None and _poll_thread.is_alive():
229
+ return
230
+ _stop_event.clear()
231
+ _poll_thread = threading.Thread(
232
+ target=_poll_loop,
233
+ name="revenium-enforcement-poll",
234
+ daemon=True,
235
+ )
236
+ _poll_thread.start()
237
+
238
+
239
+ def _get_rules() -> Tuple[list, bool]:
240
+ """Return cached rules and the initialized flag as a single snapshot.
241
+
242
+ Reading both fields under one ``_cache_lock`` acquisition prevents the
243
+ fail-closed path from torn-reading ``_cache_initialized`` against a
244
+ half-written ``_cached_rules`` from the background poller.
245
+ """
246
+ now = time.monotonic()
247
+ with _cache_lock:
248
+ age = now - _cache_timestamp
249
+ rules = list(_cached_rules)
250
+ initialized = _cache_initialized
251
+ if age > _CACHE_TTL:
252
+ if _refresh_lock.acquire(blocking=False):
253
+ try:
254
+ _refresh_cache()
255
+ with _cache_lock:
256
+ rules = list(_cached_rules)
257
+ initialized = _cache_initialized
258
+ finally:
259
+ _refresh_lock.release()
260
+ return rules, initialized
261
+
262
+
263
+ def _coerce_float(value) -> Optional[float]:
264
+ if value is None:
265
+ return None
266
+ try:
267
+ return float(value)
268
+ except (TypeError, ValueError):
269
+ return None
270
+
271
+
272
+ def check_enforcement(usage_metadata: Optional[dict] = None) -> None:
273
+ """Pre-call enforcement check.
274
+
275
+ Invoke before the upstream provider call. No-op when the circuit breaker
276
+ is disabled or no rules are tripped.
277
+
278
+ Raises:
279
+ ReveniumCostLimitExceeded: when a cost-limit rule blocks the call.
280
+ All structured fields (``rule_name``, ``current_value``,
281
+ ``threshold``, ``resets_at``, ``rule_id``) are populated when the
282
+ server provides them.
283
+ """
284
+ if is_bypass_enabled():
285
+ return
286
+ if not is_circuit_breaker_enabled():
287
+ return
288
+
289
+ _load_cache_from_disk()
290
+ _ensure_poller_running()
291
+ rules, initialized = _get_rules()
292
+
293
+ # Fail-closed mode: only block when the cache has *never* loaded. An
294
+ # empty list from a successful fetch (HTTP 204 = "no rules apply") is a
295
+ # valid initialized state and must pass through. Use the snapshot taken
296
+ # under _cache_lock so the decision can't see a torn write.
297
+ if _fail_mode_is_closed() and not initialized:
298
+ raise ReveniumCostLimitExceeded(
299
+ "Request blocked: enforcement cache is uninitialized and "
300
+ "REVENIUM_CB_FAIL_MODE=closed."
301
+ )
302
+
303
+ credential = (usage_metadata or {}).get("subscriber_credential", "")
304
+
305
+ for rule in rules:
306
+ if not isinstance(rule, dict):
307
+ continue
308
+ # A rule is tripped when the server marks it as ``breached`` (current
309
+ # nucleus schema) or ``blocked`` (legacy). Skip ``shadowMode`` rules
310
+ # even when breached — those are observe-and-log only.
311
+ is_tripped = rule.get("breached", False) or rule.get("blocked", False)
312
+ if not is_tripped:
313
+ continue
314
+ if rule.get("shadowMode", False):
315
+ continue
316
+
317
+ rule_credential = rule.get("credential", "")
318
+ if rule_credential and rule_credential != credential:
319
+ continue
320
+
321
+ rule_name = rule.get("name", "cost limit")
322
+ raise ReveniumCostLimitExceeded(
323
+ message=f"Request blocked by Revenium enforcement rule: {rule_name}",
324
+ rule_name=rule_name,
325
+ current_value=_coerce_float(rule.get("currentValue")),
326
+ threshold=_coerce_float(rule.get("threshold")),
327
+ resets_at=rule.get("resetsAt"),
328
+ rule_id=rule.get("ruleId") or rule.get("id"),
329
+ )
330
+
331
+
332
+ def stop_polling() -> None:
333
+ """Gracefully stop the background polling thread.
334
+
335
+ Reads ``_poll_thread`` under ``_poll_lock`` to avoid a TOCTOU race with
336
+ ``_ensure_poller_running`` spinning the thread up on a concurrent first
337
+ request — without the lock, shutdown could observe ``None`` and skip the
338
+ ``join`` even though a poller is alive.
339
+ """
340
+ _stop_event.set()
341
+ with _poll_lock:
342
+ thread = _poll_thread
343
+ if thread is not None:
344
+ thread.join(timeout=5)
@@ -0,0 +1,35 @@
1
+ """
2
+ Core exceptions shared across all Revenium middleware providers.
3
+
4
+ The unified SDK ships ``ReveniumCostLimitExceeded`` from ``_core`` so every
5
+ provider subpackage (openai, anthropic, google, …) raises the same exception
6
+ type and downstream callers can ``except ReveniumCostLimitExceeded`` once.
7
+ """
8
+
9
+ from typing import Optional, Union
10
+
11
+
12
+ class ReveniumCostLimitExceeded(Exception):
13
+ """Raised when a Revenium enforcement rule blocks the outbound request.
14
+
15
+ Inherits directly from ``Exception`` (not from any middleware-error base)
16
+ so per-provider ``handle_exception_safely`` decorators never swallow it —
17
+ enforcement must always reach the caller.
18
+ """
19
+
20
+ def __init__(
21
+ self,
22
+ message: str,
23
+ rule_name: Optional[str] = None,
24
+ current_value: Optional[float] = None,
25
+ threshold: Optional[float] = None,
26
+ resets_at: Optional[str] = None,
27
+ rule_id: Optional[Union[str, int]] = None,
28
+ ):
29
+ super().__init__(message)
30
+ self.message = message
31
+ self.rule_name = rule_name
32
+ self.current_value = current_value
33
+ self.threshold = threshold
34
+ self.resets_at = resets_at
35
+ self.rule_id = rule_id
@@ -10,6 +10,8 @@ Usage:
10
10
  """
11
11
  import logging
12
12
 
13
+ from revenium_middleware._core.exceptions import ReveniumCostLimitExceeded
14
+
13
15
  logger = logging.getLogger(__name__)
14
16
 
15
17
  # Conditionally import middleware (requires wrapt + openai SDK)
@@ -20,4 +22,4 @@ except ImportError:
20
22
  logger.debug("OpenAI middleware dependencies not available, middleware not loaded")
21
23
  create_wrapper = None # type: ignore
22
24
 
23
- __all__ = ["create_wrapper"]
25
+ __all__ = ["create_wrapper", "ReveniumCostLimitExceeded"]
@@ -5,6 +5,8 @@ This module defines a hierarchy of exceptions that provide better error handling
5
5
  and more specific error information for different failure scenarios.
6
6
  """
7
7
 
8
+ from revenium_middleware._core.exceptions import ReveniumCostLimitExceeded # noqa: F401
9
+
8
10
 
9
11
  class ReveniumMiddlewareError(Exception):
10
12
  """Base exception for all Revenium middleware errors."""
@@ -70,6 +72,9 @@ def handle_exception_safely(func):
70
72
  def wrapper(*args, **kwargs):
71
73
  try:
72
74
  return func(*args, **kwargs)
75
+ except ReveniumCostLimitExceeded:
76
+ # Enforcement exceptions must reach the caller — never swallow.
77
+ raise
73
78
  except ReveniumMiddlewareError as e:
74
79
  # Log middleware-specific errors
75
80
  import logging
@@ -7,6 +7,8 @@ from enum import Enum
7
7
 
8
8
  import wrapt
9
9
  from revenium_middleware import client, run_async_in_thread, shutdown_event, merge_metadata
10
+ from revenium_middleware._core.enforcement import check_enforcement
11
+ from revenium_middleware._core.exceptions import ReveniumCostLimitExceeded # noqa: F401 — re-exported via openai.exceptions
10
12
  from revenium_middleware._core.subscriber import extract_subscriber_from_metadata
11
13
  from revenium_middleware._core.fields import extract_org_and_product, extract_common_metadata, extract_agentic_job_fields, merge_extra_body
12
14
  from revenium_middleware._core.config import is_selective_metering_enabled, is_capture_prompts_enabled
@@ -804,6 +806,10 @@ def embeddings_create_wrapper(wrapped, instance, args, kwargs):
804
806
 
805
807
  # Record request time
806
808
  request_time_dt = datetime.datetime.now(datetime.timezone.utc)
809
+
810
+ # Enforcement pre-call check — may raise ReveniumCostLimitExceeded
811
+ check_enforcement(usage_metadata)
812
+
807
813
  logger.debug(
808
814
  f"Calling wrapped embeddings function with args: {args}, "
809
815
  f"kwargs: {kwargs}"
@@ -890,6 +896,10 @@ def create_wrapper(wrapped, instance, args, kwargs):
890
896
  azure_config.validate_deployment()
891
897
 
892
898
  request_time_dt = datetime.datetime.now(datetime.timezone.utc)
899
+
900
+ # Enforcement pre-call check — may raise ReveniumCostLimitExceeded
901
+ check_enforcement(usage_metadata)
902
+
893
903
  logger.debug(
894
904
  f"Calling wrapped function with args: {args}, kwargs: {kwargs}"
895
905
  )
@@ -1239,6 +1249,10 @@ def responses_create_wrapper(wrapped, instance, args, kwargs):
1239
1249
 
1240
1250
  # Record request time
1241
1251
  request_time_dt = datetime.datetime.now(datetime.timezone.utc)
1252
+
1253
+ # Enforcement pre-call check — may raise ReveniumCostLimitExceeded
1254
+ check_enforcement(usage_metadata)
1255
+
1242
1256
  logger.debug(f"Calling wrapped responses function with args: {args}, kwargs: {kwargs}")
1243
1257
 
1244
1258
  # Call the original OpenAI function
@@ -1640,6 +1654,9 @@ def async_create_wrapper(wrapped, instance, args, kwargs):
1640
1654
 
1641
1655
  request_time_dt = datetime.datetime.now(datetime.timezone.utc)
1642
1656
 
1657
+ # Enforcement pre-call check — may raise ReveniumCostLimitExceeded
1658
+ check_enforcement(usage_metadata)
1659
+
1643
1660
  async def _async_invoke():
1644
1661
  response = await wrapped(*args, **kwargs)
1645
1662
 
@@ -1703,6 +1720,9 @@ def async_embeddings_create_wrapper(wrapped, instance, args, kwargs):
1703
1720
 
1704
1721
  request_time_dt = datetime.datetime.now(datetime.timezone.utc)
1705
1722
 
1723
+ # Enforcement pre-call check — may raise ReveniumCostLimitExceeded
1724
+ check_enforcement(usage_metadata)
1725
+
1706
1726
  async def _async_invoke():
1707
1727
  response = await wrapped(*args, **kwargs)
1708
1728
 
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: revenium-python-sdk
3
- Version: 0.1.3
3
+ Version: 0.1.4
4
4
  Summary: The official Revenium Python SDK — unified AI metering middleware for OpenAI, Anthropic, Google, Ollama, LiteLLM, Perplexity, and fal.ai.
5
5
  Author-email: Revenium <support@revenium.io>
6
6
  License: MIT
@@ -1136,6 +1136,106 @@ Trace ID: abc-123
1136
1136
 
1137
1137
  ---
1138
1138
 
1139
+ ## Cost Controls / Enforcement
1140
+
1141
+ Block outbound provider requests client-side when a Revenium cost-limit rule trips. When the circuit breaker is enabled, the middleware polls enforcement rules from the Revenium API in a background daemon thread and raises `ReveniumCostLimitExceeded` **before** the upstream call, preventing spend beyond the configured limit.
1142
+
1143
+ Currently wired for the OpenAI provider (other providers land via per-provider follow-on tickets).
1144
+
1145
+ ### Enable
1146
+
1147
+ ```bash
1148
+ pip install 'revenium-python-sdk[openai]'
1149
+ ```
1150
+
1151
+ ```env
1152
+ REVENIUM_CIRCUIT_BREAKER_ENABLED=true
1153
+ REVENIUM_METERING_API_KEY=hak_your_key_here
1154
+ REVENIUM_TEAM_ID=your_hashed_team_id
1155
+ REVENIUM_ENFORCEMENT_BASE_URL=https://api.revenium.ai/profitstream # optional
1156
+ ```
1157
+
1158
+ ### Environment Variables
1159
+
1160
+ | Variable | Default | Description |
1161
+ |----------|---------|-------------|
1162
+ | `REVENIUM_CIRCUIT_BREAKER_ENABLED` | `false` | Master switch. `true` / `1` / `yes` / `on` to enable. |
1163
+ | `REVENIUM_BYPASS` | `false` | When `true`, every `check_enforcement` call short-circuits to a no-op. Useful for incident response. |
1164
+ | `REVENIUM_TEAM_ID` | — | Hashed team ID. Path component on rule fetches; required when the breaker is enabled. |
1165
+ | `REVENIUM_ENFORCEMENT_BASE_URL` | origin of `REVENIUM_METERING_BASE_URL` | Base URL for the enforcement API. Set when the enforcement API lives behind a context-path. |
1166
+ | `REVENIUM_CB_POLL_INTERVAL_SECONDS` | `60` | Background poll interval for rule refreshes. |
1167
+ | `REVENIUM_CB_FAIL_MODE` | `open` | `open` (default) lets calls through when no cache exists; `closed` raises `ReveniumCostLimitExceeded` until rules are loaded. |
1168
+ | `REVENIUM_CACHE_DIR` | — | When set, the rule cache is mirrored to `<dir>/revenium_enforcement_rules.json` so a restarted process doesn't fail-closed on the very first call. |
1169
+
1170
+ ### Public API
1171
+
1172
+ Enforcement auto-initializes when the OpenAI middleware loads:
1173
+
1174
+ ```python
1175
+ import revenium_middleware.openai # auto-instruments openai
1176
+ import openai
1177
+
1178
+ client = openai.OpenAI()
1179
+ ```
1180
+
1181
+ The pre-call check fires before every chat / embeddings / responses call. When the circuit breaker is disabled, it is a no-op. When enabled:
1182
+
1183
+ 1. A daemon thread (`revenium-enforcement-poll`) starts on first use.
1184
+ 2. It polls `GET {REVENIUM_ENFORCEMENT_BASE_URL}/v2/api/ai/enforcement-rules/{REVENIUM_TEAM_ID}` every `REVENIUM_CB_POLL_INTERVAL_SECONDS` with the `x-api-key` header.
1185
+ 3. Rules are cached in-process (120 s TTL, refresh-on-stale with thundering-herd guard).
1186
+ 4. `204 No Content` is treated as "no rules configured" — the cache is cleared.
1187
+
1188
+ ### Exception Contract
1189
+
1190
+ ```python
1191
+ from revenium_middleware.openai import ReveniumCostLimitExceeded
1192
+ ```
1193
+
1194
+ When a tripped rule matches the current request, the middleware raises before the OpenAI call is made. All structured fields are populated when the server provides them:
1195
+
1196
+ | Attribute | Type | Description |
1197
+ |-----------|------|-------------|
1198
+ | `message` | `str` | Human-readable reason, e.g. `"Request blocked by Revenium enforcement rule: monthly-gpt4-cap"` |
1199
+ | `rule_name` | `str \| None` | Server-side rule name |
1200
+ | `current_value` | `float \| None` | Current metric value at the time of the block |
1201
+ | `threshold` | `float \| None` | Configured limit |
1202
+ | `resets_at` | `str \| None` | ISO-8601 timestamp the rule next resets |
1203
+ | `rule_id` | `str \| int \| None` | Server-side rule identifier |
1204
+
1205
+ `ReveniumCostLimitExceeded` does **not** inherit from `ReveniumMiddlewareError`, so the OpenAI middleware's `handle_exception_safely` decorator never swallows it — it always reaches your `except` block.
1206
+
1207
+ ```python
1208
+ from revenium_middleware.openai import ReveniumCostLimitExceeded
1209
+ import openai
1210
+
1211
+ client = openai.OpenAI()
1212
+
1213
+ try:
1214
+ response = client.chat.completions.create(
1215
+ model="gpt-4o-mini",
1216
+ messages=[{"role": "user", "content": "Summarize the meeting notes"}],
1217
+ )
1218
+ except ReveniumCostLimitExceeded as exc:
1219
+ print(f"Cost limit reached: {exc.message}")
1220
+ print(f"Rule {exc.rule_name}: {exc.current_value} / {exc.threshold}; resets {exc.resets_at}")
1221
+ ```
1222
+
1223
+ ### Fail-Open vs Fail-Closed
1224
+
1225
+ By default (`REVENIUM_CB_FAIL_MODE=open`) enforcement failures never propagate to user code. If the rule fetch errors (network, 5xx, auth), the previous in-memory cache is preserved and a debug log line is emitted. If there is no cache yet, enforcement behaves as if no rules are configured and the request continues.
1226
+
1227
+ Set `REVENIUM_CB_FAIL_MODE=closed` to refuse calls until at least one rule fetch (or `REVENIUM_CACHE_DIR` snapshot) succeeds. Pair with `REVENIUM_CACHE_DIR` so a process restart loads the last-known rules rather than blocking every call until the first poll completes.
1228
+
1229
+ ### Shadow Mode
1230
+
1231
+ Rules with `shadowMode: true` are observe-and-log: they are skipped by `check_enforcement`. Use shadow mode on the server side to audit a rule before flipping it to enforce.
1232
+
1233
+ ### End-to-End Example
1234
+
1235
+ See [`examples/openai/openai_blocking_demo.py`](examples/openai/openai_blocking_demo.py) for a runnable end-to-end demo using a seeded budget rule.
1236
+
1237
+ ---
1238
+
1139
1239
  ## Configuration Reference
1140
1240
 
1141
1241
  ### Required Environment Variables
@@ -1238,6 +1338,12 @@ Available log levels:
1238
1338
 
1239
1339
  For detailed documentation, visit [docs.revenium.io](https://docs.revenium.io)
1240
1340
 
1341
+ ### Server-Side Cost Controls
1342
+
1343
+ Cost controls (spend limits, throttling, alerts) are managed server-side in Revenium, not in this SDK. The SDK reports usage; Revenium evaluates it against your configured cost controls.
1344
+
1345
+ The cost-controls API endpoint is `/v2/api/ai/cost-controls`. This Python SDK does not call the endpoint directly — no SDK changes are required to use cost controls. If you manage cost controls via the Revenium API, HTTP client, or `curl`, see [docs.revenium.io](https://docs.revenium.io) for the current API reference.
1346
+
1241
1347
  ## Contributing
1242
1348
 
1243
1349
  See [CONTRIBUTING.md](./CONTRIBUTING.md)
@@ -6,6 +6,8 @@ revenium_middleware/_core/__init__.py
6
6
  revenium_middleware/_core/config.py
7
7
  revenium_middleware/_core/context.py
8
8
  revenium_middleware/_core/decorators.py
9
+ revenium_middleware/_core/enforcement.py
10
+ revenium_middleware/_core/exceptions.py
9
11
  revenium_middleware/_core/fields.py
10
12
  revenium_middleware/_core/metering.py
11
13
  revenium_middleware/_core/patch_registry.py