shadowshield 0.6.2__tar.gz → 0.7.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (147) hide show
  1. {shadowshield-0.6.2 → shadowshield-0.7.0}/PKG-INFO +83 -12
  2. {shadowshield-0.6.2 → shadowshield-0.7.0}/README.md +82 -10
  3. shadowshield-0.7.0/docs/AUDIT_REPORT_2026-08-05.md +190 -0
  4. {shadowshield-0.6.2 → shadowshield-0.7.0}/docs/BENCHMARKS.md +35 -9
  5. {shadowshield-0.6.2 → shadowshield-0.7.0}/docs/PRODUCTION_READINESS.md +35 -11
  6. {shadowshield-0.6.2 → shadowshield-0.7.0}/docs/RELEASING.md +46 -9
  7. shadowshield-0.7.0/docs/UPGRADE_PLAN_2026-08-05.md +114 -0
  8. {shadowshield-0.6.2 → shadowshield-0.7.0}/docs/configuration.md +7 -5
  9. {shadowshield-0.6.2 → shadowshield-0.7.0}/docs/index.md +2 -0
  10. {shadowshield-0.6.2 → shadowshield-0.7.0}/docs/plugins.md +3 -1
  11. {shadowshield-0.6.2 → shadowshield-0.7.0}/docs/security-model.md +4 -2
  12. shadowshield-0.7.0/examples/gateway_mode.py +46 -0
  13. shadowshield-0.7.0/examples/rag_guard.py +42 -0
  14. shadowshield-0.7.0/examples/streaming_scan.py +50 -0
  15. {shadowshield-0.6.2 → shadowshield-0.7.0}/pyproject.toml +11 -3
  16. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/__init__.py +3 -1
  17. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/_security.py +112 -21
  18. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/cli.py +154 -23
  19. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/config/default.yaml +4 -0
  20. shadowshield-0.7.0/src/shadowshield/control/__init__.py +60 -0
  21. shadowshield-0.7.0/src/shadowshield/control/__main__.py +5 -0
  22. shadowshield-0.7.0/src/shadowshield/control/app.py +465 -0
  23. shadowshield-0.7.0/src/shadowshield/control/migrate.py +134 -0
  24. shadowshield-0.7.0/src/shadowshield/control/models.py +34 -0
  25. shadowshield-0.7.0/src/shadowshield/control/policy_state.py +270 -0
  26. shadowshield-0.7.0/src/shadowshield/control/state.py +655 -0
  27. shadowshield-0.7.0/src/shadowshield/core/calibration.py +201 -0
  28. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/core/config.py +13 -0
  29. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/core/engine.py +111 -13
  30. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/core/policy.py +15 -2
  31. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/core/session.py +34 -19
  32. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/core/shield.py +15 -0
  33. shadowshield-0.7.0/src/shadowshield/core/stream.py +159 -0
  34. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/detectors/anomaly.py +2 -1
  35. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/detectors/base.py +2 -0
  36. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/detectors/exfiltration.py +6 -0
  37. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/detectors/vector.py +83 -1
  38. shadowshield-0.7.0/src/shadowshield/eval/data/generalization_benchmark_v4.jsonl +58 -0
  39. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/eval/dataset.py +7 -6
  40. shadowshield-0.7.0/src/shadowshield/eval/harness.py +432 -0
  41. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/middleware/__init__.py +18 -0
  42. shadowshield-0.7.0/src/shadowshield/middleware/anthropic.py +153 -0
  43. shadowshield-0.7.0/src/shadowshield/middleware/asgi.py +261 -0
  44. shadowshield-0.7.0/src/shadowshield/middleware/litellm.py +125 -0
  45. shadowshield-0.7.0/src/shadowshield/middleware/rag.py +221 -0
  46. shadowshield-0.7.0/src/shadowshield/proxy.py +512 -0
  47. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/reporter.py +33 -1
  48. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/responders/sanitizer.py +14 -4
  49. shadowshield-0.7.0/tests/test_calibration.py +130 -0
  50. {shadowshield-0.6.2 → shadowshield-0.7.0}/tests/test_cli.py +58 -0
  51. {shadowshield-0.6.2 → shadowshield-0.7.0}/tests/test_config.py +2 -0
  52. {shadowshield-0.6.2 → shadowshield-0.7.0}/tests/test_control.py +138 -1
  53. {shadowshield-0.6.2 → shadowshield-0.7.0}/tests/test_detectors.py +17 -0
  54. shadowshield-0.7.0/tests/test_eval_harness.py +240 -0
  55. {shadowshield-0.6.2 → shadowshield-0.7.0}/tests/test_generalization_benchmarks.py +47 -1
  56. shadowshield-0.7.0/tests/test_http_security.py +326 -0
  57. shadowshield-0.7.0/tests/test_langchain_middleware.py +91 -0
  58. shadowshield-0.7.0/tests/test_middleware_breadth.py +285 -0
  59. shadowshield-0.7.0/tests/test_parallel.py +161 -0
  60. shadowshield-0.7.0/tests/test_plugins.py +133 -0
  61. {shadowshield-0.6.2 → shadowshield-0.7.0}/tests/test_policy.py +45 -9
  62. {shadowshield-0.6.2 → shadowshield-0.7.0}/tests/test_production.py +55 -0
  63. shadowshield-0.7.0/tests/test_proxy.py +256 -0
  64. shadowshield-0.7.0/tests/test_rag.py +138 -0
  65. {shadowshield-0.6.2 → shadowshield-0.7.0}/tests/test_responders.py +40 -0
  66. {shadowshield-0.6.2 → shadowshield-0.7.0}/tests/test_shield.py +49 -0
  67. {shadowshield-0.6.2 → shadowshield-0.7.0}/tests/test_site_security.py +13 -0
  68. shadowshield-0.7.0/tests/test_stream.py +136 -0
  69. {shadowshield-0.6.2 → shadowshield-0.7.0}/tests/test_telemetry.py +23 -0
  70. {shadowshield-0.6.2 → shadowshield-0.7.0}/tests/test_vector.py +41 -0
  71. shadowshield-0.7.0/tests/test_workflow_policy.py +149 -0
  72. shadowshield-0.6.2/src/shadowshield/control.py +0 -1399
  73. shadowshield-0.6.2/src/shadowshield/eval/harness.py +0 -206
  74. shadowshield-0.6.2/tests/test_http_security.py +0 -138
  75. {shadowshield-0.6.2 → shadowshield-0.7.0}/.gitignore +0 -0
  76. {shadowshield-0.6.2 → shadowshield-0.7.0}/LICENSE +0 -0
  77. {shadowshield-0.6.2 → shadowshield-0.7.0}/docs/CODE_REVIEW.md +0 -0
  78. {shadowshield-0.6.2 → shadowshield-0.7.0}/docs/COMPARISON.md +0 -0
  79. {shadowshield-0.6.2 → shadowshield-0.7.0}/docs/MARKET_LANDSCAPE.md +0 -0
  80. {shadowshield-0.6.2 → shadowshield-0.7.0}/docs/NEXT_STEPS.md +0 -0
  81. {shadowshield-0.6.2 → shadowshield-0.7.0}/docs/OWASP_LLM_TOP10.md +0 -0
  82. {shadowshield-0.6.2 → shadowshield-0.7.0}/docs/PLAN_REVIEW.md +0 -0
  83. {shadowshield-0.6.2 → shadowshield-0.7.0}/docs/REPORTER_SDK_SPEC.md +0 -0
  84. {shadowshield-0.6.2 → shadowshield-0.7.0}/docs/SAAS_STRATEGY.md +0 -0
  85. {shadowshield-0.6.2 → shadowshield-0.7.0}/docs/UPGRADE_OPPORTUNITIES.md +0 -0
  86. {shadowshield-0.6.2 → shadowshield-0.7.0}/docs/detectors.md +0 -0
  87. {shadowshield-0.6.2 → shadowshield-0.7.0}/docs/research/LANDSCAPE.md +0 -0
  88. {shadowshield-0.6.2 → shadowshield-0.7.0}/examples/agentic_security.py +0 -0
  89. {shadowshield-0.6.2 → shadowshield-0.7.0}/examples/custom_detector.py +0 -0
  90. {shadowshield-0.6.2 → shadowshield-0.7.0}/examples/langchain_integration.py +0 -0
  91. {shadowshield-0.6.2 → shadowshield-0.7.0}/examples/openai_integration.py +0 -0
  92. {shadowshield-0.6.2 → shadowshield-0.7.0}/examples/quickstart.py +0 -0
  93. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/config/__init__.py +0 -0
  94. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/core/__init__.py +0 -0
  95. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/core/canary.py +0 -0
  96. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/core/coverage.py +0 -0
  97. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/core/telemetry.py +0 -0
  98. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/core/types.py +0 -0
  99. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/detectors/__init__.py +0 -0
  100. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/detectors/alignment.py +0 -0
  101. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/detectors/canary.py +0 -0
  102. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/detectors/data/attack_corpus.txt +0 -0
  103. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/detectors/encoding.py +0 -0
  104. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/detectors/jailbreak.py +0 -0
  105. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/detectors/llm_check.py +0 -0
  106. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/detectors/pii.py +0 -0
  107. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/detectors/prompt_injection.py +0 -0
  108. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/detectors/transformer.py +0 -0
  109. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/eval/__init__.py +0 -0
  110. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/eval/data/adversarial_benchmark.jsonl +0 -0
  111. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/eval/data/builtin_benchmark.jsonl +0 -0
  112. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/eval/data/generalization_benchmark_v1.jsonl +0 -0
  113. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/eval/data/generalization_benchmark_v2.jsonl +0 -0
  114. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/eval/data/generalization_benchmark_v3.jsonl +0 -0
  115. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/integrations/__init__.py +0 -0
  116. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/integrations/agentdojo.py +0 -0
  117. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/integrations/mcp.py +0 -0
  118. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/middleware/base.py +0 -0
  119. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/middleware/decorators.py +0 -0
  120. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/middleware/langchain.py +0 -0
  121. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/middleware/openai.py +0 -0
  122. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/plugins/__init__.py +0 -0
  123. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/plugins/base.py +0 -0
  124. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/plugins/manager.py +0 -0
  125. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/py.typed +0 -0
  126. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/responders/__init__.py +0 -0
  127. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/responders/base.py +0 -0
  128. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/responders/blocker.py +0 -0
  129. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/responders/isolator.py +0 -0
  130. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/responders/rate_limiter.py +0 -0
  131. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/server.py +0 -0
  132. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/static/dashboard.html +0 -0
  133. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/utils/__init__.py +0 -0
  134. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/utils/logging.py +0 -0
  135. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/utils/scoring.py +0 -0
  136. {shadowshield-0.6.2 → shadowshield-0.7.0}/src/shadowshield/utils/text.py +0 -0
  137. {shadowshield-0.6.2 → shadowshield-0.7.0}/tests/__init__.py +0 -0
  138. {shadowshield-0.6.2 → shadowshield-0.7.0}/tests/test_adversarial.py +0 -0
  139. {shadowshield-0.6.2 → shadowshield-0.7.0}/tests/test_agentic.py +0 -0
  140. {shadowshield-0.6.2 → shadowshield-0.7.0}/tests/test_mcp_guard.py +0 -0
  141. {shadowshield-0.6.2 → shadowshield-0.7.0}/tests/test_middleware.py +0 -0
  142. {shadowshield-0.6.2 → shadowshield-0.7.0}/tests/test_multilingual.py +0 -0
  143. {shadowshield-0.6.2 → shadowshield-0.7.0}/tests/test_pii_backends.py +0 -0
  144. {shadowshield-0.6.2 → shadowshield-0.7.0}/tests/test_prompt_injection.py +0 -0
  145. {shadowshield-0.6.2 → shadowshield-0.7.0}/tests/test_server.py +0 -0
  146. {shadowshield-0.6.2 → shadowshield-0.7.0}/tests/test_spans.py +0 -0
  147. {shadowshield-0.6.2 → shadowshield-0.7.0}/tests/test_transformer.py +0 -0
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: shadowshield
3
- Version: 0.6.2
3
+ Version: 0.7.0
4
4
  Summary: Unified open-source security shield for agentic AI systems — inspired by Sentinel & ShadowClaw.
5
5
  Project-URL: Homepage, https://shadowshield.xyz
6
6
  Project-URL: Documentation, https://github.com/0xsl1m/shadowshield#readme
@@ -28,7 +28,6 @@ Requires-Dist: httpx>=0.27
28
28
  Requires-Dist: pydantic>=2.5
29
29
  Requires-Dist: pyyaml>=6.0
30
30
  Requires-Dist: structlog>=24.1
31
- Requires-Dist: tiktoken>=0.6
32
31
  Provides-Extra: all
33
32
  Requires-Dist: datasets>=2.18; extra == 'all'
34
33
  Requires-Dist: fastapi>=0.110; extra == 'all'
@@ -135,8 +134,10 @@ print(result.safe_text) # safe fallback message
135
134
  false-positive rate on hard negatives** in the bundled benchmark.
136
135
  - **Proven, reproducibly.** Ships an eval harness + offline benchmark:
137
136
  `shadowshield benchmark`. Loads public datasets (PINT/deepset/InjecAgent) too.
138
- - **Drop-in integrations.** OpenAI-compatible clients, LangChain, decorators,
139
- context managers, **async** (`ascan`). Or call `shield.scan()` directly.
137
+ - **Drop-in integrations.** A zero-code **reverse-proxy gateway** for any
138
+ OpenAI-compatible endpoint, plus OpenAI/Anthropic/LiteLLM/LangChain wrappers,
139
+ pure-ASGI middleware, RAG retriever guards, decorators, context managers, and
140
+ **async** (`ascan`). Or call `shield.scan()` directly.
140
141
  - **Extensible & lightweight.** Add a detector/responder in ~10 lines or ship a
141
142
  plugin. Tiny core dependency set; ML/PII/datasets are optional extras.
142
143
 
@@ -149,9 +150,10 @@ print(result.safe_text) # safe fallback message
149
150
  > in-distribution **regression baseline, not a SOTA claim**. We publish the humbling
150
151
  > external numbers on purpose — a credible security tool shows its homework.
151
152
  > The frozen blind semantic snapshots are harder still: v1 reaches 26.7% recall /
152
- > 13.3% FPR, v2 reaches 0% / 10%, and v3 reaches 30% / 30%; the 90-row aggregate
153
- > is 22.2% / 20%. Run `shadowshield benchmark --generalization all`; these gaps
154
- > are public by design.
153
+ > 13.3% FPR, v2 reaches 0% / 10%, v3 reaches 30% / 30%, and the 58-example v4
154
+ > snapshot (indirect tool-result, multilingual, and semantic-pretext attacks)
155
+ > reaches 58.6% / 6.9%; the 148-row aggregate is 36.5% / 14.9%. Run
156
+ > `shadowshield benchmark --generalization all`; these gaps are public by design.
155
157
 
156
158
  ---
157
159
 
@@ -207,8 +209,8 @@ pip install "shadowshield[dashboard]" # + FastAPI HTTP server & dashboard
207
209
  pip install "shadowshield[all]" # everything
208
210
  ```
209
211
 
210
- Core deps are intentionally small: `pydantic`, `structlog`, `pyyaml`, `httpx`,
211
- `tiktoken`. The ML classifier, Presidio PII, dataset loaders, and dashboard live
212
+ Core deps are intentionally small: `pydantic`, `structlog`, `pyyaml`, and
213
+ `httpx`. The ML classifier, Presidio PII, dataset loaders, and dashboard live
212
214
  behind extras — the default install pulls **no** heavy ML stack.
213
215
 
214
216
  ---
@@ -297,7 +299,9 @@ shadowshield init > shield.yaml # write an annotated default config
297
299
  shadowshield benchmark # run the bundled offline benchmark
298
300
  shadowshield benchmark --adversarial
299
301
  shadowshield benchmark --generalization all # all frozen blind semantic snapshots
302
+ shadowshield calibrate --input benchmark.json --output calibration.json # isotonic score calibration
300
303
  shadowshield serve # HTTP server + live dashboard (needs [dashboard])
304
+ shadowshield proxy --upstream https://api.openai.com --port 8100 # gateway mode (needs [dashboard])
301
305
  ```
302
306
 
303
307
  ### 9. HTTP server (any language / a browser dashboard)
@@ -317,6 +321,62 @@ Endpoints: `GET /health` (liveness), `GET /ready` (readiness), `POST /scan`,
317
321
  `api_keys` is supplied; local-only trusted embeddings must explicitly pass
318
322
  `allow_insecure_local=True`.
319
323
 
324
+ ### 10. Gateway mode — guardrails without code changes
325
+
326
+ Put ShadowShield *in front of* any OpenAI-compatible endpoint and point your
327
+ existing SDK at the proxy instead. Chat messages are scanned pre-flight (a
328
+ blocked request returns an OpenAI-style `403` and never reaches the upstream),
329
+ completions are scanned post-flight, and malicious **SSE streams are cut
330
+ mid-flight** with a conventional `finish_reason="content_filter"` chunk:
331
+
332
+ ```bash
333
+ pip install "shadowshield[dashboard]"
334
+ shadowshield proxy --upstream https://api.openai.com --port 8100 \
335
+ --api-key "$GATEWAY_KEY" # proxy auth; upstream key stays in Authorization
336
+ ```
337
+
338
+ ```python
339
+ client = OpenAI(base_url="http://localhost:8100/v1", api_key=os.environ["OPENAI_API_KEY"])
340
+ ```
341
+
342
+ The same protection embeds in-process via `ShieldASGIMiddleware`
343
+ (`shadowshield.middleware.asgi`) for any ASGI app — SSE passes through there,
344
+ so use the proxy when you stream.
345
+
346
+ ### 11. Streaming completions — cut the stream mid-flight
347
+
348
+ `StreamScanner` scans a completion *while it streams* (bounded memory,
349
+ carry-over window so split signatures still match) and returns a terminal
350
+ verdict the moment the stream must stop:
351
+
352
+ ```python
353
+ scanner = shield.stream_scanner(scan_interval_chars=256)
354
+ for event in openai_stream: # any streaming SDK
355
+ terminal = scanner.feed(event.delta_text)
356
+ if terminal is not None:
357
+ break # close the stream NOW
358
+ final = scanner.finalize()
359
+ ```
360
+
361
+ ### 12. Anthropic, LiteLLM, and RAG pipelines
362
+
363
+ ```python
364
+ from shadowshield.middleware import (
365
+ ShieldedAnthropicClient, shielded_completion, scan_retrieved_chunks,
366
+ )
367
+
368
+ anthropic_client = ShieldedAnthropicClient(Anthropic(), shield) # messages.create guarded
369
+ completion = shielded_completion(shield) # wraps litellm.completion
370
+
371
+ # Retrieved chunks are untrusted: drop poisoned documents before the prompt.
372
+ report = scan_retrieved_chunks(shield, retrieved_chunks, on_threat="drop")
373
+ prompt = build_prompt(query, report.safe_chunks)
374
+ ```
375
+
376
+ Duck-typed wrappers for LlamaIndex (`ShieldedLlamaIndexRetriever`) and Haystack
377
+ (`ShieldedHaystackRetriever`) filter poisoned nodes/documents inside each
378
+ framework's retriever contract.
379
+
320
380
  ### Production container
321
381
 
322
382
  The included container runs the full control plane as a non-root user with a
@@ -330,7 +390,7 @@ export SHADOWSHIELD_ADMIN_KEY="$(openssl rand -hex 32)"
330
390
  export SHADOWSHIELD_POLICY_KEY="$(openssl rand -hex 32)"
331
391
  export SHADOWSHIELD_POLICY_STATE_KEY="$(openssl rand -hex 32)"
332
392
  export SHADOWSHIELD_IMAGE_DIGEST="$(curl -fsSL \
333
- https://github.com/0xsl1m/shadowshield/releases/download/v0.6.2/container-digest.txt)"
393
+ https://github.com/0xsl1m/shadowshield/releases/download/v0.7.0/container-digest.txt)"
334
394
  docker compose pull
335
395
  docker compose up -d
336
396
  ```
@@ -345,20 +405,31 @@ it beyond localhost. See the
345
405
  [production-readiness roadmap](docs/PRODUCTION_READINESS.md) for launch gates,
346
406
  known scale limits, and the operator checklist.
347
407
 
408
+ **Kubernetes:** `deploy/helm/shadowshield/` ships a chart with the same
409
+ hardening (digest-pinned image — install fails without it, read-only root
410
+ filesystem, all capabilities dropped, non-root, seccomp, bounded `/tmp`,
411
+ external secrets, policy-state PVC, health probes):
412
+
413
+ ```bash
414
+ helm install shadowshield deploy/helm/shadowshield \
415
+ --set image.digest=sha256:<release-digest> \
416
+ --set secrets.existingSecret=shadowshield-secrets
417
+ ```
418
+
348
419
  Upgrading a control-plane volume from 0.6.0 requires an offline re-key because
349
420
  0.6.0 authenticated durable state with the policy-signing key:
350
421
 
351
422
  ```bash
352
423
  # Load the existing scan/admin keys first so the new Compose file can resolve.
353
424
  export SHADOWSHIELD_IMAGE_DIGEST="$(curl -fsSL \
354
- https://github.com/0xsl1m/shadowshield/releases/download/v0.6.2/container-digest.txt)"
425
+ https://github.com/0xsl1m/shadowshield/releases/download/v0.7.0/container-digest.txt)"
355
426
  export SHADOWSHIELD_POLICY_KEY="<existing-0.6.0-policy-key>"
356
427
  export SHADOWSHIELD_POLICY_STATE_KEY="$(openssl rand -hex 32)"
357
428
  # Stop every writer and snapshot the volume before running the migration.
358
429
  docker compose stop shadowshield
359
430
  docker compose run --rm --no-deps shadowshield \
360
431
  shadowshield migrate-policy-state --path /var/lib/shadowshield/policy-state.json
361
- # Preserve the reported .pre-0.6.1.bak file, then start 0.6.2.
432
+ # Preserve the reported .pre-0.6.1.bak file, then start 0.7.0.
362
433
  docker compose up -d
363
434
  ```
364
435
 
@@ -62,8 +62,10 @@ print(result.safe_text) # safe fallback message
62
62
  false-positive rate on hard negatives** in the bundled benchmark.
63
63
  - **Proven, reproducibly.** Ships an eval harness + offline benchmark:
64
64
  `shadowshield benchmark`. Loads public datasets (PINT/deepset/InjecAgent) too.
65
- - **Drop-in integrations.** OpenAI-compatible clients, LangChain, decorators,
66
- context managers, **async** (`ascan`). Or call `shield.scan()` directly.
65
+ - **Drop-in integrations.** A zero-code **reverse-proxy gateway** for any
66
+ OpenAI-compatible endpoint, plus OpenAI/Anthropic/LiteLLM/LangChain wrappers,
67
+ pure-ASGI middleware, RAG retriever guards, decorators, context managers, and
68
+ **async** (`ascan`). Or call `shield.scan()` directly.
67
69
  - **Extensible & lightweight.** Add a detector/responder in ~10 lines or ship a
68
70
  plugin. Tiny core dependency set; ML/PII/datasets are optional extras.
69
71
 
@@ -76,9 +78,10 @@ print(result.safe_text) # safe fallback message
76
78
  > in-distribution **regression baseline, not a SOTA claim**. We publish the humbling
77
79
  > external numbers on purpose — a credible security tool shows its homework.
78
80
  > The frozen blind semantic snapshots are harder still: v1 reaches 26.7% recall /
79
- > 13.3% FPR, v2 reaches 0% / 10%, and v3 reaches 30% / 30%; the 90-row aggregate
80
- > is 22.2% / 20%. Run `shadowshield benchmark --generalization all`; these gaps
81
- > are public by design.
81
+ > 13.3% FPR, v2 reaches 0% / 10%, v3 reaches 30% / 30%, and the 58-example v4
82
+ > snapshot (indirect tool-result, multilingual, and semantic-pretext attacks)
83
+ > reaches 58.6% / 6.9%; the 148-row aggregate is 36.5% / 14.9%. Run
84
+ > `shadowshield benchmark --generalization all`; these gaps are public by design.
82
85
 
83
86
  ---
84
87
 
@@ -134,8 +137,8 @@ pip install "shadowshield[dashboard]" # + FastAPI HTTP server & dashboard
134
137
  pip install "shadowshield[all]" # everything
135
138
  ```
136
139
 
137
- Core deps are intentionally small: `pydantic`, `structlog`, `pyyaml`, `httpx`,
138
- `tiktoken`. The ML classifier, Presidio PII, dataset loaders, and dashboard live
140
+ Core deps are intentionally small: `pydantic`, `structlog`, `pyyaml`, and
141
+ `httpx`. The ML classifier, Presidio PII, dataset loaders, and dashboard live
139
142
  behind extras — the default install pulls **no** heavy ML stack.
140
143
 
141
144
  ---
@@ -224,7 +227,9 @@ shadowshield init > shield.yaml # write an annotated default config
224
227
  shadowshield benchmark # run the bundled offline benchmark
225
228
  shadowshield benchmark --adversarial
226
229
  shadowshield benchmark --generalization all # all frozen blind semantic snapshots
230
+ shadowshield calibrate --input benchmark.json --output calibration.json # isotonic score calibration
227
231
  shadowshield serve # HTTP server + live dashboard (needs [dashboard])
232
+ shadowshield proxy --upstream https://api.openai.com --port 8100 # gateway mode (needs [dashboard])
228
233
  ```
229
234
 
230
235
  ### 9. HTTP server (any language / a browser dashboard)
@@ -244,6 +249,62 @@ Endpoints: `GET /health` (liveness), `GET /ready` (readiness), `POST /scan`,
244
249
  `api_keys` is supplied; local-only trusted embeddings must explicitly pass
245
250
  `allow_insecure_local=True`.
246
251
 
252
+ ### 10. Gateway mode — guardrails without code changes
253
+
254
+ Put ShadowShield *in front of* any OpenAI-compatible endpoint and point your
255
+ existing SDK at the proxy instead. Chat messages are scanned pre-flight (a
256
+ blocked request returns an OpenAI-style `403` and never reaches the upstream),
257
+ completions are scanned post-flight, and malicious **SSE streams are cut
258
+ mid-flight** with a conventional `finish_reason="content_filter"` chunk:
259
+
260
+ ```bash
261
+ pip install "shadowshield[dashboard]"
262
+ shadowshield proxy --upstream https://api.openai.com --port 8100 \
263
+ --api-key "$GATEWAY_KEY" # proxy auth; upstream key stays in Authorization
264
+ ```
265
+
266
+ ```python
267
+ client = OpenAI(base_url="http://localhost:8100/v1", api_key=os.environ["OPENAI_API_KEY"])
268
+ ```
269
+
270
+ The same protection embeds in-process via `ShieldASGIMiddleware`
271
+ (`shadowshield.middleware.asgi`) for any ASGI app — SSE passes through there,
272
+ so use the proxy when you stream.
273
+
274
+ ### 11. Streaming completions — cut the stream mid-flight
275
+
276
+ `StreamScanner` scans a completion *while it streams* (bounded memory,
277
+ carry-over window so split signatures still match) and returns a terminal
278
+ verdict the moment the stream must stop:
279
+
280
+ ```python
281
+ scanner = shield.stream_scanner(scan_interval_chars=256)
282
+ for event in openai_stream: # any streaming SDK
283
+ terminal = scanner.feed(event.delta_text)
284
+ if terminal is not None:
285
+ break # close the stream NOW
286
+ final = scanner.finalize()
287
+ ```
288
+
289
+ ### 12. Anthropic, LiteLLM, and RAG pipelines
290
+
291
+ ```python
292
+ from shadowshield.middleware import (
293
+ ShieldedAnthropicClient, shielded_completion, scan_retrieved_chunks,
294
+ )
295
+
296
+ anthropic_client = ShieldedAnthropicClient(Anthropic(), shield) # messages.create guarded
297
+ completion = shielded_completion(shield) # wraps litellm.completion
298
+
299
+ # Retrieved chunks are untrusted: drop poisoned documents before the prompt.
300
+ report = scan_retrieved_chunks(shield, retrieved_chunks, on_threat="drop")
301
+ prompt = build_prompt(query, report.safe_chunks)
302
+ ```
303
+
304
+ Duck-typed wrappers for LlamaIndex (`ShieldedLlamaIndexRetriever`) and Haystack
305
+ (`ShieldedHaystackRetriever`) filter poisoned nodes/documents inside each
306
+ framework's retriever contract.
307
+
247
308
  ### Production container
248
309
 
249
310
  The included container runs the full control plane as a non-root user with a
@@ -257,7 +318,7 @@ export SHADOWSHIELD_ADMIN_KEY="$(openssl rand -hex 32)"
257
318
  export SHADOWSHIELD_POLICY_KEY="$(openssl rand -hex 32)"
258
319
  export SHADOWSHIELD_POLICY_STATE_KEY="$(openssl rand -hex 32)"
259
320
  export SHADOWSHIELD_IMAGE_DIGEST="$(curl -fsSL \
260
- https://github.com/0xsl1m/shadowshield/releases/download/v0.6.2/container-digest.txt)"
321
+ https://github.com/0xsl1m/shadowshield/releases/download/v0.7.0/container-digest.txt)"
261
322
  docker compose pull
262
323
  docker compose up -d
263
324
  ```
@@ -272,20 +333,31 @@ it beyond localhost. See the
272
333
  [production-readiness roadmap](docs/PRODUCTION_READINESS.md) for launch gates,
273
334
  known scale limits, and the operator checklist.
274
335
 
336
+ **Kubernetes:** `deploy/helm/shadowshield/` ships a chart with the same
337
+ hardening (digest-pinned image — install fails without it, read-only root
338
+ filesystem, all capabilities dropped, non-root, seccomp, bounded `/tmp`,
339
+ external secrets, policy-state PVC, health probes):
340
+
341
+ ```bash
342
+ helm install shadowshield deploy/helm/shadowshield \
343
+ --set image.digest=sha256:<release-digest> \
344
+ --set secrets.existingSecret=shadowshield-secrets
345
+ ```
346
+
275
347
  Upgrading a control-plane volume from 0.6.0 requires an offline re-key because
276
348
  0.6.0 authenticated durable state with the policy-signing key:
277
349
 
278
350
  ```bash
279
351
  # Load the existing scan/admin keys first so the new Compose file can resolve.
280
352
  export SHADOWSHIELD_IMAGE_DIGEST="$(curl -fsSL \
281
- https://github.com/0xsl1m/shadowshield/releases/download/v0.6.2/container-digest.txt)"
353
+ https://github.com/0xsl1m/shadowshield/releases/download/v0.7.0/container-digest.txt)"
282
354
  export SHADOWSHIELD_POLICY_KEY="<existing-0.6.0-policy-key>"
283
355
  export SHADOWSHIELD_POLICY_STATE_KEY="$(openssl rand -hex 32)"
284
356
  # Stop every writer and snapshot the volume before running the migration.
285
357
  docker compose stop shadowshield
286
358
  docker compose run --rm --no-deps shadowshield \
287
359
  shadowshield migrate-policy-state --path /var/lib/shadowshield/policy-state.json
288
- # Preserve the reported .pre-0.6.1.bak file, then start 0.6.2.
360
+ # Preserve the reported .pre-0.6.1.bak file, then start 0.7.0.
289
361
  docker compose up -d
290
362
  ```
291
363
 
@@ -0,0 +1,190 @@
1
+ # ShadowShield — Comprehensive Audit Report
2
+
3
+ **Project:** shadowshield v0.6.3 (`C:\Users\jhwil\Documents\OpenClaw VPS\shadowshield`)
4
+ **Audit date:** 2026-08-05 · **Auditor:** Kimi Work (automated code audit)
5
+ **Scope:** full repository — source, tests, packaging, CI/CD, container, docs, dependency & secret hygiene
6
+
7
+ ---
8
+
9
+ ## 1. Executive summary
10
+
11
+ ShadowShield is a unified open-source security shield for agentic AI systems
12
+ (prompt-injection / jailbreak / PII / exfiltration detection with a
13
+ detect → decide → respond engine, plus an optional FastAPI control plane).
14
+
15
+ **Overall verdict: strong.** This is an unusually disciplined codebase for a
16
+ 0.x project: the full test suite passes, strict mypy and Ruff are clean, the
17
+ threat model is documented and consistently implemented, and the release
18
+ pipeline is among the most hardened seen in open-source Python (hash-locked
19
+ builds, reproducible-build gate, pinned actions, Trivy gates, OIDC publishing,
20
+ fail-closed container startup tests). Findings are limited to one
21
+ medium-severity API design issue and a set of low-severity gaps.
22
+
23
+ **Verification results from this audit run (2026-08-05):**
24
+
25
+ | Check | Result |
26
+ |---|---|
27
+ | Test suite (`pytest tests`) | **320 passed, 2 skipped** (opt-in real-model tests), 38.8 s |
28
+ | Ruff lint | **Clean** (`All checks passed!`) |
29
+ | mypy (strict, 54 source files) | **Clean** (`no issues found`) |
30
+ | Coverage | **86%** total (branch), gate `fail_under = 80` met |
31
+ | pip-audit (installed environment) | **No known vulnerabilities** |
32
+ | Dangerous-sink scan (`eval/exec/pickle/subprocess/shell=True/yaml.load`) | **None in `src/`** |
33
+ | Committed-secret scan (AWS/GitHub/OpenAI key patterns, private keys) | **None** (2 hits are intentional test fixtures) |
34
+ | Git working tree | Clean; 146 tracked files; build artifacts correctly ignored |
35
+
36
+ ---
37
+
38
+ ## 2. Strengths confirmed
39
+
40
+ ### 2.1 Detection core
41
+ - **ReDoS-conscious regexes.** Signature patterns use bounded quantifiers
42
+ (`[\w\s,'’]{0,40}?` etc.); multilingual regex groups are pre-filtered by
43
+ cheap vocabulary cues before running (`_candidate_signatures`,
44
+ `prompt_injection.py:487`). Exfiltration detector explicitly documents a
45
+ rewrite away from a quadratic pattern.
46
+ - **Defense in depth.** Normalization (zero-width, homoglyph, bidi stripping),
47
+ base64/hex payload decoding with re-scan and severity bump, canary tokens for
48
+ detecting *successful* injections, optional transformer/vector/LLM-judge
49
+ layers that are lazy-imported and never pulled in by the base install.
50
+ - **Bounded outputs.** `MAX_FINDINGS_PER_DETECTOR`, truncated matched text
51
+ (`m.group(0)[:160]`), content-free telemetry — detection failures cannot
52
+ amplify into memory or leak content into logs.
53
+
54
+ ### 2.2 HTTP control plane (`_security.py`, `control.py`, `server.py`)
55
+ - **Fail-closed by default:** the app factory raises if API/admin keys are
56
+ missing; insecure mode requires an explicit `allow_insecure_local=True`.
57
+ - **Early authentication** before body reads; constant-time key comparison via
58
+ `hmac.compare_digest` (byte-safe for non-ASCII input).
59
+ - **Intake hardening:** 1 MiB body cap, 8,192-frame cap, single 15 s total read
60
+ deadline (defeats slow/chunked-body starvation of admission slots),
61
+ 503-with-`Retry-After` concurrency cap (16) on scan paths, `OPTIONS`
62
+ authenticated like other requests.
63
+ - **Credential hygiene enforced at startup:** scan keys, admin keys, policy
64
+ signing key, and policy-state key must all be pairwise distinct
65
+ (constant-time overlap check); policy-state key has a minimum length.
66
+ - **Browser hardening:** CSP, `nosniff`, `DENY` framing, `no-store`, COOP,
67
+ conditional HSTS; docs/redoc/openapi endpoints disabled.
68
+
69
+ ### 2.3 Policy push (`core/policy.py`) — the standout design
70
+ Signed (HMAC-SHA256, pluggable verifier) config bundles with a **structural
71
+ protection floor**: allow-listed fields only, always-on detectors cannot be
72
+ disabled or de-weighted below baseline, block-threshold ceiling, aggregate
73
+ degradation cap, fail-safe (never fail-open) semantics with durable HMAC'd
74
+ anti-replay state. A compromised control plane cannot remotely weaken the
75
+ fleet — this is the correct threat model, implemented end-to-end.
76
+
77
+ ### 2.4 Supply chain & CI
78
+ - All GitHub Actions **SHA-pinned**; `persist-credentials: false`; read-only
79
+ token permissions; actionlint gate.
80
+ - **Reproducible-build gate** (double build + `cmp`), Twine metadata check,
81
+ installed-wheel smoke test.
82
+ - Hash-locked build/container lockfiles (`uv pip compile --generate-hashes`),
83
+ `--require-hashes` installs in the Dockerfile; base image **digest-pinned**.
84
+ - Per-extra dependency-audit matrix (dynamically discovered), Trivy SBOM +
85
+ fail-on-fixable-CRITICAL/HIGH gate, PyPI OIDC publishing, SLSA/CycloneDX
86
+ attestations, Dependabot, pre-commit.
87
+ - Container: non-root user, read-only root fs, `cap_drop: ALL`,
88
+ `no-new-privileges`, pids/mem/cpu limits, healthcheck; CI proves the image
89
+ **refuses to start without keys** and exercises 401/200/413 paths end-to-end.
90
+
91
+ ### 2.5 Honest posture
92
+ `docs/PRODUCTION_READINESS.md` reports blind-generalization detection honestly
93
+ (v3: 30% ASR / 30% FPR — labeled **Beta**, not oversold), separates
94
+ library-ready vs operator-owned items, and the changelog is precise about what
95
+ each release changed.
96
+
97
+ ---
98
+
99
+ ## 3. Findings
100
+
101
+ ### M-1 (Medium) — `apply_bundle` accepts unsigned bundles when no verifier is passed
102
+ `src/shadowshield/core/policy.py:250` — `if verifier is not None and not
103
+ verifier(bundle): raise PolicyRejected`. When the caller omits `verifier`, an
104
+ **unsigned bundle is silently applied**, contradicting the module's own
105
+ documented invariant ("an unsigned/badly-signed bundle is rejected"). The
106
+ control plane compensates (`control.py:1286` rejects remote updates when no
107
+ verifier exists unless `allow_insecure_local`), so the deployed path is safe —
108
+ but the **library API itself is fail-open by default**, and any future caller
109
+ that forgets the verifier gets no protection.
110
+
111
+ **Recommendation:** make verification mandatory by default — e.g. raise unless
112
+ `verifier` is provided, or add an explicit `allow_unsigned: bool = False`
113
+ opt-in so skipping signature checks is a deliberate, greppable decision.
114
+
115
+ ### L-1 (Low) — Coverage gaps in glue/integration code
116
+ 86% overall, but: `middleware/langchain.py` **0%**, `plugins/manager.py` 33%,
117
+ `integrations/agentdojo.py` 23%, `middleware/base.py` 57%, `detectors/pii.py`
118
+ 77%. LangChain middleware and the plugin manager are user-facing extension
119
+ points; their error paths are exactly where misuse bugs live. The AgentDojo
120
+ adapter is acceptable to leave low (needs API keys), but langchain middleware
121
+ and plugin-manager state transitions deserve unit tests with fakes.
122
+
123
+ ### L-2 (Low) — `control.py` is a 1,448-line god-file
124
+ App factory, auth, policy endpoints, metrics/Prometheus, dashboard HTML, and
125
+ CLI glue in one module. It is well organized internally, but this size raises
126
+ audit cost and regression risk per edit. Consider splitting into
127
+ `control/auth.py`, `control/policy_api.py`, `control/metrics.py`,
128
+ `control/dashboard.py`.
129
+
130
+ ### L-3 (Low) — Reporter transport hardening
131
+ `reporter.py:_http_transport`: TLS verification relies on httpx defaults
132
+ (fine), but there is no scheme enforcement — an `http://` endpoint would ship
133
+ telemetry plus the `x-api-key` header in cleartext. Add an https-only check
134
+ (or loud warning) at construction. The bounded queue with silent drop-counting
135
+ is good; consider emitting a log line when drops begin so operators notice
136
+ collector outages.
137
+
138
+ ### L-4 (Low) — Sanitizer overlapping spans
139
+ `sanitizer.py` replaces spans right-to-left with index clamping — correct for
140
+ nested/disjoint spans, but two overlapping detector spans can yield nested
141
+ `[redacted:…]` placeholders. Harmless (no corruption), but a span-merge pass
142
+ would produce cleaner output.
143
+
144
+ ### L-5 (Info) — Local pip-audit of `container.lock` fails offline
145
+ `pip-audit -r requirements/container.lock` errors locally because the lock
146
+ contains a direct-URL `colorama` entry combined with `--require-hashes`. CI
147
+ audits it with `--no-deps`, which passes — so this is a reproduction quirk,
148
+ not a defect. Worth one line in `requirements/` docs for anyone auditing
149
+ locally.
150
+
151
+ ### L-6 (Info) — Local environment drift
152
+ `uvicorn` is not installed in the local `.venv` despite being in the
153
+ `dashboard` extra (tests pass regardless). Also `.coverage`/`coverage.xml`/
154
+ `dist/`/`production-dist/` artifacts exist locally (all correctly gitignored).
155
+ Housekeeping only.
156
+
157
+ ---
158
+
159
+ ## 4. Non-findings (checked, no issue)
160
+
161
+ - **Committed secrets:** none. Two regex hits are deliberate detector test
162
+ fixtures (`tests/test_detectors.py:131`, `tests/test_telemetry.py:19`).
163
+ - **`.env.example`:** empty values by design; `compose.yaml` uses
164
+ `${VAR:?…}` so copying the example can never boot with placeholder creds.
165
+ - **Dangerous dynamic execution:** no `eval/exec/pickle/subprocess/shell=True`
166
+ in `src/`; YAML usage is safe; FastAPI apps disable introspection endpoints.
167
+ - **Skipped tests:** both skips are opt-in real-model tests covered by the
168
+ dedicated `ml-integration` CI job, not silent gaps.
169
+ - **Site headers:** `site/vercel.json` ships a strict CSP with per-script
170
+ hashes, HSTS preload, and a restrictive Permissions-Policy.
171
+
172
+ ## 5. Recommended next actions (priority order)
173
+
174
+ 1. **M-1:** make bundle signature verification non-skippable by default in
175
+ `core/policy.py` (small change, closes the only fail-open path).
176
+ 2. **L-1:** add unit tests for `middleware/langchain.py` and
177
+ `plugins/manager.py` error paths.
178
+ 3. **L-2:** split `control.py` before it grows further.
179
+ 4. **L-3:** enforce https (or warn) on `Reporter` endpoints.
180
+ 5. Keep the blind-benchmark program running; the honest Beta label is a
181
+ strength — protect it.
182
+
183
+ ---
184
+
185
+ *Verification commands: `.venv/Scripts/python.exe -m pytest tests -q`,
186
+ `-m ruff check .`, `-m mypy`, `-m coverage report`,
187
+ `uv tool run pip-audit --path .venv/Lib/site-packages`,
188
+ plus targeted source review of `_security.py`, `core/policy.py`,
189
+ `core/engine.py`, `responders/sanitizer.py`, `reporter.py`, `control.py`,
190
+ CI workflows, Dockerfile, and `compose.yaml`.*
@@ -26,14 +26,40 @@ pip install "shadowshield[transformers]"
26
26
  shadowshield benchmark --hf deepset/prompt-injections --split test --transformer
27
27
  ```
28
28
 
29
+ The CLI warms every configured detector before starting the latency clock. Its
30
+ JSON and text reports include a runtime-integrity section covering warmup,
31
+ readiness, and bounded per-detector failure counts. A benchmark with a warmup
32
+ failure, a detector that remains unready, or any detector error exits non-zero
33
+ and is marked **unreliable**; its confusion counts remain available for
34
+ diagnosis but must not be published as quality evidence. CLI benchmark shields
35
+ disable normal audit emission so logging I/O does not distort latency or corrupt
36
+ the command's machine-readable JSON stdout.
37
+
38
+ Programmatic callers can request the same behavior with
39
+ `evaluate_shield(shield, examples, warmup=True)`. The default remains
40
+ `warmup=False` for API compatibility. In both modes, readiness is checked after
41
+ the scans and per-scan detector error counters are aggregated.
42
+
43
+ Per-category output retains the original `total`, `flagged`, and flag-rate
44
+ fields, and additionally reports class-conditional TP/FP/TN/FN, recall, FPR,
45
+ and balanced accuracy. Recall is `null`/`n/a` for categories with no attack
46
+ rows; FPR is `null`/`n/a` for categories with no benign rows; balanced accuracy
47
+ is defined only when both classes are present.
48
+
49
+ Aggregate and per-category recall/FPR also include two-sided 95% Wilson score
50
+ intervals. Wilson intervals remain informative at small sample sizes and at
51
+ observed rates of 0% or 100%; an interval is `null`/`n/a` when that category has
52
+ no rows for the corresponding class. Existing scalar metric fields are
53
+ unchanged, and the interval fields are additive.
54
+
29
55
  ## 1. Bundled benchmark (in-distribution — a regression baseline)
30
56
 
31
57
  75 curated examples (40 attack / 35 benign, incl. 16 NotInject-style hard
32
58
  negatives). `balanced` mode:
33
59
 
34
- | recall | FPR | precision | F1 | p50 |
60
+ | recall (95% CI) | FPR (95% CI) | precision | F1 | p50 |
35
61
  |---:|---:|---:|---:|---:|
36
- | 100% | 0% | 100% | 100% | 0.16 ms |
62
+ | 100% [91.2%, 100%] | 0% [0%, 9.9%] | 100% | 100% | 0.16 ms |
37
63
 
38
64
  **This is a regression baseline and a smoke test — NOT a claim of real-world
39
65
  accuracy.** 100% on our own set just means we don't regress on the attack
@@ -45,9 +71,9 @@ The 36-row adversarial catalogue includes obfuscation, multilingual and indirect
45
71
  attacks, plus benign trigger-heavy counterexamples. It improved from
46
72
  15/2/16/3 to 18/0/18/0 (TP/FP/TN/FN):
47
73
 
48
- | recall | FPR | precision | F1 |
74
+ | recall (95% CI) | FPR (95% CI) | precision | F1 |
49
75
  |---:|---:|---:|---:|
50
- | 100% | 0% | 100% | 100% |
76
+ | 100% [82.4%, 100%] | 0% [0%, 17.6%] | 100% | 100% |
51
77
 
52
78
  This is still a curated regression set. The signatures were developed with these
53
79
  cases visible, so its perfect score is not a generalization claim.
@@ -67,12 +93,12 @@ v1 SHA-256 `b3281ba1a42d266bb930bbb41943016d47b38dbc822ff7cff5131f3448a0248f`;
67
93
  v2 SHA-256 `aa8b8c81c00a55bb65180e15ff743b6241d24845b3886e8e60b52b9b23db47fa`;
68
94
  v3 SHA-256 `2285031e8143572311a522a4b6ec1a39c96a34ac2b42785f17818e8a145342bf`.
69
95
 
70
- | snapshot | rows | TP/FP/TN/FN | recall | FPR | balanced accuracy |
96
+ | snapshot | rows | TP/FP/TN/FN | recall (95% CI) | FPR (95% CI) | balanced accuracy |
71
97
  |---|---:|---:|---:|---:|---:|
72
- | v1 | 30 | 4/2/13/11 | 26.7% | 13.3% | 56.7% |
73
- | v2 | 20 | 0/1/9/10 | 0% | 10% | 45% |
74
- | v3 | 40 | 6/6/14/14 | 30% | 30% | 50% |
75
- | **aggregate** | **90** | **10/9/36/35** | **22.2%** | **20%** | **51.1%** |
98
+ | v1 | 30 | 4/2/13/11 | 26.7% [10.9%, 52.0%] | 13.3% [3.7%, 37.9%] | 56.7% |
99
+ | v2 | 20 | 0/1/9/10 | 0% [0%, 27.8%] | 10% [1.8%, 40.4%] | 45% |
100
+ | v3 | 40 | 6/6/14/14 | 30% [14.5%, 51.9%] | 30% [14.5%, 51.9%] | 50% |
101
+ | **aggregate** | **90** | **10/9/36/35** | **22.2% [12.5%, 36.3%]** | **20% [10.9%, 33.8%]** | **51.1%** |
76
102
 
77
103
  The v3 snapshot was opened only after a detector candidate and acceptance bar
78
104
  were frozen. That candidate raised aggregate recall to 55.6% but failed the