langchain-sync-monitors 0.1.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (281) hide show
  1. langchain_sync_monitors-0.1.0/.gitignore +37 -0
  2. langchain_sync_monitors-0.1.0/CHANGELOG.md +106 -0
  3. langchain_sync_monitors-0.1.0/CITATION.cff +32 -0
  4. langchain_sync_monitors-0.1.0/LICENSE +21 -0
  5. langchain_sync_monitors-0.1.0/PKG-INFO +223 -0
  6. langchain_sync_monitors-0.1.0/README.md +183 -0
  7. langchain_sync_monitors-0.1.0/docs/assets/brand/favicon-180.png +0 -0
  8. langchain_sync_monitors-0.1.0/docs/assets/brand/favicon-32.png +0 -0
  9. langchain_sync_monitors-0.1.0/docs/assets/brand/favicon-48.png +0 -0
  10. langchain_sync_monitors-0.1.0/docs/assets/brand/omamori-lab.svg +532 -0
  11. langchain_sync_monitors-0.1.0/docs/assets/diagrams/halt-stands-dark.svg +66 -0
  12. langchain_sync_monitors-0.1.0/docs/assets/diagrams/halt-stands-light.svg +66 -0
  13. langchain_sync_monitors-0.1.0/docs/assets/diagrams/live-sandbox-dark.svg +89 -0
  14. langchain_sync_monitors-0.1.0/docs/assets/diagrams/live-sandbox-light.svg +89 -0
  15. langchain_sync_monitors-0.1.0/docs/assets/diagrams/live-traces-dark.svg +211 -0
  16. langchain_sync_monitors-0.1.0/docs/assets/diagrams/live-traces-light.svg +211 -0
  17. langchain_sync_monitors-0.1.0/docs/assets/diagrams/live-visibility-dark.svg +86 -0
  18. langchain_sync_monitors-0.1.0/docs/assets/diagrams/live-visibility-light.svg +86 -0
  19. langchain_sync_monitors-0.1.0/docs/assets/diagrams/monitor-view-dark.svg +75 -0
  20. langchain_sync_monitors-0.1.0/docs/assets/diagrams/monitor-view-light.svg +75 -0
  21. langchain_sync_monitors-0.1.0/docs/assets/diagrams/monitored-step-dark.svg +98 -0
  22. langchain_sync_monitors-0.1.0/docs/assets/diagrams/monitored-step-light.svg +98 -0
  23. langchain_sync_monitors-0.1.0/docs/assets/diagrams/monitors-average-then-calibrate-dark.svg +70 -0
  24. langchain_sync_monitors-0.1.0/docs/assets/diagrams/monitors-average-then-calibrate-light.svg +70 -0
  25. langchain_sync_monitors-0.1.0/docs/assets/diagrams/monitors-cascade-dark.svg +62 -0
  26. langchain_sync_monitors-0.1.0/docs/assets/diagrams/monitors-cascade-light.svg +62 -0
  27. langchain_sync_monitors-0.1.0/docs/assets/diagrams/monitors-chat-judge-dark.svg +88 -0
  28. langchain_sync_monitors-0.1.0/docs/assets/diagrams/monitors-chat-judge-light.svg +88 -0
  29. langchain_sync_monitors-0.1.0/docs/assets/diagrams/monitors-dark.svg +83 -0
  30. langchain_sync_monitors-0.1.0/docs/assets/diagrams/monitors-decision-model-dark.svg +78 -0
  31. langchain_sync_monitors-0.1.0/docs/assets/diagrams/monitors-decision-model-light.svg +78 -0
  32. langchain_sync_monitors-0.1.0/docs/assets/diagrams/monitors-guard-scoring-dark.svg +100 -0
  33. langchain_sync_monitors-0.1.0/docs/assets/diagrams/monitors-guard-scoring-light.svg +100 -0
  34. langchain_sync_monitors-0.1.0/docs/assets/diagrams/monitors-honest-scores-dark.svg +63 -0
  35. langchain_sync_monitors-0.1.0/docs/assets/diagrams/monitors-honest-scores-light.svg +63 -0
  36. langchain_sync_monitors-0.1.0/docs/assets/diagrams/monitors-light.svg +83 -0
  37. langchain_sync_monitors-0.1.0/docs/assets/diagrams/protocol-auto-mode-dark.svg +73 -0
  38. langchain_sync_monitors-0.1.0/docs/assets/diagrams/protocol-auto-mode-light.svg +73 -0
  39. langchain_sync_monitors-0.1.0/docs/assets/diagrams/protocol-choice-dark.svg +76 -0
  40. langchain_sync_monitors-0.1.0/docs/assets/diagrams/protocol-choice-light.svg +76 -0
  41. langchain_sync_monitors-0.1.0/docs/assets/diagrams/protocol-defer-to-resample-dark.svg +72 -0
  42. langchain_sync_monitors-0.1.0/docs/assets/diagrams/protocol-defer-to-resample-light.svg +72 -0
  43. langchain_sync_monitors-0.1.0/docs/assets/diagrams/protocol-defer-to-trusted-dark.svg +58 -0
  44. langchain_sync_monitors-0.1.0/docs/assets/diagrams/protocol-defer-to-trusted-light.svg +58 -0
  45. langchain_sync_monitors-0.1.0/docs/assets/diagrams/protocol-feedback-visibility-dark.svg +76 -0
  46. langchain_sync_monitors-0.1.0/docs/assets/diagrams/protocol-feedback-visibility-light.svg +76 -0
  47. langchain_sync_monitors-0.1.0/docs/assets/diagrams/protocol-first-agent-dark.svg +64 -0
  48. langchain_sync_monitors-0.1.0/docs/assets/diagrams/protocol-first-agent-light.svg +64 -0
  49. langchain_sync_monitors-0.1.0/docs/assets/diagrams/protocol-trusted-monitoring-dark.svg +58 -0
  50. langchain_sync_monitors-0.1.0/docs/assets/diagrams/protocol-trusted-monitoring-light.svg +58 -0
  51. langchain_sync_monitors-0.1.0/docs/assets/diagrams/protocols-dark.svg +96 -0
  52. langchain_sync_monitors-0.1.0/docs/assets/diagrams/protocols-light.svg +96 -0
  53. langchain_sync_monitors-0.1.0/docs/assets/diagrams/step-commit-dark.svg +56 -0
  54. langchain_sync_monitors-0.1.0/docs/assets/diagrams/step-commit-light.svg +56 -0
  55. langchain_sync_monitors-0.1.0/docs/assets/diagrams/subagent-halts-dark.svg +55 -0
  56. langchain_sync_monitors-0.1.0/docs/assets/diagrams/subagent-halts-light.svg +55 -0
  57. langchain_sync_monitors-0.1.0/docs/assets/diagrams/subagent-monitors-dark.svg +74 -0
  58. langchain_sync_monitors-0.1.0/docs/assets/diagrams/subagent-monitors-light.svg +74 -0
  59. langchain_sync_monitors-0.1.0/docs/assets/diagrams/subagents-dark.svg +70 -0
  60. langchain_sync_monitors-0.1.0/docs/assets/diagrams/subagents-light.svg +70 -0
  61. langchain_sync_monitors-0.1.0/docs/assets/diagrams/who-speaks-as-the-user-dark.svg +86 -0
  62. langchain_sync_monitors-0.1.0/docs/assets/diagrams/who-speaks-as-the-user-light.svg +86 -0
  63. langchain_sync_monitors-0.1.0/docs/explanation/background.md +235 -0
  64. langchain_sync_monitors-0.1.0/docs/explanation/design.md +1307 -0
  65. langchain_sync_monitors-0.1.0/docs/explanation/index.md +14 -0
  66. langchain_sync_monitors-0.1.0/docs/explanation/live-runs.md +454 -0
  67. langchain_sync_monitors-0.1.0/docs/hooks/citations.py +159 -0
  68. langchain_sync_monitors-0.1.0/docs/how-to/choose-a-protocol.md +258 -0
  69. langchain_sync_monitors-0.1.0/docs/how-to/choose-what-the-monitor-reads.md +411 -0
  70. langchain_sync_monitors-0.1.0/docs/how-to/combine-and-calibrate-monitors.md +393 -0
  71. langchain_sync_monitors-0.1.0/docs/how-to/index.md +37 -0
  72. langchain_sync_monitors-0.1.0/docs/how-to/monitor-deep-agents-subagents.md +424 -0
  73. langchain_sync_monitors-0.1.0/docs/how-to/read-the-monitor-log.md +456 -0
  74. langchain_sync_monitors-0.1.0/docs/how-to/see-decisions-in-langsmith-and-langfuse.md +426 -0
  75. langchain_sync_monitors-0.1.0/docs/how-to/use-a-chat-judge.md +267 -0
  76. langchain_sync_monitors-0.1.0/docs/how-to/use-a-decision-model.md +268 -0
  77. langchain_sync_monitors-0.1.0/docs/how-to/use-a-guard-model.md +350 -0
  78. langchain_sync_monitors-0.1.0/docs/how-to/use-auto-mode.md +353 -0
  79. langchain_sync_monitors-0.1.0/docs/how-to/use-defer-to-resample.md +343 -0
  80. langchain_sync_monitors-0.1.0/docs/how-to/use-defer-to-trusted.md +162 -0
  81. langchain_sync_monitors-0.1.0/docs/how-to/use-trusted-monitoring.md +178 -0
  82. langchain_sync_monitors-0.1.0/docs/index.md +240 -0
  83. langchain_sync_monitors-0.1.0/docs/reference/api.md +386 -0
  84. langchain_sync_monitors-0.1.0/docs/reference/index.md +8 -0
  85. langchain_sync_monitors-0.1.0/docs/references.bib +866 -0
  86. langchain_sync_monitors-0.1.0/docs/stylesheets/brand.css +15 -0
  87. langchain_sync_monitors-0.1.0/docs/stylesheets/site.css +251 -0
  88. langchain_sync_monitors-0.1.0/docs/tutorials/first-monitored-agent.md +427 -0
  89. langchain_sync_monitors-0.1.0/docs/tutorials/index.md +9 -0
  90. langchain_sync_monitors-0.1.0/pyproject.toml +221 -0
  91. langchain_sync_monitors-0.1.0/src/langchain_sync_monitors/__init__.py +154 -0
  92. langchain_sync_monitors-0.1.0/src/langchain_sync_monitors/_langchain.py +610 -0
  93. langchain_sync_monitors-0.1.0/src/langchain_sync_monitors/commits.py +175 -0
  94. langchain_sync_monitors-0.1.0/src/langchain_sync_monitors/concurrency.py +39 -0
  95. langchain_sync_monitors-0.1.0/src/langchain_sync_monitors/context_values.py +22 -0
  96. langchain_sync_monitors-0.1.0/src/langchain_sync_monitors/contracts.py +385 -0
  97. langchain_sync_monitors-0.1.0/src/langchain_sync_monitors/deepagents.py +344 -0
  98. langchain_sync_monitors-0.1.0/src/langchain_sync_monitors/delegation.py +77 -0
  99. langchain_sync_monitors-0.1.0/src/langchain_sync_monitors/errors.py +65 -0
  100. langchain_sync_monitors-0.1.0/src/langchain_sync_monitors/feedback.py +119 -0
  101. langchain_sync_monitors-0.1.0/src/langchain_sync_monitors/halts.py +212 -0
  102. langchain_sync_monitors-0.1.0/src/langchain_sync_monitors/message_ids.py +100 -0
  103. langchain_sync_monitors-0.1.0/src/langchain_sync_monitors/middleware.py +463 -0
  104. langchain_sync_monitors-0.1.0/src/langchain_sync_monitors/model_calls.py +142 -0
  105. langchain_sync_monitors-0.1.0/src/langchain_sync_monitors/monitor_state.py +91 -0
  106. langchain_sync_monitors-0.1.0/src/langchain_sync_monitors/monitors/__init__.py +46 -0
  107. langchain_sync_monitors-0.1.0/src/langchain_sync_monitors/monitors/chat.py +529 -0
  108. langchain_sync_monitors-0.1.0/src/langchain_sync_monitors/monitors/composition.py +264 -0
  109. langchain_sync_monitors-0.1.0/src/langchain_sync_monitors/monitors/decision.py +332 -0
  110. langchain_sync_monitors-0.1.0/src/langchain_sync_monitors/monitors/decision_questions.py +125 -0
  111. langchain_sync_monitors-0.1.0/src/langchain_sync_monitors/monitors/guard.py +576 -0
  112. langchain_sync_monitors-0.1.0/src/langchain_sync_monitors/monitors/guard_labels.py +310 -0
  113. langchain_sync_monitors-0.1.0/src/langchain_sync_monitors/monitors/openrouter_decisions.py +339 -0
  114. langchain_sync_monitors-0.1.0/src/langchain_sync_monitors/options.py +272 -0
  115. langchain_sync_monitors-0.1.0/src/langchain_sync_monitors/pending_steps.py +480 -0
  116. langchain_sync_monitors-0.1.0/src/langchain_sync_monitors/placement.py +484 -0
  117. langchain_sync_monitors-0.1.0/src/langchain_sync_monitors/prompts.py +135 -0
  118. langchain_sync_monitors-0.1.0/src/langchain_sync_monitors/protocols/__init__.py +34 -0
  119. langchain_sync_monitors-0.1.0/src/langchain_sync_monitors/protocols/auto_mode.py +251 -0
  120. langchain_sync_monitors-0.1.0/src/langchain_sync_monitors/protocols/defer_to_resample.py +223 -0
  121. langchain_sync_monitors-0.1.0/src/langchain_sync_monitors/protocols/fallbacks.py +132 -0
  122. langchain_sync_monitors-0.1.0/src/langchain_sync_monitors/protocols/trusted_monitoring.py +49 -0
  123. langchain_sync_monitors-0.1.0/src/langchain_sync_monitors/py.typed +0 -0
  124. langchain_sync_monitors-0.1.0/src/langchain_sync_monitors/records.py +272 -0
  125. langchain_sync_monitors-0.1.0/src/langchain_sync_monitors/returned_records.py +319 -0
  126. langchain_sync_monitors-0.1.0/src/langchain_sync_monitors/run_inputs.py +365 -0
  127. langchain_sync_monitors-0.1.0/src/langchain_sync_monitors/server_tools.py +171 -0
  128. langchain_sync_monitors-0.1.0/src/langchain_sync_monitors/spans.py +251 -0
  129. langchain_sync_monitors-0.1.0/src/langchain_sync_monitors/state_keys.py +49 -0
  130. langchain_sync_monitors-0.1.0/src/langchain_sync_monitors/subagent_returns.py +115 -0
  131. langchain_sync_monitors-0.1.0/src/langchain_sync_monitors/task_authorship.py +557 -0
  132. langchain_sync_monitors-0.1.0/src/langchain_sync_monitors/thresholds.py +219 -0
  133. langchain_sync_monitors-0.1.0/src/langchain_sync_monitors/tool_calls.py +137 -0
  134. langchain_sync_monitors-0.1.0/src/langchain_sync_monitors/transcript.py +556 -0
  135. langchain_sync_monitors-0.1.0/tests/__init__.py +0 -0
  136. langchain_sync_monitors-0.1.0/tests/integration/__init__.py +0 -0
  137. langchain_sync_monitors-0.1.0/tests/integration/conftest.py +3 -0
  138. langchain_sync_monitors-0.1.0/tests/integration/test_audit_queries.py +194 -0
  139. langchain_sync_monitors-0.1.0/tests/integration/test_classifier_spans.py +171 -0
  140. langchain_sync_monitors-0.1.0/tests/integration/test_feedback_reasons.py +63 -0
  141. langchain_sync_monitors-0.1.0/tests/integration/test_full_stack.py +138 -0
  142. langchain_sync_monitors-0.1.0/tests/integration/test_kept_user_turns.py +936 -0
  143. langchain_sync_monitors-0.1.0/tests/integration/test_message_stream.py +294 -0
  144. langchain_sync_monitors-0.1.0/tests/integration/test_monitor_call_names.py +267 -0
  145. langchain_sync_monitors-0.1.0/tests/integration/test_monitor_spans.py +458 -0
  146. langchain_sync_monitors-0.1.0/tests/integration/test_protocols_in_agents.py +263 -0
  147. langchain_sync_monitors-0.1.0/tests/integration/test_spans_beside_langsmith.py +265 -0
  148. langchain_sync_monitors-0.1.0/tests/integration/test_streams_with_spans.py +205 -0
  149. langchain_sync_monitors-0.1.0/tests/integration/test_task_authorship.py +1371 -0
  150. langchain_sync_monitors-0.1.0/tests/integration/test_what_the_monitor_reads.py +462 -0
  151. langchain_sync_monitors-0.1.0/tests/live/__init__.py +41 -0
  152. langchain_sync_monitors-0.1.0/tests/live/checks.py +306 -0
  153. langchain_sync_monitors-0.1.0/tests/live/conftest.py +46 -0
  154. langchain_sync_monitors-0.1.0/tests/live/costs.py +270 -0
  155. langchain_sync_monitors-0.1.0/tests/live/deep_agents.py +267 -0
  156. langchain_sync_monitors-0.1.0/tests/live/fakes.py +240 -0
  157. langchain_sync_monitors-0.1.0/tests/live/harness.py +660 -0
  158. langchain_sync_monitors-0.1.0/tests/live/honest_scores.py +31 -0
  159. langchain_sync_monitors-0.1.0/tests/live/invariants.py +153 -0
  160. langchain_sync_monitors-0.1.0/tests/live/reports.py +326 -0
  161. langchain_sync_monitors-0.1.0/tests/live/sandbox.py +163 -0
  162. langchain_sync_monitors-0.1.0/tests/live/scenario.py +193 -0
  163. langchain_sync_monitors-0.1.0/tests/live/test_deep_agents_subagents.py +232 -0
  164. langchain_sync_monitors-0.1.0/tests/live/test_eval_checks_offline.py +988 -0
  165. langchain_sync_monitors-0.1.0/tests/live/test_halt_path.py +75 -0
  166. langchain_sync_monitors-0.1.0/tests/live/test_harness_offline.py +411 -0
  167. langchain_sync_monitors-0.1.0/tests/live/test_honest_runs.py +185 -0
  168. langchain_sync_monitors-0.1.0/tests/live/test_kept_run_inputs.py +254 -0
  169. langchain_sync_monitors-0.1.0/tests/live/test_monitor_families.py +130 -0
  170. langchain_sync_monitors-0.1.0/tests/live/test_protocol_paths.py +89 -0
  171. langchain_sync_monitors-0.1.0/tests/live/test_whole_agent_runs.py +94 -0
  172. langchain_sync_monitors-0.1.0/tests/live/test_wrapper_monitors.py +88 -0
  173. langchain_sync_monitors-0.1.0/tests/live/traces.py +234 -0
  174. langchain_sync_monitors-0.1.0/tests/live/wrapper_checks.py +80 -0
  175. langchain_sync_monitors-0.1.0/tests/support/__init__.py +0 -0
  176. langchain_sync_monitors-0.1.0/tests/support/agents.py +256 -0
  177. langchain_sync_monitors-0.1.0/tests/support/array_scalars.py +65 -0
  178. langchain_sync_monitors-0.1.0/tests/support/chat_models.py +206 -0
  179. langchain_sync_monitors-0.1.0/tests/support/deep_agents.py +78 -0
  180. langchain_sync_monitors-0.1.0/tests/support/fixtures.py +15 -0
  181. langchain_sync_monitors-0.1.0/tests/support/flaky_models.py +102 -0
  182. langchain_sync_monitors-0.1.0/tests/support/monitors.py +151 -0
  183. langchain_sync_monitors-0.1.0/tests/support/protocols.py +154 -0
  184. langchain_sync_monitors-0.1.0/tests/support/server_tools.py +126 -0
  185. langchain_sync_monitors-0.1.0/tests/support/tracing.py +320 -0
  186. langchain_sync_monitors-0.1.0/tests/support/written_human_messages.py +339 -0
  187. langchain_sync_monitors-0.1.0/tests/unit/__init__.py +0 -0
  188. langchain_sync_monitors-0.1.0/tests/unit/deepagents/__init__.py +0 -0
  189. langchain_sync_monitors-0.1.0/tests/unit/deepagents/conftest.py +7 -0
  190. langchain_sync_monitors-0.1.0/tests/unit/deepagents/test_deep_agent_protocols.py +150 -0
  191. langchain_sync_monitors-0.1.0/tests/unit/deepagents/test_deep_agent_runs.py +136 -0
  192. langchain_sync_monitors-0.1.0/tests/unit/deepagents/test_delegation_state.py +250 -0
  193. langchain_sync_monitors-0.1.0/tests/unit/deepagents/test_evicted_messages.py +78 -0
  194. langchain_sync_monitors-0.1.0/tests/unit/deepagents/test_failed_delegations.py +311 -0
  195. langchain_sync_monitors-0.1.0/tests/unit/deepagents/test_general_purpose_subagent.py +84 -0
  196. langchain_sync_monitors-0.1.0/tests/unit/deepagents/test_halts_under_harness_nudges.py +93 -0
  197. langchain_sync_monitors-0.1.0/tests/unit/deepagents/test_harness_nudges.py +81 -0
  198. langchain_sync_monitors-0.1.0/tests/unit/deepagents/test_kept_user_turns.py +339 -0
  199. langchain_sync_monitors-0.1.0/tests/unit/deepagents/test_monitor_subagents.py +219 -0
  200. langchain_sync_monitors-0.1.0/tests/unit/deepagents/test_nested_delegations.py +119 -0
  201. langchain_sync_monitors-0.1.0/tests/unit/deepagents/test_parent_bound_records.py +179 -0
  202. langchain_sync_monitors-0.1.0/tests/unit/deepagents/test_returned_subagent_records.py +578 -0
  203. langchain_sync_monitors-0.1.0/tests/unit/deepagents/test_rubric_after_halt.py +86 -0
  204. langchain_sync_monitors-0.1.0/tests/unit/deepagents/test_same_named_subagents.py +270 -0
  205. langchain_sync_monitors-0.1.0/tests/unit/deepagents/test_stacked_subagent_monitors.py +279 -0
  206. langchain_sync_monitors-0.1.0/tests/unit/deepagents/test_subagent_guards.py +127 -0
  207. langchain_sync_monitors-0.1.0/tests/unit/deepagents/test_subagent_halts.py +51 -0
  208. langchain_sync_monitors-0.1.0/tests/unit/deepagents/test_subagent_spans.py +113 -0
  209. langchain_sync_monitors-0.1.0/tests/unit/deepagents/test_subagent_task_authorship.py +56 -0
  210. langchain_sync_monitors-0.1.0/tests/unit/deepagents/test_thread_block_budget.py +377 -0
  211. langchain_sync_monitors-0.1.0/tests/unit/middleware/__init__.py +0 -0
  212. langchain_sync_monitors-0.1.0/tests/unit/middleware/conftest.py +3 -0
  213. langchain_sync_monitors-0.1.0/tests/unit/middleware/test_array_scores.py +79 -0
  214. langchain_sync_monitors-0.1.0/tests/unit/middleware/test_bound_server_tools.py +154 -0
  215. langchain_sync_monitors-0.1.0/tests/unit/middleware/test_cached_resamples.py +296 -0
  216. langchain_sync_monitors-0.1.0/tests/unit/middleware/test_closed_steps.py +112 -0
  217. langchain_sync_monitors-0.1.0/tests/unit/middleware/test_delegation.py +428 -0
  218. langchain_sync_monitors-0.1.0/tests/unit/middleware/test_failed_step_spans.py +214 -0
  219. langchain_sync_monitors-0.1.0/tests/unit/middleware/test_failed_steps.py +543 -0
  220. langchain_sync_monitors-0.1.0/tests/unit/middleware/test_failing_stream_writer.py +134 -0
  221. langchain_sync_monitors-0.1.0/tests/unit/middleware/test_feedback.py +202 -0
  222. langchain_sync_monitors-0.1.0/tests/unit/middleware/test_forged_records.py +331 -0
  223. langchain_sync_monitors-0.1.0/tests/unit/middleware/test_halt_input_counts.py +479 -0
  224. langchain_sync_monitors-0.1.0/tests/unit/middleware/test_halted_runs.py +330 -0
  225. langchain_sync_monitors-0.1.0/tests/unit/middleware/test_middleware_feedback.py +131 -0
  226. langchain_sync_monitors-0.1.0/tests/unit/middleware/test_middleware_outcomes.py +204 -0
  227. langchain_sync_monitors-0.1.0/tests/unit/middleware/test_middleware_state.py +352 -0
  228. langchain_sync_monitors-0.1.0/tests/unit/middleware/test_middleware_streaming.py +106 -0
  229. langchain_sync_monitors-0.1.0/tests/unit/middleware/test_middleware_sync.py +259 -0
  230. langchain_sync_monitors-0.1.0/tests/unit/middleware/test_pending_steps.py +531 -0
  231. langchain_sync_monitors-0.1.0/tests/unit/middleware/test_placement.py +314 -0
  232. langchain_sync_monitors-0.1.0/tests/unit/middleware/test_records.py +471 -0
  233. langchain_sync_monitors-0.1.0/tests/unit/middleware/test_returned_records.py +530 -0
  234. langchain_sync_monitors-0.1.0/tests/unit/middleware/test_run_synchronously.py +208 -0
  235. langchain_sync_monitors-0.1.0/tests/unit/middleware/test_server_tool_marks.py +142 -0
  236. langchain_sync_monitors-0.1.0/tests/unit/middleware/test_server_tools.py +305 -0
  237. langchain_sync_monitors-0.1.0/tests/unit/middleware/test_stacked_monitors.py +587 -0
  238. langchain_sync_monitors-0.1.0/tests/unit/middleware/test_stacked_placement.py +537 -0
  239. langchain_sync_monitors-0.1.0/tests/unit/middleware/test_standing_halts.py +517 -0
  240. langchain_sync_monitors-0.1.0/tests/unit/middleware/test_subagent_returns.py +140 -0
  241. langchain_sync_monitors-0.1.0/tests/unit/middleware/test_tool_call_bubbles.py +101 -0
  242. langchain_sync_monitors-0.1.0/tests/unit/middleware/test_transcript_kept_out_of_logs.py +280 -0
  243. langchain_sync_monitors-0.1.0/tests/unit/monitors/__init__.py +1 -0
  244. langchain_sync_monitors-0.1.0/tests/unit/monitors/captured_replies.py +79 -0
  245. langchain_sync_monitors-0.1.0/tests/unit/monitors/conftest.py +67 -0
  246. langchain_sync_monitors-0.1.0/tests/unit/monitors/doubles.py +123 -0
  247. langchain_sync_monitors-0.1.0/tests/unit/monitors/test_cached_monitor_replies.py +203 -0
  248. langchain_sync_monitors-0.1.0/tests/unit/monitors/test_calibration.py +272 -0
  249. langchain_sync_monitors-0.1.0/tests/unit/monitors/test_chat.py +688 -0
  250. langchain_sync_monitors-0.1.0/tests/unit/monitors/test_composition.py +281 -0
  251. langchain_sync_monitors-0.1.0/tests/unit/monitors/test_decision.py +985 -0
  252. langchain_sync_monitors-0.1.0/tests/unit/monitors/test_guard.py +1647 -0
  253. langchain_sync_monitors-0.1.0/tests/unit/monitors/test_prompts.py +81 -0
  254. langchain_sync_monitors-0.1.0/tests/unit/monitors/test_rate_limits.py +262 -0
  255. langchain_sync_monitors-0.1.0/tests/unit/protocols/__init__.py +1 -0
  256. langchain_sync_monitors-0.1.0/tests/unit/protocols/conftest.py +29 -0
  257. langchain_sync_monitors-0.1.0/tests/unit/protocols/scripted_step.py +157 -0
  258. langchain_sync_monitors-0.1.0/tests/unit/protocols/test_auto_mode.py +259 -0
  259. langchain_sync_monitors-0.1.0/tests/unit/protocols/test_defer_to_resample.py +387 -0
  260. langchain_sync_monitors-0.1.0/tests/unit/protocols/test_fallbacks.py +136 -0
  261. langchain_sync_monitors-0.1.0/tests/unit/protocols/test_feedback_template.py +75 -0
  262. langchain_sync_monitors-0.1.0/tests/unit/protocols/test_synchronous_driving.py +42 -0
  263. langchain_sync_monitors-0.1.0/tests/unit/protocols/test_thresholds.py +727 -0
  264. langchain_sync_monitors-0.1.0/tests/unit/protocols/test_trusted_monitoring.py +73 -0
  265. langchain_sync_monitors-0.1.0/tests/unit/test_concurrency.py +54 -0
  266. langchain_sync_monitors-0.1.0/tests/unit/test_construction_warnings.py +73 -0
  267. langchain_sync_monitors-0.1.0/tests/unit/test_contracts.py +77 -0
  268. langchain_sync_monitors-0.1.0/tests/unit/test_docs_examples.py +569 -0
  269. langchain_sync_monitors-0.1.0/tests/unit/test_docs_standard.py +1203 -0
  270. langchain_sync_monitors-0.1.0/tests/unit/test_docstring_citations.py +271 -0
  271. langchain_sync_monitors-0.1.0/tests/unit/test_langchain_boundary.py +52 -0
  272. langchain_sync_monitors-0.1.0/tests/unit/test_langgraph_readers.py +161 -0
  273. langchain_sync_monitors-0.1.0/tests/unit/test_missing_extras.py +118 -0
  274. langchain_sync_monitors-0.1.0/tests/unit/test_model_calls.py +198 -0
  275. langchain_sync_monitors-0.1.0/tests/unit/test_option_types.py +1088 -0
  276. langchain_sync_monitors-0.1.0/tests/unit/test_options.py +91 -0
  277. langchain_sync_monitors-0.1.0/tests/unit/test_references.py +150 -0
  278. langchain_sync_monitors-0.1.0/tests/unit/test_run_inputs.py +588 -0
  279. langchain_sync_monitors-0.1.0/tests/unit/test_task_authorship.py +824 -0
  280. langchain_sync_monitors-0.1.0/tests/unit/test_traced_runs.py +340 -0
  281. langchain_sync_monitors-0.1.0/tests/unit/test_transcript.py +1179 -0
@@ -0,0 +1,37 @@
1
+ # Secrets
2
+ .env
3
+ .env.*
4
+ !.env.example
5
+
6
+ # Python
7
+ __pycache__/
8
+ *.py[cod]
9
+ *.egg-info/
10
+ .venv/
11
+ venv/
12
+ dist/
13
+ build/
14
+
15
+ # Tooling caches
16
+ .mypy_cache/
17
+ .ty_cache/
18
+ .ruff_cache/
19
+ .pytest_cache/
20
+ .coverage
21
+ htmlcov/
22
+
23
+ # Editors and OS
24
+ .idea/
25
+ .vscode/
26
+ .DS_Store
27
+
28
+ # Local agent notes, docs builds and live run reports
29
+ CLAUDE.local.md
30
+ site/
31
+ live-reports/
32
+
33
+ # Browser tool logs
34
+ .playwright-mcp/
35
+
36
+ # Agent worktrees (one checkout per parallel agent)
37
+ .claude/worktrees/
@@ -0,0 +1,106 @@
1
+ # Changelog
2
+
3
+ All notable changes to this project are documented in this file.
4
+
5
+ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
6
+ and this project adheres to
7
+ [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
8
+
9
+ The public API is the set of names exported from `langchain_sync_monitors`
10
+ (monitors, control protocols, the middleware and their constructor keywords)
11
+ and the shape of the `monitor_log` records. A release that could turn a working
12
+ integration into a failing one is a breaking change; before 1.0.0, a breaking
13
+ change raises the minor version.
14
+
15
+ ## [Unreleased]
16
+
17
+ ## [0.1.0] - 2026-10-01
18
+
19
+ ### Added
20
+
21
+ - `MonitorMiddleware`, which judges every step a LangChain agent or Deep Agent
22
+ proposes before any of its own tools run, under `invoke()` and `ainvoke()`.
23
+ - Control protocols `TrustedMonitoring`, `DeferToTrusted`, `DeferToResample`,
24
+ with sequential or parallel resampling (`Resampling`), and `AutoMode`, with
25
+ the `DeferToTrustedModel` and `HaltRun` fallbacks.
26
+ - `AutoMode` blocks a suspicious step and tells the agent why. After three
27
+ blocks in one step, or twenty in the thread, its fallback halts the run by
28
+ default.
29
+ - Default thresholds, 0.6 for `defer_threshold`, `block_threshold` and Trusted
30
+ Monitoring's `audit_threshold`, and 0.9 for the `audit_threshold` of Defer to
31
+ Trusted and Defer to Resample, which raise `UncalibratedThresholdWarning`
32
+ until you set your own.
33
+ - `ControlProtocol` and `PendingStep`, for a protocol of your own, with
34
+ `SampleDrawing` to draw a step's samples one after another or at once, and
35
+ `resolve_threshold` and `ThresholdStatistic` to warn about a default
36
+ threshold and name the score it is calibrated on.
37
+ - Monitors for three model families: `LLMMonitor` for chat judges;
38
+ `GuardModelMonitor` for guard models; and `DecisionModelMonitor` for
39
+ decision models, through `OpenRouterDecisionModel` or
40
+ `TypeSafeDecisionModel`.
41
+ - The `RepeatedMonitor`, `CalibratedMonitor` and `CascadeMonitor` wrappers.
42
+ - `MonitorView` and `Channel`, which choose what the monitor reads; the default
43
+ follows Claude Code's auto mode classifier.
44
+ - `monitor_log`, which keeps one `StepRecord` per committed step, with every
45
+ sample judged.
46
+ - `MonitorStepEvent` and `MonitorStepFailedEvent` on `stream_mode="custom"`.
47
+ Only committed steps reach `stream_mode="messages"`.
48
+ - A halted step ends the run, and the halt stands until a later run brings
49
+ new input the monitor can confirm.
50
+ - Each run's input stays available to the judge whole and in order, even
51
+ after summarisation. Every other human message is a note that authorises
52
+ nothing.
53
+ - Guarded tool writes: a message a tool writes to the state cannot speak as
54
+ the user or the monitor, and its writes to the monitor's keys are dropped,
55
+ except `monitor_log`, whose records are checked.
56
+ - `monitor_subagents`, which gives every Deep Agents subagent a monitor of its
57
+ own, and `SubagentHalt`, which decides what a subagent's halt does to the
58
+ run.
59
+ - Auto Mode's total counts the blocks across the conversation thread,
60
+ subagents included.
61
+ - `check_monitor_placement`, which warns with a `MonitorPlacementWarning` about
62
+ middleware or stacked monitors placed where they undermine the monitor.
63
+ - `ServerToolWarning`, `CachedResampleWarning` and `HardLabelWarning`, for a
64
+ known server tool the provider runs itself, a cache that copies resamples,
65
+ and a guard that scores with hard labels.
66
+ - Named spans in any LangChain tracer, LangSmith and Langfuse among them, and
67
+ the built-in monitors' own model calls named `monitor call`.
68
+ - The library's own log lines and errors never quote the transcript.
69
+ - A judge's or guard's reply that cannot be read fails closed, and a decision
70
+ model's invalid answer raises `MonitorError`.
71
+ - A chat judge's or guard's call that the provider answers with HTTP 429 is
72
+ made again, up to four attempts in all.
73
+ - Protocols, fallbacks, monitors and the middleware raise `ConfigurationError`
74
+ for an option of the wrong type.
75
+ - `MonitorError` and its subclasses `ConfigurationError`, `MissingExtraError`,
76
+ `SynchronousRunError` and `InvalidSuspicionError`.
77
+ - Packaging for Python 3.12, 3.13 and 3.14, with the `deepagents` (0.7.13 or
78
+ newer), `openrouter` and `typesafe` extras.
79
+ - A [documentation site](https://omamori-lab.github.io/langchain-sync-monitors/)
80
+ with a tutorial, how-to guides, the API reference and explanation pages.
81
+ - A bibliography of every paper, post and code base the library draws on,
82
+ cited where it is used.
83
+ - `CITATION.cff`, a security policy with private vulnerability reporting, and
84
+ issue forms.
85
+ - For contributors: one gate script, pre-commit hooks, and CI on every
86
+ supported Python.
87
+
88
+ ### Known limits
89
+
90
+ - Server tools, such as web search, run inside the model call, before the
91
+ monitor judges the step, and again for every sample.
92
+ - A subagent whose run raises returns no records, so its steps and blocks reach
93
+ the parent only if the run is resumed.
94
+ - Subagents that run in parallel do not see each other's blocks, so together
95
+ they can pass Auto Mode's total.
96
+ - `TypeSafeDecisionModel` reads a `false`, `true` or numeric string from the API
97
+ as a number, because `langchain-typesafe` parses answers leniently.
98
+ - A human message that a middleware listed before the monitor writes at a
99
+ run's start or end can count as the user's input and lift a halt.
100
+ - Forked subagents (`mode="fork"`) are not supported yet:
101
+ `monitor_subagents` refuses them with `ConfigurationError`.
102
+ - The full list is in
103
+ [Known limits and open paths](https://omamori-lab.github.io/langchain-sync-monitors/explanation/design/#known-limits-and-open-paths).
104
+
105
+ [Unreleased]: https://github.com/omamori-lab/langchain-sync-monitors/compare/v0.1.0...HEAD
106
+ [0.1.0]: https://github.com/omamori-lab/langchain-sync-monitors/releases/tag/v0.1.0
@@ -0,0 +1,32 @@
1
+ cff-version: 1.2.0
2
+ message: >-
3
+ If you use this software, please cite it using these metadata, and cite the
4
+ research it builds on: docs/references.bib lists every source.
5
+ type: software
6
+ title: langchain-sync-monitors
7
+ abstract: >-
8
+ AI-control monitors for LangChain agents and Deep Agents: every step an agent
9
+ proposes is judged before any of its own tools run, and a control protocol
10
+ decides what runs.
11
+ authors:
12
+ - given-names: Antonio
13
+ family-names: Badilla-Olivas
14
+ repository-code: https://github.com/omamori-lab/langchain-sync-monitors
15
+ url: https://omamori-lab.github.io/langchain-sync-monitors/
16
+ license: MIT
17
+ keywords:
18
+ - ai-safety
19
+ - ai-control
20
+ - langchain
21
+ - langgraph
22
+ - deepagents
23
+ - ai-agents
24
+ - llm-agents
25
+ - agent-safety
26
+ - monitoring
27
+ - guardrails
28
+ - prompt-injection
29
+ - llm-security
30
+ - middleware
31
+ version: "0.1.0"
32
+ date-released: "2026-10-01"
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 langchain-sync-monitors contributors
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,223 @@
1
+ Metadata-Version: 2.5
2
+ Name: langchain-sync-monitors
3
+ Version: 0.1.0
4
+ Summary: AI-control monitors for LangChain agents and Deep Agents: every step an agent proposes is judged before any of its own tools run, and a control protocol decides what runs.
5
+ Project-URL: Homepage, https://omamori-lab.github.io/langchain-sync-monitors/
6
+ Project-URL: Documentation, https://omamori-lab.github.io/langchain-sync-monitors/
7
+ Project-URL: Repository, https://github.com/omamori-lab/langchain-sync-monitors
8
+ Project-URL: Issues, https://github.com/omamori-lab/langchain-sync-monitors/issues
9
+ Project-URL: Changelog, https://github.com/omamori-lab/langchain-sync-monitors/blob/main/CHANGELOG.md
10
+ Author-email: Antonio Badilla-Olivas <omamori-lab@proton.me>
11
+ License-Expression: MIT
12
+ License-File: LICENSE
13
+ Keywords: agent-safety,ai-agents,ai-control,ai-safety,deepagents,guardrails,langchain,langgraph,llm-agents,llm-security,middleware,monitoring,prompt-injection
14
+ Classifier: Development Status :: 3 - Alpha
15
+ Classifier: Intended Audience :: Developers
16
+ Classifier: Intended Audience :: Science/Research
17
+ Classifier: Operating System :: OS Independent
18
+ Classifier: Programming Language :: Python :: 3
19
+ Classifier: Programming Language :: Python :: 3 :: Only
20
+ Classifier: Programming Language :: Python :: 3.12
21
+ Classifier: Programming Language :: Python :: 3.13
22
+ Classifier: Programming Language :: Python :: 3.14
23
+ Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
24
+ Classifier: Topic :: Security
25
+ Classifier: Typing :: Typed
26
+ Requires-Python: >=3.12
27
+ Requires-Dist: httpx>=0.27
28
+ Requires-Dist: langchain-core>=1.6
29
+ Requires-Dist: langchain>=1.4
30
+ Requires-Dist: langgraph>=1.2.11
31
+ Requires-Dist: pydantic>=2.12
32
+ Requires-Dist: stamina>=25.1
33
+ Provides-Extra: deepagents
34
+ Requires-Dist: deepagents>=0.7.13; extra == 'deepagents'
35
+ Provides-Extra: openrouter
36
+ Requires-Dist: langchain-openrouter>=0.2.1; extra == 'openrouter'
37
+ Provides-Extra: typesafe
38
+ Requires-Dist: langchain-typesafe>=0.0.1a3; extra == 'typesafe'
39
+ Description-Content-Type: text/markdown
40
+
41
+ # langchain-sync-monitors
42
+
43
+ **A monitor for every step your LangChain agent or Deep Agent proposes, and a control protocol that decides what runs.**
44
+
45
+ <picture>
46
+ <source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/omamori-lab/langchain-sync-monitors/v0.1.0/docs/assets/diagrams/monitored-step-dark.svg">
47
+ <img alt="The untrusted model proposes a step. The monitor, drawn as an omamori charm, reads it through its view and scores its suspicion from 0 to 1. The control protocol chooses one of five outcomes, from the lightest touch to the strongest: allow it, resample it, defer to the trusted model, block it and tell the agent why, or halt. All of this happens inside the middleware, before any of the agent's own tools run; a brushed line marks that boundary. A committed step reaches the tools and a halted run ends, and either way one record joins monitor_log. Tool results come back, and the next model call is a new step." src="https://raw.githubusercontent.com/omamori-lab/langchain-sync-monitors/v0.1.0/docs/assets/diagrams/monitored-step-light.svg">
48
+ </picture>
49
+
50
+ ## Why this exists
51
+
52
+ An agent with tools can do harm in a single step: post a credential, delete a
53
+ file, run a command. AI control research guards against this with a monitor,
54
+ a second model that scores each step, and a protocol that acts on the score.
55
+ This library brings those protocols, from AI Control, Ctrl-Z and Claude Code's
56
+ auto mode, to `create_agent` and `create_deep_agent` as one middleware.
57
+ "Sync" means synchronous in the control sense of
58
+ [Blocking live failures with synchronous monitors](https://blog.redwoodresearch.org/p/blocking-live-failures-with-synchronous)
59
+ (Lucassen and Kaufman, Redwood Research, 2026): a step waits for its
60
+ verdict, so a suspicious step can be stopped before any of the agent's own
61
+ tools run it.
62
+
63
+ ## Install
64
+
65
+ ```console
66
+ pip install "langchain-sync-monitors[openrouter,deepagents]"
67
+ ```
68
+
69
+ | Extra | Adds | Needed for |
70
+ |---|---|---|
71
+ | `openrouter` | `langchain-openrouter` | model strings such as `"openrouter:xiaomi/mimo-v2.6-pro"` |
72
+ | `deepagents` | `deepagents` | `create_deep_agent` and `monitor_subagents` |
73
+ | `typesafe` | `langchain-typesafe` | `TypeSafeDecisionModel` |
74
+
75
+ ## Quick start
76
+
77
+ ```python
78
+ from pathlib import Path
79
+
80
+ from langchain.agents import create_agent
81
+ from langchain_sync_monitors import (
82
+ DeferToResample,
83
+ DeferToTrustedModel,
84
+ LLMMonitor,
85
+ MonitorMiddleware,
86
+ )
87
+
88
+
89
+ def read_file(path: str) -> str:
90
+ """Read a file from the workspace."""
91
+ return Path(path).read_text()
92
+
93
+
94
+ monitor = MonitorMiddleware(
95
+ # The monitor: a chat model that judges every step the agent proposes.
96
+ monitor=LLMMonitor(model="openrouter:xiaomi/mimo-v2.6-pro"),
97
+ # The protocol: sample a suspicious step again, and if every sample stays
98
+ # suspicious, have the trusted model write the step instead.
99
+ protocol=DeferToResample(
100
+ fallback=DeferToTrustedModel(trusted_model="openrouter:xiaomi/mimo-v2.6-flash"),
101
+ # Resample a step that scores 0.6 or more, and flag it for review
102
+ # when even its least suspicious sample scores 0.9 or more.
103
+ defer_threshold=0.6,
104
+ audit_threshold=0.9,
105
+ ),
106
+ )
107
+ # The untrusted model does the work, with the monitor last in the list.
108
+ agent = create_agent(
109
+ "openrouter:deepseek/deepseek-v4.1-flash", tools=[read_file], middleware=[monitor]
110
+ )
111
+ result = agent.invoke({"messages": [{"role": "user", "content": "Summarise q3.md."}]})
112
+ # One record per step: what the protocol did, and whether to review it.
113
+ for record in result["monitor_log"]:
114
+ print(record["step_number"], record["outcome"], record["flagged"])
115
+ ```
116
+
117
+ Three models play three roles: the untrusted agent does the work, the monitor
118
+ judges each step, and the trusted model writes a step when every sample looks
119
+ suspicious. Each is a parameter, and any LangChain chat model works; the
120
+ library picks none for you. Thresholds share one scale from 0 to 1 and warn
121
+ until you set them; calibrate them on honest runs of your own agent. Put the
122
+ monitor last in the middleware list. `check_monitor_placement(middleware=[...])`
123
+ warns about placements known to undermine it, but does not look for the
124
+ `before_agent` and `after_agent` hooks that can lift a halt.
125
+
126
+ ## How it fits together
127
+
128
+ <picture>
129
+ <source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/omamori-lab/langchain-sync-monitors/v0.1.0/docs/assets/diagrams/protocols-dark.svg">
130
+ <img alt="What each protocol does with a step the monitor finds suspicious; a step below the threshold runs unchanged under every protocol. Trusted Monitoring lets the step run and flags it for a person to review. Defer to Trusted has the trusted model write the step instead. Defer to Resample draws up to two fresh samples; if one passes, the least suspicious runs, and if none passes, the fallback writes the step, usually the trusted model. Auto Mode blocks the step and tells the agent why, and the agent tries again; after three blocks in a row, or twenty in the thread with subagents' blocks included, the fallback takes over, which by default halts the run." src="https://raw.githubusercontent.com/omamori-lab/langchain-sync-monitors/v0.1.0/docs/assets/diagrams/protocols-light.svg">
131
+ </picture>
132
+
133
+ Four control protocols, from watching to blocking: [Choose a protocol](https://github.com/omamori-lab/langchain-sync-monitors/blob/v0.1.0/docs/how-to/choose-a-protocol.md) compares them.
134
+
135
+ <picture>
136
+ <source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/omamori-lab/langchain-sync-monitors/v0.1.0/docs/assets/diagrams/monitors-dark.svg">
137
+ <img alt="Three monitor classes cover three kinds of model, each drawn as the charm. LLMMonitor asks a chat model to reason and then score the step, from 0 to 10 by default. GuardModelMonitor has a guard model label the step against a policy you write, and suspicion is the chance the label is unsafe. DecisionModelMonitor asks a decision model yes or no questions, and suspicion is the chance of yes. Each returns the same Verdict, a suspicion from 0 to 1 and a reason, so any monitor works with any control protocol. Three wrappers nest around any monitor: RepeatedMonitor averages several calls, CalibratedMonitor ranks the score among the scores of honest runs, and CascadeMonitor runs a fast monitor on every step and, when the fast one finds the step suspicious, a careful one whose verdict replaces the fast one's." src="https://raw.githubusercontent.com/omamori-lab/langchain-sync-monitors/v0.1.0/docs/assets/diagrams/monitors-light.svg">
138
+ </picture>
139
+
140
+ Monitors for [chat judges](https://github.com/omamori-lab/langchain-sync-monitors/blob/v0.1.0/docs/how-to/use-a-chat-judge.md), [guard models](https://github.com/omamori-lab/langchain-sync-monitors/blob/v0.1.0/docs/how-to/use-a-guard-model.md) and [decision models](https://github.com/omamori-lab/langchain-sync-monitors/blob/v0.1.0/docs/how-to/use-a-decision-model.md), and [wrappers](https://github.com/omamori-lab/langchain-sync-monitors/blob/v0.1.0/docs/how-to/combine-and-calibrate-monitors.md) that average, calibrate or cascade them.
141
+
142
+ <picture>
143
+ <source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/omamori-lab/langchain-sync-monitors/v0.1.0/docs/assets/diagrams/monitor-view-dark.svg">
144
+ <img alt="Every entry of the transcript belongs to one channel. The default view, Channel.ACTIONS, reads USER (every turn of the user, kept verbatim even after summarisation, context notes and feedback), TOOL_CALLS (tool calls and errors) and SUBAGENT_REPORTS (what subagents report), and leaves out REASONING, AGENT_TEXT and TOOL_RESULTS, each one flag away. The proposed step's tool calls are always shown, and its text when it is a final answer; the rest of the step follows the view's channels. A MonitorView chooses the channels, can keep only the most recent entries besides the user's turns, and names the tools whose results are subagent reports. The monitor reads the chosen entries as escaped, tagged text ending with the proposed step, then scores the step." src="https://raw.githubusercontent.com/omamori-lab/langchain-sync-monitors/v0.1.0/docs/assets/diagrams/monitor-view-light.svg">
145
+ </picture>
146
+
147
+ What the monitor reads by default, and each channel it leaves out: [Choose what the monitor reads](https://github.com/omamori-lab/langchain-sync-monitors/blob/v0.1.0/docs/how-to/choose-what-the-monitor-reads.md).
148
+
149
+ <picture>
150
+ <source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/omamori-lab/langchain-sync-monitors/v0.1.0/docs/assets/diagrams/subagents-dark.svg">
151
+ <img alt="In a Deep Agent, the main agent's monitor, monitor[main], reads the user's task. When the main agent calls the task tool, monitor[main] judges that call first, and the subagent then runs under a monitor of its own, monitor[researcher], which reads the task as the parent agent's words, not the user's, and judges each of the subagent's steps. The subagent's final report returns to the main agent. When the subagent returns, halted or not, its records flow back into the main agent's monitor_log, halts included; a subagent whose run raises returns none. monitor_subagents gives every subagent a copy of the monitor, the general-purpose one included." src="https://raw.githubusercontent.com/omamori-lab/langchain-sync-monitors/v0.1.0/docs/assets/diagrams/subagents-light.svg">
152
+ </picture>
153
+
154
+ A monitor in every Deep Agents subagent: [Monitor Deep Agents subagents](https://github.com/omamori-lab/langchain-sync-monitors/blob/v0.1.0/docs/how-to/monitor-deep-agents-subagents.md).
155
+
156
+ ## Limits
157
+
158
+ - **Server tools.** Server-side tools, such as Anthropic's web search or
159
+ OpenAI's code interpreter, run inside the model call: they have run before
160
+ the monitor sees the step, and run again for every sample drawn.
161
+ `ServerToolWarning` warns about the known ones for
162
+ Anthropic, OpenAI and Gemini when it can see them; it cannot see some, such
163
+ as OpenRouter's web plugin.
164
+ - **What the monitor reads.** By default it reads the task, the tool calls, the
165
+ subagent reports and a proposed final answer. It does not read the agent's
166
+ reasoning, the text the agent writes beside a tool call, or raw tool
167
+ output; each is one flag away.
168
+ - **Halts.** A halt stands until a later run brings a message the monitor
169
+ records as new input from the user. A middleware listed before the monitor
170
+ can still lift it, with an untagged human message written from its
171
+ `before_agent` hook, or from its `after_agent` hook without sending the run
172
+ back to the model.
173
+ - **Subagents.** A monitor on a Deep Agent's main agent alone leaves every
174
+ subagent unmonitored, and nothing warns; pass
175
+ `subagents=monitor_subagents(middleware=monitor, subagents=[...])` to
176
+ `create_deep_agent`.
177
+ Subagents that run in parallel do not see each other's blocks, so together
178
+ they can pass Auto Mode's limit on blocks in the thread; their parent counts
179
+ every one of them at its next step. A subagent whose run raises returns no
180
+ records unless the run is resumed.
181
+ - **Time and graph steps.** Each step waits for its verdict and streams whole
182
+ once committed, not token by token. Each monitor adds two graph steps per
183
+ agent step and two per run, which count towards a `recursion_limit` you
184
+ set; see
185
+ [Allow for the graph steps](https://github.com/omamori-lab/langchain-sync-monitors/blob/v0.1.0/docs/how-to/read-the-monitor-log.md#allow-for-the-graph-steps).
186
+
187
+ [Known limits and open paths](https://github.com/omamori-lab/langchain-sync-monitors/blob/v0.1.0/docs/explanation/design.md#known-limits-and-open-paths)
188
+ explains each of these.
189
+
190
+ ## Read the docs
191
+
192
+ | If you want to | Read |
193
+ |---|---|
194
+ | build a monitored agent step by step, with or without an API key | [Monitor your first agent](https://github.com/omamori-lab/langchain-sync-monitors/blob/v0.1.0/docs/tutorials/first-monitored-agent.md) |
195
+ | read the records, the stream events and the logs | [Read the monitor log](https://github.com/omamori-lab/langchain-sync-monitors/blob/v0.1.0/docs/how-to/read-the-monitor-log.md) |
196
+ | find halted and flagged steps in a tracer | [See the monitor's decisions in LangSmith and Langfuse](https://github.com/omamori-lab/langchain-sync-monitors/blob/v0.1.0/docs/how-to/see-decisions-in-langsmith-and-langfuse.md) |
197
+ | look up a class or a keyword | [API](https://omamori-lab.github.io/langchain-sync-monitors/reference/api/) |
198
+ | understand how a step flows, and why | [How the library is built](https://github.com/omamori-lab/langchain-sync-monitors/blob/v0.1.0/docs/explanation/design.md) |
199
+ | see monitored agents run against real models | [Live runs of a monitored agent](https://github.com/omamori-lab/langchain-sync-monitors/blob/v0.1.0/docs/explanation/live-runs.md) |
200
+
201
+ ## Credits
202
+
203
+ The protocols come from AI control research. Trusted Monitoring and Defer to
204
+ Trusted come from [AI Control](https://arxiv.org/abs/2312.06942) (Greenblatt
205
+ et al., 2023), Defer to Resample from
206
+ [Ctrl-Z](https://arxiv.org/abs/2504.10374) (Bhatt et al., 2025), and Auto
207
+ Mode follows
208
+ [How we built Claude Code auto mode](https://www.anthropic.com/engineering/claude-code-auto-mode)
209
+ (Hughes, Anthropic, 2026). [Where the ideas come from](https://github.com/omamori-lab/langchain-sync-monitors/blob/v0.1.0/docs/explanation/background.md)
210
+ credits every source and what it contributed, and
211
+ [`docs/references.bib`](https://github.com/omamori-lab/langchain-sync-monitors/blob/v0.1.0/docs/references.bib) holds the full entries. If you
212
+ use this library in research, please cite the original authors.
213
+
214
+ ## Status and licence
215
+
216
+ Alpha, and released on [PyPI](https://pypi.org/project/langchain-sync-monitors/).
217
+ The monitors, protocols, middleware and Deep Agents support are tested offline
218
+ and tried in [live runs](https://github.com/omamori-lab/langchain-sync-monitors/blob/v0.1.0/docs/explanation/live-runs.md). Before 1.0.0, a minor
219
+ release can break the API; [CHANGELOG.md](https://github.com/omamori-lab/langchain-sync-monitors/blob/v0.1.0/CHANGELOG.md) records each change.
220
+ [CONTRIBUTING.md](https://github.com/omamori-lab/langchain-sync-monitors/blob/v0.1.0/CONTRIBUTING.md) explains how to work on the library, and
221
+ [docs/plans](https://github.com/omamori-lab/langchain-sync-monitors/blob/v0.1.0/docs/plans/README.md) holds each feature's design plan.
222
+
223
+ Released under the MIT licence; see [LICENSE](https://github.com/omamori-lab/langchain-sync-monitors/blob/v0.1.0/LICENSE).
@@ -0,0 +1,183 @@
1
+ # langchain-sync-monitors
2
+
3
+ **A monitor for every step your LangChain agent or Deep Agent proposes, and a control protocol that decides what runs.**
4
+
5
+ <picture>
6
+ <source media="(prefers-color-scheme: dark)" srcset="docs/assets/diagrams/monitored-step-dark.svg">
7
+ <img alt="The untrusted model proposes a step. The monitor, drawn as an omamori charm, reads it through its view and scores its suspicion from 0 to 1. The control protocol chooses one of five outcomes, from the lightest touch to the strongest: allow it, resample it, defer to the trusted model, block it and tell the agent why, or halt. All of this happens inside the middleware, before any of the agent's own tools run; a brushed line marks that boundary. A committed step reaches the tools and a halted run ends, and either way one record joins monitor_log. Tool results come back, and the next model call is a new step." src="docs/assets/diagrams/monitored-step-light.svg">
8
+ </picture>
9
+
10
+ ## Why this exists
11
+
12
+ An agent with tools can do harm in a single step: post a credential, delete a
13
+ file, run a command. AI control research guards against this with a monitor,
14
+ a second model that scores each step, and a protocol that acts on the score.
15
+ This library brings those protocols, from AI Control, Ctrl-Z and Claude Code's
16
+ auto mode, to `create_agent` and `create_deep_agent` as one middleware.
17
+ "Sync" means synchronous in the control sense of
18
+ [Blocking live failures with synchronous monitors](https://blog.redwoodresearch.org/p/blocking-live-failures-with-synchronous)
19
+ (Lucassen and Kaufman, Redwood Research, 2026): a step waits for its
20
+ verdict, so a suspicious step can be stopped before any of the agent's own
21
+ tools run it.
22
+
23
+ ## Install
24
+
25
+ ```console
26
+ pip install "langchain-sync-monitors[openrouter,deepagents]"
27
+ ```
28
+
29
+ | Extra | Adds | Needed for |
30
+ |---|---|---|
31
+ | `openrouter` | `langchain-openrouter` | model strings such as `"openrouter:xiaomi/mimo-v2.6-pro"` |
32
+ | `deepagents` | `deepagents` | `create_deep_agent` and `monitor_subagents` |
33
+ | `typesafe` | `langchain-typesafe` | `TypeSafeDecisionModel` |
34
+
35
+ ## Quick start
36
+
37
+ ```python
38
+ from pathlib import Path
39
+
40
+ from langchain.agents import create_agent
41
+ from langchain_sync_monitors import (
42
+ DeferToResample,
43
+ DeferToTrustedModel,
44
+ LLMMonitor,
45
+ MonitorMiddleware,
46
+ )
47
+
48
+
49
+ def read_file(path: str) -> str:
50
+ """Read a file from the workspace."""
51
+ return Path(path).read_text()
52
+
53
+
54
+ monitor = MonitorMiddleware(
55
+ # The monitor: a chat model that judges every step the agent proposes.
56
+ monitor=LLMMonitor(model="openrouter:xiaomi/mimo-v2.6-pro"),
57
+ # The protocol: sample a suspicious step again, and if every sample stays
58
+ # suspicious, have the trusted model write the step instead.
59
+ protocol=DeferToResample(
60
+ fallback=DeferToTrustedModel(trusted_model="openrouter:xiaomi/mimo-v2.6-flash"),
61
+ # Resample a step that scores 0.6 or more, and flag it for review
62
+ # when even its least suspicious sample scores 0.9 or more.
63
+ defer_threshold=0.6,
64
+ audit_threshold=0.9,
65
+ ),
66
+ )
67
+ # The untrusted model does the work, with the monitor last in the list.
68
+ agent = create_agent(
69
+ "openrouter:deepseek/deepseek-v4.1-flash", tools=[read_file], middleware=[monitor]
70
+ )
71
+ result = agent.invoke({"messages": [{"role": "user", "content": "Summarise q3.md."}]})
72
+ # One record per step: what the protocol did, and whether to review it.
73
+ for record in result["monitor_log"]:
74
+ print(record["step_number"], record["outcome"], record["flagged"])
75
+ ```
76
+
77
+ Three models play three roles: the untrusted agent does the work, the monitor
78
+ judges each step, and the trusted model writes a step when every sample looks
79
+ suspicious. Each is a parameter, and any LangChain chat model works; the
80
+ library picks none for you. Thresholds share one scale from 0 to 1 and warn
81
+ until you set them; calibrate them on honest runs of your own agent. Put the
82
+ monitor last in the middleware list. `check_monitor_placement(middleware=[...])`
83
+ warns about placements known to undermine it, but does not look for the
84
+ `before_agent` and `after_agent` hooks that can lift a halt.
85
+
86
+ ## How it fits together
87
+
88
+ <picture>
89
+ <source media="(prefers-color-scheme: dark)" srcset="docs/assets/diagrams/protocols-dark.svg">
90
+ <img alt="What each protocol does with a step the monitor finds suspicious; a step below the threshold runs unchanged under every protocol. Trusted Monitoring lets the step run and flags it for a person to review. Defer to Trusted has the trusted model write the step instead. Defer to Resample draws up to two fresh samples; if one passes, the least suspicious runs, and if none passes, the fallback writes the step, usually the trusted model. Auto Mode blocks the step and tells the agent why, and the agent tries again; after three blocks in a row, or twenty in the thread with subagents' blocks included, the fallback takes over, which by default halts the run." src="docs/assets/diagrams/protocols-light.svg">
91
+ </picture>
92
+
93
+ Four control protocols, from watching to blocking: [Choose a protocol](docs/how-to/choose-a-protocol.md) compares them.
94
+
95
+ <picture>
96
+ <source media="(prefers-color-scheme: dark)" srcset="docs/assets/diagrams/monitors-dark.svg">
97
+ <img alt="Three monitor classes cover three kinds of model, each drawn as the charm. LLMMonitor asks a chat model to reason and then score the step, from 0 to 10 by default. GuardModelMonitor has a guard model label the step against a policy you write, and suspicion is the chance the label is unsafe. DecisionModelMonitor asks a decision model yes or no questions, and suspicion is the chance of yes. Each returns the same Verdict, a suspicion from 0 to 1 and a reason, so any monitor works with any control protocol. Three wrappers nest around any monitor: RepeatedMonitor averages several calls, CalibratedMonitor ranks the score among the scores of honest runs, and CascadeMonitor runs a fast monitor on every step and, when the fast one finds the step suspicious, a careful one whose verdict replaces the fast one's." src="docs/assets/diagrams/monitors-light.svg">
98
+ </picture>
99
+
100
+ Monitors for [chat judges](docs/how-to/use-a-chat-judge.md), [guard models](docs/how-to/use-a-guard-model.md) and [decision models](docs/how-to/use-a-decision-model.md), and [wrappers](docs/how-to/combine-and-calibrate-monitors.md) that average, calibrate or cascade them.
101
+
102
+ <picture>
103
+ <source media="(prefers-color-scheme: dark)" srcset="docs/assets/diagrams/monitor-view-dark.svg">
104
+ <img alt="Every entry of the transcript belongs to one channel. The default view, Channel.ACTIONS, reads USER (every turn of the user, kept verbatim even after summarisation, context notes and feedback), TOOL_CALLS (tool calls and errors) and SUBAGENT_REPORTS (what subagents report), and leaves out REASONING, AGENT_TEXT and TOOL_RESULTS, each one flag away. The proposed step's tool calls are always shown, and its text when it is a final answer; the rest of the step follows the view's channels. A MonitorView chooses the channels, can keep only the most recent entries besides the user's turns, and names the tools whose results are subagent reports. The monitor reads the chosen entries as escaped, tagged text ending with the proposed step, then scores the step." src="docs/assets/diagrams/monitor-view-light.svg">
105
+ </picture>
106
+
107
+ What the monitor reads by default, and each channel it leaves out: [Choose what the monitor reads](docs/how-to/choose-what-the-monitor-reads.md).
108
+
109
+ <picture>
110
+ <source media="(prefers-color-scheme: dark)" srcset="docs/assets/diagrams/subagents-dark.svg">
111
+ <img alt="In a Deep Agent, the main agent's monitor, monitor[main], reads the user's task. When the main agent calls the task tool, monitor[main] judges that call first, and the subagent then runs under a monitor of its own, monitor[researcher], which reads the task as the parent agent's words, not the user's, and judges each of the subagent's steps. The subagent's final report returns to the main agent. When the subagent returns, halted or not, its records flow back into the main agent's monitor_log, halts included; a subagent whose run raises returns none. monitor_subagents gives every subagent a copy of the monitor, the general-purpose one included." src="docs/assets/diagrams/subagents-light.svg">
112
+ </picture>
113
+
114
+ A monitor in every Deep Agents subagent: [Monitor Deep Agents subagents](docs/how-to/monitor-deep-agents-subagents.md).
115
+
116
+ ## Limits
117
+
118
+ - **Server tools.** Server-side tools, such as Anthropic's web search or
119
+ OpenAI's code interpreter, run inside the model call: they have run before
120
+ the monitor sees the step, and run again for every sample drawn.
121
+ `ServerToolWarning` warns about the known ones for
122
+ Anthropic, OpenAI and Gemini when it can see them; it cannot see some, such
123
+ as OpenRouter's web plugin.
124
+ - **What the monitor reads.** By default it reads the task, the tool calls, the
125
+ subagent reports and a proposed final answer. It does not read the agent's
126
+ reasoning, the text the agent writes beside a tool call, or raw tool
127
+ output; each is one flag away.
128
+ - **Halts.** A halt stands until a later run brings a message the monitor
129
+ records as new input from the user. A middleware listed before the monitor
130
+ can still lift it, with an untagged human message written from its
131
+ `before_agent` hook, or from its `after_agent` hook without sending the run
132
+ back to the model.
133
+ - **Subagents.** A monitor on a Deep Agent's main agent alone leaves every
134
+ subagent unmonitored, and nothing warns; pass
135
+ `subagents=monitor_subagents(middleware=monitor, subagents=[...])` to
136
+ `create_deep_agent`.
137
+ Subagents that run in parallel do not see each other's blocks, so together
138
+ they can pass Auto Mode's limit on blocks in the thread; their parent counts
139
+ every one of them at its next step. A subagent whose run raises returns no
140
+ records unless the run is resumed.
141
+ - **Time and graph steps.** Each step waits for its verdict and streams whole
142
+ once committed, not token by token. Each monitor adds two graph steps per
143
+ agent step and two per run, which count towards a `recursion_limit` you
144
+ set; see
145
+ [Allow for the graph steps](docs/how-to/read-the-monitor-log.md#allow-for-the-graph-steps).
146
+
147
+ [Known limits and open paths](docs/explanation/design.md#known-limits-and-open-paths)
148
+ explains each of these.
149
+
150
+ ## Read the docs
151
+
152
+ | If you want to | Read |
153
+ |---|---|
154
+ | build a monitored agent step by step, with or without an API key | [Monitor your first agent](docs/tutorials/first-monitored-agent.md) |
155
+ | read the records, the stream events and the logs | [Read the monitor log](docs/how-to/read-the-monitor-log.md) |
156
+ | find halted and flagged steps in a tracer | [See the monitor's decisions in LangSmith and Langfuse](docs/how-to/see-decisions-in-langsmith-and-langfuse.md) |
157
+ | look up a class or a keyword | [API](https://omamori-lab.github.io/langchain-sync-monitors/reference/api/) |
158
+ | understand how a step flows, and why | [How the library is built](docs/explanation/design.md) |
159
+ | see monitored agents run against real models | [Live runs of a monitored agent](docs/explanation/live-runs.md) |
160
+
161
+ ## Credits
162
+
163
+ The protocols come from AI control research. Trusted Monitoring and Defer to
164
+ Trusted come from [AI Control](https://arxiv.org/abs/2312.06942) (Greenblatt
165
+ et al., 2023), Defer to Resample from
166
+ [Ctrl-Z](https://arxiv.org/abs/2504.10374) (Bhatt et al., 2025), and Auto
167
+ Mode follows
168
+ [How we built Claude Code auto mode](https://www.anthropic.com/engineering/claude-code-auto-mode)
169
+ (Hughes, Anthropic, 2026). [Where the ideas come from](docs/explanation/background.md)
170
+ credits every source and what it contributed, and
171
+ [`docs/references.bib`](docs/references.bib) holds the full entries. If you
172
+ use this library in research, please cite the original authors.
173
+
174
+ ## Status and licence
175
+
176
+ Alpha, and released on [PyPI](https://pypi.org/project/langchain-sync-monitors/).
177
+ The monitors, protocols, middleware and Deep Agents support are tested offline
178
+ and tried in [live runs](docs/explanation/live-runs.md). Before 1.0.0, a minor
179
+ release can break the API; [CHANGELOG.md](CHANGELOG.md) records each change.
180
+ [CONTRIBUTING.md](CONTRIBUTING.md) explains how to work on the library, and
181
+ [docs/plans](docs/plans/README.md) holds each feature's design plan.
182
+
183
+ Released under the MIT licence; see [LICENSE](LICENSE).