aicordon-langchain 0.1.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (29) hide show
  1. aicordon_langchain-0.1.0/.gitignore +41 -0
  2. aicordon_langchain-0.1.0/CHANGELOG.md +29 -0
  3. aicordon_langchain-0.1.0/LICENSE +202 -0
  4. aicordon_langchain-0.1.0/NOTES.md +112 -0
  5. aicordon_langchain-0.1.0/PKG-INFO +265 -0
  6. aicordon_langchain-0.1.0/README.md +242 -0
  7. aicordon_langchain-0.1.0/eval/_pools.py +113 -0
  8. aicordon_langchain-0.1.0/eval/costturn.py +185 -0
  9. aicordon_langchain-0.1.0/eval/measure_agent.py +125 -0
  10. aicordon_langchain-0.1.0/eval/measure_ingest.py +100 -0
  11. aicordon_langchain-0.1.0/eval/measure_tools.py +126 -0
  12. aicordon_langchain-0.1.0/eval/result-agent.json +12 -0
  13. aicordon_langchain-0.1.0/eval/result-costturn.json +82 -0
  14. aicordon_langchain-0.1.0/eval/result-ingest-disclose.json +13 -0
  15. aicordon_langchain-0.1.0/eval/result-ingest.json +13 -0
  16. aicordon_langchain-0.1.0/eval/result-tools.json +13 -0
  17. aicordon_langchain-0.1.0/example/agent.py +75 -0
  18. aicordon_langchain-0.1.0/example/ingest.py +44 -0
  19. aicordon_langchain-0.1.0/pyproject.toml +50 -0
  20. aicordon_langchain-0.1.0/src/aicordon_langchain/__init__.py +22 -0
  21. aicordon_langchain-0.1.0/src/aicordon_langchain/_common.py +27 -0
  22. aicordon_langchain-0.1.0/src/aicordon_langchain/chain.py +143 -0
  23. aicordon_langchain-0.1.0/src/aicordon_langchain/documents.py +92 -0
  24. aicordon_langchain-0.1.0/src/aicordon_langchain/middleware.py +320 -0
  25. aicordon_langchain-0.1.0/tests/conftest.py +75 -0
  26. aicordon_langchain-0.1.0/tests/test_chain.py +82 -0
  27. aicordon_langchain-0.1.0/tests/test_documents.py +88 -0
  28. aicordon_langchain-0.1.0/tests/test_middleware_request.py +143 -0
  29. aicordon_langchain-0.1.0/tests/test_middleware_tools.py +121 -0
@@ -0,0 +1,41 @@
1
+ # Secrets never live next to code. This line is a fuse, not a description of what is here.
2
+ .env
3
+ .env.*
4
+
5
+ __pycache__/
6
+ *.py[cod]
7
+ .venv/
8
+ dist/
9
+ build/
10
+ *.egg-info/
11
+
12
+ # The injection bank and any datasets never travel: the bank is both a resource of the project and
13
+ # its only honest held-out test. Insurance against an accidental `git add`, not a precaution in the
14
+ # abstract — memory is not the thing to rely on here.
15
+ datasets/
16
+ data/pools/
17
+ *.jsonl
18
+
19
+ # Three exceptions, each of them ours end to end and rebuildable from the script beside it: the ten
20
+ # demonstration exchanges and the two ground-truth manifests. Without them a clone can neither read
21
+ # the example sets nor run `show_material.py` / `show_request.py` over them.
22
+ !integrations/examples/dialogues.jsonl
23
+ !integrations/examples/dialog_manifest.jsonl
24
+ !integrations/examples/manifest.jsonl
25
+
26
+ # The sources the base is built from stay out: the slot dictionaries and the rule file are at once
27
+ # the recipe for the base and the instructions for walking around it. The product does not claim its
28
+ # rules are readable; what ships is the built base, one file (`src/aicordon/picket/data/*.bin`).
29
+ src/aicordon/picket/data/slots/
30
+ src/aicordon/picket/data/*.json
31
+
32
+ # Research notes live in the research repository and do not travel. This line is not caution in the
33
+ # abstract: an append to RESULTS.md from this directory has already dropped fifty lines of
34
+ # measurement prose in here once. Better the mistake is caught by git than by eye.
35
+ RESULTS.md
36
+ HANDOFF.md
37
+ PLAN*.md
38
+
39
+ # Working notes and release plans: ours, not part of what ships. Same rule as the research
40
+ # notes above -- the repository holds the product, not the scaffolding around a release.
41
+ TODO/
@@ -0,0 +1,29 @@
1
+ # Changelog
2
+
3
+ ## 0.1.0 — 2026-08-26
4
+
5
+ First release. AI Cordon Picket for LangChain, in four places a string reaches a model.
6
+
7
+ | entry point | host contract | reads | with |
8
+ |---|---|---|---|
9
+ | `PromptInjectionFilter` | `BaseDocumentTransformer` | documents at ingest | `ipi` |
10
+ | `ToolOutputFilter` | `AgentMiddleware.wrap_tool_call` | what a tool handed back | `ipi` |
11
+ | `PromptInjectionGuard` | `AgentMiddleware.wrap_model_call` | the turn an agent will answer | `dpi` |
12
+ | `PromptInjectionValidator` | `Runnable` | the request in a chain | `dpi` |
13
+
14
+ The policy is `aicordon.guard`, shipped inside the detector: modes, cut boundaries and metadata are
15
+ the same here as in `aicordon-haystack`, and the acceptance measurements reproduce that package's
16
+ numbers on the same corpora.
17
+
18
+ **The tool result is the surface a RAG pipeline has no equivalent of.** A page a tool fetches enters
19
+ the conversation with nothing between it and the model, and it is material by any reading — so it is
20
+ read with the `ipi` rules at the point the tool returns, where cutting is what the policy was
21
+ measured on. `drop` there withholds the text and keeps the message: every tool call must be answered
22
+ by a result carrying its id.
23
+
24
+ **Nothing rewrites a request.** `PromptInjectionGuard` and `PromptInjectionValidator` take
25
+ `annotate`, `drop` and `fail` only, and a chain link takes no `drop` at all — a `Runnable` returns a
26
+ value and the next link is the model, so it raises or marks and lets a `RunnableBranch` decide.
27
+
28
+ Against `langchain-core` 1.6, `langchain` 1.3 and `langgraph` 1.2. The contract each entry point
29
+ stands on, and the four traps behind these choices, are in `NOTES.md`.
@@ -0,0 +1,202 @@
1
+
2
+ Apache License
3
+ Version 2.0, January 2004
4
+ http://www.apache.org/licenses/
5
+
6
+ TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
7
+
8
+ 1. Definitions.
9
+
10
+ "License" shall mean the terms and conditions for use, reproduction,
11
+ and distribution as defined by Sections 1 through 9 of this document.
12
+
13
+ "Licensor" shall mean the copyright owner or entity authorized by
14
+ the copyright owner that is granting the License.
15
+
16
+ "Legal Entity" shall mean the union of the acting entity and all
17
+ other entities that control, are controlled by, or are under common
18
+ control with that entity. For the purposes of this definition,
19
+ "control" means (i) the power, direct or indirect, to cause the
20
+ direction or management of such entity, whether by contract or
21
+ otherwise, or (ii) ownership of fifty percent (50%) or more of the
22
+ outstanding shares, or (iii) beneficial ownership of such entity.
23
+
24
+ "You" (or "Your") shall mean an individual or Legal Entity
25
+ exercising permissions granted by this License.
26
+
27
+ "Source" form shall mean the preferred form for making modifications,
28
+ including but not limited to software source code, documentation
29
+ source, and configuration files.
30
+
31
+ "Object" form shall mean any form resulting from mechanical
32
+ transformation or translation of a Source form, including but
33
+ not limited to compiled object code, generated documentation,
34
+ and conversions to other media types.
35
+
36
+ "Work" shall mean the work of authorship, whether in Source or
37
+ Object form, made available under the License, as indicated by a
38
+ copyright notice that is included in or attached to the work
39
+ (an example is provided in the Appendix below).
40
+
41
+ "Derivative Works" shall mean any work, whether in Source or Object
42
+ form, that is based on (or derived from) the Work and for which the
43
+ editorial revisions, annotations, elaborations, or other modifications
44
+ represent, as a whole, an original work of authorship. For the purposes
45
+ of this License, Derivative Works shall not include works that remain
46
+ separable from, or merely link (or bind by name) to the interfaces of,
47
+ the Work and Derivative Works thereof.
48
+
49
+ "Contribution" shall mean any work of authorship, including
50
+ the original version of the Work and any modifications or additions
51
+ to that Work or Derivative Works thereof, that is intentionally
52
+ submitted to Licensor for inclusion in the Work by the copyright owner
53
+ or by an individual or Legal Entity authorized to submit on behalf of
54
+ the copyright owner. For the purposes of this definition, "submitted"
55
+ means any form of electronic, verbal, or written communication sent
56
+ to the Licensor or its representatives, including but not limited to
57
+ communication on electronic mailing lists, source code control systems,
58
+ and issue tracking systems that are managed by, or on behalf of, the
59
+ Licensor for the purpose of discussing and improving the Work, but
60
+ excluding communication that is conspicuously marked or otherwise
61
+ designated in writing by the copyright owner as "Not a Contribution."
62
+
63
+ "Contributor" shall mean Licensor and any individual or Legal Entity
64
+ on behalf of whom a Contribution has been received by Licensor and
65
+ subsequently incorporated within the Work.
66
+
67
+ 2. Grant of Copyright License. Subject to the terms and conditions of
68
+ this License, each Contributor hereby grants to You a perpetual,
69
+ worldwide, non-exclusive, no-charge, royalty-free, irrevocable
70
+ copyright license to reproduce, prepare Derivative Works of,
71
+ publicly display, publicly perform, sublicense, and distribute the
72
+ Work and such Derivative Works in Source or Object form.
73
+
74
+ 3. Grant of Patent License. Subject to the terms and conditions of
75
+ this License, each Contributor hereby grants to You a perpetual,
76
+ worldwide, non-exclusive, no-charge, royalty-free, irrevocable
77
+ (except as stated in this section) patent license to make, have made,
78
+ use, offer to sell, sell, import, and otherwise transfer the Work,
79
+ where such license applies only to those patent claims licensable
80
+ by such Contributor that are necessarily infringed by their
81
+ Contribution(s) alone or by combination of their Contribution(s)
82
+ with the Work to which such Contribution(s) was submitted. If You
83
+ institute patent litigation against any entity (including a
84
+ cross-claim or counterclaim in a lawsuit) alleging that the Work
85
+ or a Contribution incorporated within the Work constitutes direct
86
+ or contributory patent infringement, then any patent licenses
87
+ granted to You under this License for that Work shall terminate
88
+ as of the date such litigation is filed.
89
+
90
+ 4. Redistribution. You may reproduce and distribute copies of the
91
+ Work or Derivative Works thereof in any medium, with or without
92
+ modifications, and in Source or Object form, provided that You
93
+ meet the following conditions:
94
+
95
+ (a) You must give any other recipients of the Work or
96
+ Derivative Works a copy of this License; and
97
+
98
+ (b) You must cause any modified files to carry prominent notices
99
+ stating that You changed the files; and
100
+
101
+ (c) You must retain, in the Source form of any Derivative Works
102
+ that You distribute, all copyright, patent, trademark, and
103
+ attribution notices from the Source form of the Work,
104
+ excluding those notices that do not pertain to any part of
105
+ the Derivative Works; and
106
+
107
+ (d) If the Work includes a "NOTICE" text file as part of its
108
+ distribution, then any Derivative Works that You distribute must
109
+ include a readable copy of the attribution notices contained
110
+ within such NOTICE file, excluding those notices that do not
111
+ pertain to any part of the Derivative Works, in at least one
112
+ of the following places: within a NOTICE text file distributed
113
+ as part of the Derivative Works; within the Source form or
114
+ documentation, if provided along with the Derivative Works; or,
115
+ within a display generated by the Derivative Works, if and
116
+ wherever such third-party notices normally appear. The contents
117
+ of the NOTICE file are for informational purposes only and
118
+ do not modify the License. You may add Your own attribution
119
+ notices within Derivative Works that You distribute, alongside
120
+ or as an addendum to the NOTICE text from the Work, provided
121
+ that such additional attribution notices cannot be construed
122
+ as modifying the License.
123
+
124
+ You may add Your own copyright statement to Your modifications and
125
+ may provide additional or different license terms and conditions
126
+ for use, reproduction, or distribution of Your modifications, or
127
+ for any such Derivative Works as a whole, provided Your use,
128
+ reproduction, and distribution of the Work otherwise complies with
129
+ the conditions stated in this License.
130
+
131
+ 5. Submission of Contributions. Unless You explicitly state otherwise,
132
+ any Contribution intentionally submitted for inclusion in the Work
133
+ by You to the Licensor shall be under the terms and conditions of
134
+ this License, without any additional terms or conditions.
135
+ Notwithstanding the above, nothing herein shall supersede or modify
136
+ the terms of any separate license agreement you may have executed
137
+ with Licensor regarding such Contributions.
138
+
139
+ 6. Trademarks. This License does not grant permission to use the trade
140
+ names, trademarks, service marks, or product names of the Licensor,
141
+ except as required for reasonable and customary use in describing the
142
+ origin of the Work and reproducing the content of the NOTICE file.
143
+
144
+ 7. Disclaimer of Warranty. Unless required by applicable law or
145
+ agreed to in writing, Licensor provides the Work (and each
146
+ Contributor provides its Contributions) on an "AS IS" BASIS,
147
+ WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
148
+ implied, including, without limitation, any warranties or conditions
149
+ of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
150
+ PARTICULAR PURPOSE. You are solely responsible for determining the
151
+ appropriateness of using or redistributing the Work and assume any
152
+ risks associated with Your exercise of permissions under this License.
153
+
154
+ 8. Limitation of Liability. In no event and under no legal theory,
155
+ whether in tort (including negligence), contract, or otherwise,
156
+ unless required by applicable law (such as deliberate and grossly
157
+ negligent acts) or agreed to in writing, shall any Contributor be
158
+ liable to You for damages, including any direct, indirect, special,
159
+ incidental, or consequential damages of any character arising as a
160
+ result of this License or out of the use or inability to use the
161
+ Work (including but not limited to damages for loss of goodwill,
162
+ work stoppage, computer failure or malfunction, or any and all
163
+ other commercial damages or losses), even if such Contributor
164
+ has been advised of the possibility of such damages.
165
+
166
+ 9. Accepting Warranty or Additional Liability. While redistributing
167
+ the Work or Derivative Works thereof, You may choose to offer,
168
+ and charge a fee for, acceptance of support, warranty, indemnity,
169
+ or other liability obligations and/or rights consistent with this
170
+ License. However, in accepting such obligations, You may act only
171
+ on Your own behalf and on Your sole responsibility, not on behalf
172
+ of any other Contributor, and only if You agree to indemnify,
173
+ defend, and hold each Contributor harmless for any liability
174
+ incurred by, or claims asserted against, such Contributor by reason
175
+ of your accepting any such warranty or additional liability.
176
+
177
+ END OF TERMS AND CONDITIONS
178
+
179
+ APPENDIX: How to apply the Apache License to your work.
180
+
181
+ To apply the Apache License to your work, attach the following
182
+ boilerplate notice, with the fields enclosed by brackets "[]"
183
+ replaced with your own identifying information. (Don't include
184
+ the brackets!) The text should be enclosed in the appropriate
185
+ comment syntax for the file format. We also recommend that a
186
+ file or class name and description of purpose be included on the
187
+ same "printed page" as the copyright notice for easier
188
+ identification within third-party archives.
189
+
190
+ Copyright 2026 Mikhail Gribov
191
+
192
+ Licensed under the Apache License, Version 2.0 (the "License");
193
+ you may not use this file except in compliance with the License.
194
+ You may obtain a copy of the License at
195
+
196
+ http://www.apache.org/licenses/LICENSE-2.0
197
+
198
+ Unless required by applicable law or agreed to in writing, software
199
+ distributed under the License is distributed on an "AS IS" BASIS,
200
+ WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
201
+ See the License for the specific language governing permissions and
202
+ limitations under the License.
@@ -0,0 +1,112 @@
1
+ # LangChain: the contract, established from the sources
2
+
3
+ Checked 2026-08-26. The clone is `langchain-ai/langchain` at `502b2b4`, which carries
4
+ `langchain-core` 1.6.1 and `langchain` 1.3.17 in development; the wrapper is written against the
5
+ released `langchain-core` 1.6.0, `langchain` 1.3.17 and `langgraph` 1.2.11, which is what the venv
6
+ here has. Repository MIT, the framework's own.
7
+
8
+ **The framework moved under the name.** `langchain` 1.x is an agent library: `create_agent` plus
9
+ middleware. What used to be `langchain` — chains, retrievers, the whole 0.3 surface — was renamed
10
+ `langchain-classic` (1.0.8) and left where it was. So an integration written for "LangChain" today
11
+ has to answer for two shapes, the agent and the chain, and this package carries one entry per shape.
12
+
13
+ ## Where we sit
14
+
15
+ Four slots, because a string reaches a model by four routes and the role it plays differs.
16
+
17
+ | slot | host contract | ours | rules |
18
+ |---|---|---|---|
19
+ | documents at ingest | `BaseDocumentTransformer.transform_documents` | `PromptInjectionFilter` | `ipi` |
20
+ | a tool's answer | `AgentMiddleware.wrap_tool_call` | `ToolOutputFilter` | `ipi` |
21
+ | the turn an agent will answer | `AgentMiddleware.wrap_model_call` | `PromptInjectionGuard` | `dpi` |
22
+ | the request in a chain | `Runnable` between prompt and model | `PromptInjectionValidator` | `dpi` |
23
+
24
+ The ingest slot is a line of the caller's own code, not a pipeline object: LangChain has no
25
+ `IngestionPipeline`, the loader hands you a list and you pass it on. That makes the transformer the
26
+ plainest of the four and the only one with no host lifecycle to respect.
27
+
28
+ ## What a middleware is obliged to do
29
+
30
+ Source of truth: `langchain/agents/middleware/types.py` and `langchain/agents/factory.py`, plus
31
+ `middleware/pii.py`, which is the closest thing in the tree to what we are doing.
32
+
33
+ * Subclass `AgentMiddleware`, implement the hooks you need, pass the instance in
34
+ `create_agent(middleware=[...])`. First in the list is the outermost layer.
35
+ * `before_model` / `after_model` are nodes: they return state updates or `None`.
36
+ * `wrap_model_call(request, handler)` and `wrap_tool_call(request, handler)` are wrappers: they may
37
+ call the handler more than once, or not at all.
38
+ * Every hook has an async twin (`awrap_tool_call`, `abefore_model`, …). See trap 4 — the asymmetry
39
+ there is sharp.
40
+ * State that goes through a checkpointer must be JSON-serialisable. Our metadata is bools, strings
41
+ and lists.
42
+
43
+ There is **no serialisation of the middleware object itself** — a middleware is constructed in the
44
+ caller's code and never rebuilt from a dict, so the trap that cost the Haystack wrapper a whole
45
+ mechanism (`to_dict`/`from_dict`, settings silently restored to their defaults) does not exist here.
46
+ The `BaseDocumentTransformer` is not `Serializable` either.
47
+
48
+ ## Traps the mock-ups caught (2026-08-26)
49
+
50
+ Each was run against a control on `langchain` 1.3.17; the probes are in the session scratchpad and
51
+ the behaviour each describes is covered by a test in `tests/`.
52
+
53
+ 1. **`jump_to` is ignored unless the hook is decorated.** A `before_model` may end a run by
54
+ returning `{"jump_to": "end"}` — but only if it also carries `@hook_config(can_jump_to=["end"])`.
55
+ The decorator is what makes `_add_middleware_edge` build a conditional edge; without it the
56
+ builder emits `graph.add_edge(name, default_destination)`, the value sits in state unread, and
57
+ the model is called with the attack. Nothing is raised, nothing is logged. Measured: undecorated,
58
+ the model was called once; decorated, zero times. **We use `wrap_model_call` instead**, where
59
+ refusing means not calling the handler — there is no edge to forget.
60
+
61
+ 2. **A tool may answer with a `Command`, and then its text is not in a `ToolMessage`.** A tool that
62
+ writes agent state returns `Command(update={"messages": [ToolMessage(...)]})`, and a wrapper
63
+ matching on `ToolMessage` alone passes it through unread — the text still reaches the model, and
64
+ the metadata does not even record that nobody looked. Measured with one tool of each kind in the
65
+ same run: the plain one was rewritten, the `Command` one was not. We walk `update["messages"]`.
66
+
67
+ 3. **A message copy without the original `id` is appended, not replaced.** The `messages` channel
68
+ reduces by id. Annotating a turn by building a fresh `HumanMessage` leaves BOTH in the list, so
69
+ the model is sent the turn twice — and in a redacting design it would be sent the clean copy and
70
+ the original side by side. Measured: with the id, two messages in the final state; without it,
71
+ three. `model_copy(update=...)` keeps the id, so that is what we use everywhere.
72
+
73
+ 4. **Sync-only hooks are not symmetrical under `ainvoke`.** A sync `before_model` or
74
+ `wrap_model_call` is run in an executor and works. A sync `wrap_tool_call` raises
75
+ `NotImplementedError` from the tools node, mid-run. Both halves of every hook are written out
76
+ here, and both are tested.
77
+
78
+ 5. **`str(message.content)` is not the message text.** Content is either a string or a list of
79
+ blocks; `str()` of the list is a Python repr, so a check over it reads punctuation and dictionary
80
+ keys that nobody sent, and an edit to it destroys the message. `langchain-core` 1.x has
81
+ `message.text`, which joins the text blocks and ignores images and files — that is what we read.
82
+ (`middleware/pii.py` in the framework itself does the `str(content)` thing.)
83
+
84
+ 6. **The two vocabularies disagree about role names.** `HumanMessage.type` is `human` where the
85
+ policy says `user`, `AIMessage.type` is `ai` where it says `assistant`. A role map keyed on the
86
+ host's word matches nothing in the default policy: the guard reads no message at all while
87
+ reporting each one read and clean. `ROLE_OF` in `_common.py` translates; the test that would fail
88
+ without it is `test_the_human_message_is_read_as_the_user_role`.
89
+
90
+ ## Decisions that are ours, not the host's
91
+
92
+ **`drop` on a tool result withholds the text, not the message.** Every tool call must be answered by
93
+ a result carrying its `tool_call_id`; a call left unanswered makes the provider error or re-ask. So
94
+ the message stays and the model is told the output was withheld — which is also the honest thing to
95
+ tell it.
96
+
97
+ **A chain link may not `drop`.** A `Runnable` returns a value and the next link is the model; there
98
+ is no arrangement in which it declines the call and answers instead. `PromptInjectionValidator`
99
+ raises on the mode rather than imitating it, and points at the two real ways to have it: `fail`, or
100
+ a `RunnableBranch` on `.flagged`. In an agent the decision has a proper home.
101
+
102
+ **The request guard reads what arrived since the model last spoke.** `wrap_model_call` runs once per
103
+ model call, so reading the whole history each time would re-bill the opening turn on every step of
104
+ a loop. The tail after the last `AIMessage` is what is new; on the first call there is none, and the
105
+ tail is the whole list — right for a run resumed with a history nobody has read yet.
106
+
107
+ ## What carries over to another framework
108
+
109
+ The policy (`aicordon.guard`) carried over from Haystack unchanged: modes, cut boundaries, metadata,
110
+ the role map, the refusal of editing modes on the request side. What had to be written here is the
111
+ translation and the host's contract — and one surface Haystack has no equivalent of, the tool
112
+ result, which is where an agent takes in material.
@@ -0,0 +1,265 @@
1
+ Metadata-Version: 2.5
2
+ Name: aicordon-langchain
3
+ Version: 0.1.0
4
+ Summary: Check what an LLM is given for prompt injection: material at ingest, tool output, and the turn it answers
5
+ Project-URL: Homepage, https://github.com/AICordon/aicordon/blob/main/integrations/langchain/README.md
6
+ Project-URL: Repository, https://github.com/AICordon/aicordon
7
+ Project-URL: Issues, https://github.com/AICordon/aicordon/issues
8
+ Project-URL: Changelog, https://github.com/AICordon/aicordon/blob/main/integrations/langchain/CHANGELOG.md
9
+ Author-email: Mikhail Gribov <mihail.gribov.rs@gmail.com>
10
+ License-Expression: Apache-2.0
11
+ License-File: LICENSE
12
+ Keywords: agents,guardrails,jailbreak,langchain,llm-security,prompt-injection,rag
13
+ Classifier: Development Status :: 4 - Beta
14
+ Classifier: Intended Audience :: Developers
15
+ Classifier: Programming Language :: Python :: 3
16
+ Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
17
+ Classifier: Topic :: Security
18
+ Requires-Python: >=3.10
19
+ Requires-Dist: aicordon>=1.1.1
20
+ Requires-Dist: langchain-core>=1.0.0
21
+ Requires-Dist: langchain>=1.3.0
22
+ Description-Content-Type: text/markdown
23
+
24
+ # AI Cordon Picket for LangChain
25
+
26
+ Check what an LLM is given for prompt injection, in each place it can arrive:
27
+
28
+ | entry point | reads | with |
29
+ |---|---|---|
30
+ | `PromptInjectionFilter` | **material**: documents at ingest, before they are chunked and embedded | Picket's `ipi` rules |
31
+ | `ToolOutputFilter` | **material**: what a tool handed back, before the model reads it | `ipi` |
32
+ | `PromptInjectionGuard` | **the request**: the turn an agent is about to answer | Picket's `dpi` rules |
33
+ | `PromptInjectionValidator` | **the request**: the same, in a chain that is not an agent | `dpi` |
34
+
35
+ The two rule sets are disjoint, and neither is a stricter version of the other — this is not a
36
+ sensitivity knob. Pick by role: material is what the model works on, the request is what it answers.
37
+ Your code knows which is which; it puts them in different places when it assembles the call.
38
+
39
+ The check is a rule, not a model: no GPU, no network, no key, a few hundred kilobytes of base, and a
40
+ fraction of a millisecond per turn on one core — see [what it costs](#what-it-costs).
41
+
42
+ ## Installation
43
+
44
+ ```bash
45
+ pip install aicordon-langchain
46
+ ```
47
+
48
+ ## Material at ingest
49
+
50
+ ```python
51
+ from aicordon_langchain import PromptInjectionFilter
52
+ from langchain_community.document_loaders import DirectoryLoader
53
+ from langchain_text_splitters import RecursiveCharacterTextSplitter
54
+
55
+ docs = DirectoryLoader("kb/").load()
56
+ docs = PromptInjectionFilter(mode="redact").transform_documents(docs) # <- here
57
+ chunks = RecursiveCharacterTextSplitter().split_documents(docs)
58
+ store.add_documents(chunks)
59
+ ```
60
+
61
+ The filter sits **before the splitter**: a cut here takes the injection out of the chunks, the
62
+ embeddings and the store at once, with no offsets to reconcile across chunk boundaries.
63
+
64
+ ### What it does with a finding
65
+
66
+ | `mode` | the document | the length |
67
+ |---|---|---|
68
+ | `annotate` | indexed unchanged, the finding recorded in metadata | unchanged |
69
+ | `blank` | every character of the block becomes `blank_char` (default `*`) | **preserved** |
70
+ | `mask` | the block is replaced by `mask_with` | changes |
71
+ | `redact` *(default)* | the block is cut out | changes |
72
+ | `drop` | not indexed | — |
73
+ | `fail` | the run stops on the first finding | — |
74
+
75
+ `blank` is for pipelines that carry offsets, page maps or diffs downstream and cannot have a
76
+ document change length under them.
77
+
78
+ The cut takes **the whole utterance the span sits in** — the sentence, across the lines a
79
+ wrapper broke it over: the span points at the injection, but what must leave the index is
80
+ everything it was saying. A short line that ends without punctuation is a bullet or a table
81
+ row and is left alone, so a list is not eaten item by item.
82
+
83
+ In `drop` mode the host's contract lets `transform_documents` return the survivors and nothing else,
84
+ so use `split` when the removed pile should stay visible:
85
+
86
+ ```python
87
+ kept, rejected = PromptInjectionFilter(mode="drop").split(docs)
88
+ ```
89
+
90
+ ## Tool output in an agent
91
+
92
+ ```python
93
+ from aicordon_langchain import ToolOutputFilter
94
+ from langchain.agents import create_agent
95
+
96
+ agent = create_agent(
97
+ model="openai:gpt-5.5",
98
+ tools=[fetch_page, read_ticket],
99
+ middleware=[ToolOutputFilter(mode="redact")], # every tool, or tools=["fetch_page"]
100
+ )
101
+ ```
102
+
103
+ A page a tool brings back is material by any reading — the model is to work on it, not answer it —
104
+ and it enters the conversation with nothing between it and the model. The filter reads it at the
105
+ point the tool returns, which is the same ingest point a document has and the only place where
106
+ cutting is what the policy was measured on.
107
+
108
+ **`drop` here withholds the text, not the message.** Every tool call must be answered by a result
109
+ carrying its id, so the message stays and the model is told the output was withheld — which is also
110
+ the honest thing to tell it.
111
+
112
+ ## The request in an agent
113
+
114
+ ```python
115
+ from aicordon_langchain import PromptInjectionGuard
116
+
117
+ agent = create_agent(
118
+ model="openai:gpt-5.5",
119
+ tools=[fetch_page],
120
+ middleware=[PromptInjectionGuard()], # mode="drop" by default
121
+ )
122
+ ```
123
+
124
+ On a flagged turn the model is **not called**, and the agent answers with the guard's own message
125
+ instead — `mode="annotate"` calls the model and records the finding on the answer, `mode="fail"`
126
+ raises `InjectionFound`. What is read is the request: by default the user's turns, and nothing else.
127
+ The system message is the operator's own text, and an operator who wants to steer their own model
128
+ does not need an injection to do it.
129
+
130
+ Both middlewares in one agent, each on its own side:
131
+
132
+ ```python
133
+ middleware=[PromptInjectionGuard(), ToolOutputFilter(mode="redact")]
134
+ ```
135
+
136
+ ## The request in a chain
137
+
138
+ ```python
139
+ from aicordon_langchain import PromptInjectionValidator
140
+
141
+ chain = prompt | PromptInjectionValidator() | model # raises InjectionFound on a finding
142
+ ```
143
+
144
+ A link in a chain returns a value and the next link is the model; there is no arrangement in which
145
+ it declines the call and answers instead. So it raises, or it marks and lets the chain decide:
146
+
147
+ ```python
148
+ guard = PromptInjectionValidator(mode="annotate")
149
+ chain = prompt | RunnableBranch((guard.flagged, refusal), model)
150
+ ```
151
+
152
+ Inside an agent the same decision has a proper home — prefer `PromptInjectionGuard` when there is
153
+ an agent to put it in.
154
+
155
+ ## Nothing is rewritten on the request side
156
+
157
+ `PromptInjectionGuard` and `PromptInjectionValidator` accept `annotate`, `drop` and `fail` only; ask
158
+ either for `redact` and it raises. Material can lose a paragraph and stay usable. Take a clause out
159
+ of what somebody asked for and the model answers a question nobody put, with the user seeing an
160
+ answer rather than a notice — and the cut itself is fitted to the wrong shape, because a typed
161
+ attack is not spliced into a turn, it *is* the turn.
162
+
163
+ ## What is written where
164
+
165
+ Metadata is written on every text that was read, including the clean ones: "read, clean" and "not
166
+ read" are different facts, and a field that appeared only on a finding could not be filtered on.
167
+
168
+ | surface | where it lands | keys |
169
+ |---|---|---|
170
+ | documents | `Document.metadata` | `ipi_flagged`, `ipi_action`, `ipi_base`, and on a finding `ipi_threats`, `ipi_spans`, `ipi_removed_chars` |
171
+ | tool output | `ToolMessage.response_metadata` | `ipi_flagged`, `ipi_action`, `ipi_base`, `ipi_threats` |
172
+ | the agent's request | the answer's `response_metadata` | `picket_blocked` or `picket_request_flagged`, `picket_request_threats`, `picket_messages` (keyed by message id) |
173
+ | the chain's request | each read message's `additional_kwargs` | `picket_flagged`, `picket_action`, `picket_base`, `picket_threats`, `picket_spans` |
174
+
175
+ Findings are also logged through the standard library logger under `aicordon_langchain.*` at
176
+ `warning` level, in every mode. Writing it down is not a policy choice.
177
+
178
+ ## Measured
179
+
180
+ Not the detector's recall — that ships with the detector — but what your line delivers with this
181
+ package in it and without.
182
+
183
+ **Material at ingest.** [Quadrat-IPI v1.0.1](https://huggingface.co/datasets/mihailgribov/quadrat-ipi),
184
+ 2000 injected and 2000 clean documents, `mode="redact"`; how much of a planted payload still reaches
185
+ the splitter:
186
+
187
+ | | whole corpus | injections that ask the model to **reveal** something |
188
+ |---|---|---|
189
+ | payload gets through intact, without the filter | 100% | 100% |
190
+ | payload gets through intact, with it | **85.2%** | **43.4%** |
191
+ | payload gone without a trace | 13.5% | **51.3%** |
192
+ | clean documents dropped or trimmed | 1 of 2000 | 1 of 2000 |
193
+
194
+ Both columns matter: the first is an arbitrary stream, the second is where the rule is strong.
195
+
196
+ **Material through a tool.** The same corpus, 1000 injected and 1000 clean pages, fetched by a tool
197
+ inside a real agent loop — and read off the message list the MODEL was handed, not off the filter's
198
+ own return value:
199
+
200
+ | | |
201
+ |---|---|
202
+ | payload reaching the model intact, without the filter | 100% (1000 of 1000) |
203
+ | payload reaching the model intact, with it | **85.4%** |
204
+ | payload gone without a trace | 13.1% |
205
+ | clean pages withheld or trimmed | 0 of 1000 |
206
+
207
+ **The request.** Held-out forum jailbreaks from
208
+ [TrustAIRLab in-the-wild](https://huggingface.co/datasets/TrustAIRLab/in-the-wild-jailbreak-prompts)
209
+ (537, near-duplicates of the fitting half removed) against 20 000 real user turns from
210
+ [WildChat](https://huggingface.co/datasets/allenai/WildChat-1M), an agent in `mode="drop"`:
211
+
212
+ | | |
213
+ |---|---|
214
+ | attacks reaching the model, without the guard | 100% (537 of 537) |
215
+ | attacks reaching the model, with it | **65.2%** |
216
+ | turns not answered, out of 20 000 real ones | 0.070% (14) |
217
+ | verdicts differing from the bare detector | **0** |
218
+
219
+ WildChat carries no attack labels and real jailbreaks sit inside it, so "turns not answered" is an
220
+ upper bound on the cost to a real user, not a false-alarm rate. The detector's working point, on a
221
+ labelled pool, is in its report.
222
+
223
+ The last row outranks the other two: a wrapper may neither lose text nor add its own, and a figure
224
+ taken while it does would be describing a different string than the one the user sent.
225
+
226
+ Reproduce all three with `eval/measure_ingest.py`, `eval/measure_tools.py` and
227
+ `eval/measure_agent.py`.
228
+
229
+ ## What it costs
230
+
231
+ Checking a turn costs **0.32 ms at the median length**, and 1.78 ± 0.06 ms averaged over ordinary
232
+ traffic — 3000 real WildChat turns, each timed five times (`eval/costturn.py`). Loading the base
233
+ costs 21 ms, once per process. Of that 1.78 ms, 1.75 is the rule engine itself and 0.03 is
234
+ everything this package adds.
235
+
236
+ The average is five times the median: cost follows length, and a chat pool has a long tail. Find
237
+ your row:
238
+
239
+ | turn length | turns in the pool | cost |
240
+ |---|---|---|
241
+ | under 200 characters | 1943 | 0.21 ms |
242
+ | 200–500 | 411 | 0.71 ms |
243
+ | 500–1500 | 321 | 1.72 ms |
244
+ | 1500–4000 | 182 | 4.62 ms |
245
+ | over 4000 | 143 | 15.63 ms |
246
+
247
+ For scale, an agent step through LangGraph costs 0.57 ms before any middleware is installed. A
248
+ document at ingest costs 11 ms — documents are long, and cost follows length there too.
249
+
250
+ Cost drifts with machine load. These were taken in one run by one procedure — the only way two
251
+ figures compare.
252
+
253
+ No GPU, no network call, no key. The rule base is a few hundred kilobytes and loads once.
254
+
255
+ ## Not a prefilter
256
+
257
+ Silence from a rule is not a verdict. Picket reports what it recognises; what it does not recognise
258
+ it says nothing about, and no finding does not mean no injection. It belongs where a cheap, local,
259
+ deterministic check is worth having on everything you ingest — not as the only thing between a model
260
+ and the web.
261
+
262
+ ## Licence
263
+
264
+ Apache-2.0. The detector itself is [`aicordon`](https://pypi.org/project/aicordon/); the policy the
265
+ wrappers share lives there as `aicordon.guard`.