langchain-diffbot 0.1.0__tar.gz → 0.2.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -14,3 +14,6 @@ build/
14
14
 
15
15
  # Claude Code harness runtime artifacts
16
16
  **/.claude/scheduled_tasks.lock
17
+
18
+ # Claude Code personal/local settings
19
+ .claude/settings.local.json
@@ -0,0 +1,379 @@
1
+ Metadata-Version: 2.5
2
+ Name: langchain-diffbot
3
+ Version: 0.2.0
4
+ Summary: LangChain integration for the Diffbot Knowledge Graph and Extract APIs
5
+ Project-URL: Homepage, https://www.diffbot.com/
6
+ Project-URL: Repository, https://github.com/diffbot/langchain-diffbot
7
+ Project-URL: Issues, https://github.com/diffbot/langchain-diffbot/issues
8
+ Author: Diffbot
9
+ License: MIT
10
+ License-File: LICENSE
11
+ Classifier: Development Status :: 4 - Beta
12
+ Classifier: Intended Audience :: Developers
13
+ Classifier: License :: OSI Approved :: MIT License
14
+ Classifier: Programming Language :: Python :: 3
15
+ Classifier: Programming Language :: Python :: 3.10
16
+ Classifier: Programming Language :: Python :: 3.11
17
+ Classifier: Programming Language :: Python :: 3.12
18
+ Classifier: Programming Language :: Python :: 3.13
19
+ Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
20
+ Requires-Python: <4.0,>=3.10
21
+ Requires-Dist: diffbot>=3.0.0
22
+ Requires-Dist: httpx<1.0,>=0.27
23
+ Requires-Dist: langchain-core<2.0,>=1.0
24
+ Provides-Extra: examples
25
+ Requires-Dist: fastapi<1.0,>=0.115; extra == 'examples'
26
+ Requires-Dist: langchain-anthropic<2.0,>=1.4; extra == 'examples'
27
+ Requires-Dist: langchain<2.0,>=1.3; extra == 'examples'
28
+ Requires-Dist: langsmith<1.0,>=0.1; extra == 'examples'
29
+ Requires-Dist: python-dotenv<2.0,>=1.0; extra == 'examples'
30
+ Requires-Dist: uvicorn[standard]<1.0,>=0.30; extra == 'examples'
31
+ Description-Content-Type: text/markdown
32
+
33
+ # langchain-diffbot
34
+
35
+ [![CI](https://github.com/diffbot/langchain-diffbot/actions/workflows/ci.yml/badge.svg)](https://github.com/diffbot/langchain-diffbot/actions/workflows/ci.yml)
36
+
37
+ **LangChain docs:** [https://docs.langchain.com/oss/python/integrations/providers/diffbot](https://docs.langchain.com/oss/python/integrations/providers/diffbot)
38
+
39
+ A thin LangChain integration over the official [`diffbot-python`](https://github.com/diffbot/diffbot-python) SDK. Every Diffbot API gets the closest LangChain primitive:
40
+
41
+ | Diffbot API | LangChain class(es) |
42
+ | --- | --- |
43
+ | Knowledge Graph (DQL) | `DiffbotKnowledgeGraphRetriever`, `DiffbotKnowledgeGraphTool` |
44
+ | Web Search | `DiffbotWebSearchRetriever`, `DiffbotWebSearchTool` |
45
+ | Extract (Analyze) | `DiffbotExtractTool`, `DiffbotExtractLoader` |
46
+ | NLP entities | `DiffbotEntitiesTool` |
47
+ | Crawl | `DiffbotCrawlLoader` |
48
+ | LLM RAG (`ask`) | `ChatDiffbot` (with native streaming), `DiffbotAskTool` |
49
+
50
+ ## Installation
51
+
52
+ ```bash
53
+ pip install langchain-diffbot
54
+ ```
55
+
56
+ ## Authentication & clients
57
+
58
+ Get an API token at https://app.diffbot.com/get-started/.
59
+
60
+ Every component takes a pre-built SDK client — you build a `diffbot.Diffbot` (sync) and/or `diffbot.DiffbotAsync` (async) and pass it via `client=` / `async_client=`. That's the only way to give a component HTTP access, and it keeps configuration in one place: customize the client (token, `timeout`, `transport=`, custom URLs) however the SDK allows, and share one client across many components to reuse a single connection pool. The component uses the client as-is and never closes it — you own its lifecycle.
61
+
62
+ ```python
63
+ import os
64
+ from diffbot import Diffbot
65
+
66
+ db = Diffbot(token=os.environ["DIFFBOT_API_TOKEN"])
67
+ ```
68
+
69
+ Pick your execution mode by which client you build: `Diffbot` for the sync surface (`invoke`, `stream`, `load`), `DiffbotAsync` for the async surface (`ainvoke`, `astream`, `alazy_load`). Pass both if a component is used both ways.
70
+
71
+ ## Quickstart — Knowledge Graph retriever
72
+
73
+ ```python
74
+ import os
75
+ from diffbot import Diffbot
76
+ from langchain_diffbot import DiffbotKnowledgeGraphRetriever
77
+
78
+ db = Diffbot(token=os.environ["DIFFBOT_API_TOKEN"])
79
+ retriever = DiffbotKnowledgeGraphRetriever(client=db, k=5)
80
+ docs = retriever.invoke("type:Organization industries:\"Artificial Intelligence\" location.city.name:\"Boston\"")
81
+ for d in docs:
82
+ print(d.metadata["name"], "—", d.page_content[:120])
83
+ ```
84
+
85
+ The query string is a [DQL (Diffbot Query Language)](https://docs.diffbot.com/reference/dql-quickstart) expression.
86
+
87
+ ## Shaping the output
88
+
89
+ Diffbot KG entities and web-search results are large. Dumping them straight into an LLM prompt can blow past per-minute input-token limits in a single call. Both retrievers expose three shaping knobs:
90
+
91
+ ```python
92
+ import os
93
+ from diffbot import Diffbot
94
+ from langchain_core.documents import Document
95
+ from langchain_diffbot import DiffbotKnowledgeGraphRetriever
96
+
97
+ db = Diffbot(token=os.environ["DIFFBOT_API_TOKEN"])
98
+
99
+ # 1. Project only the top-level fields you care about. Drops everything else
100
+ # from `metadata`. Recommended for agent / tool-use scenarios.
101
+ retriever = DiffbotKnowledgeGraphRetriever(
102
+ client=db,
103
+ k=5,
104
+ fields=["id", "type", "name", "homepageUri", "nbEmployees"],
105
+ )
106
+
107
+ # 2. Choose which field becomes `page_content`. First non-empty value wins.
108
+ retriever = DiffbotKnowledgeGraphRetriever(
109
+ client=db,
110
+ content_fields=["summary", "description", "name"],
111
+ )
112
+
113
+ # 3. For total control, pass a `document_mapper` that turns a raw entity
114
+ # dict into whatever Document shape you want.
115
+ def mapper(entity: dict) -> Document:
116
+ return Document(
117
+ page_content=entity.get("summary", ""),
118
+ metadata={"id": entity["id"], "name": entity["name"]},
119
+ )
120
+
121
+ retriever = DiffbotKnowledgeGraphRetriever(client=db, document_mapper=mapper)
122
+ ```
123
+
124
+ `fields` and `content_fields` are ignored when `document_mapper` is set. The same knobs work on `DiffbotWebSearchRetriever`.
125
+
126
+ ## Web search
127
+
128
+ ```python
129
+ import os
130
+ from diffbot import Diffbot
131
+ from langchain_diffbot import DiffbotWebSearchRetriever
132
+
133
+ db = Diffbot(token=os.environ["DIFFBOT_API_TOKEN"])
134
+ web = DiffbotWebSearchRetriever(client=db, k=5, fields=["title", "pageUrl", "score"])
135
+ docs = web.invoke("diffbot knowledge graph llm grounding")
136
+ ```
137
+
138
+ ## Extract a URL
139
+
140
+ ```python
141
+ import os
142
+ from diffbot import Diffbot
143
+ from langchain_diffbot import DiffbotExtractTool, DiffbotExtractLoader
144
+
145
+ db = Diffbot(token=os.environ["DIFFBOT_API_TOKEN"])
146
+
147
+ # Single URL
148
+ tool = DiffbotExtractTool(client=db)
149
+ page = tool.invoke({"url": "https://www.diffbot.com/products/extract/"})
150
+
151
+ # Batch — yields one Document per URL (the same client is reused)
152
+ loader = DiffbotExtractLoader(client=db, urls=["https://example.com", "https://diffbot.com"])
153
+ for doc in loader.lazy_load():
154
+ print(doc.metadata["title"], doc.page_content[:200])
155
+ ```
156
+
157
+ `DiffbotExtractTool` returns a structured `{"error": ..., "errorCode": ...}` dict when Diffbot reports an extraction failure (200 with `errorCode`), so agents can react and try another URL instead of catching an exception. Auth / rate-limit errors propagate as `diffbot.errors.AuthError` / `RateLimitError`.
158
+
159
+ ## Crawl a site
160
+
161
+ `DiffbotCrawlLoader` drives a Diffbot crawl job and yields one `Document` per crawled URL. The `page_content` is the URL itself (the crawl API surfaces URLs, not page contents) — chain it with `DiffbotExtractLoader` to fetch the content of each URL.
162
+
163
+ ```python
164
+ import os
165
+ from diffbot import Diffbot
166
+ from langchain_diffbot import DiffbotCrawlLoader
167
+
168
+ db = Diffbot(token=os.environ["DIFFBOT_API_TOKEN"])
169
+ loader = DiffbotCrawlLoader(client=db, site="https://www.diffbot.com")
170
+ for doc in loader.lazy_load():
171
+ print(doc.metadata["url"], doc.metadata["status"])
172
+ ```
173
+
174
+ ## ChatDiffbot
175
+
176
+ ```python
177
+ import os
178
+ from diffbot import Diffbot
179
+ from langchain.messages import HumanMessage
180
+ from langchain_diffbot import ChatDiffbot
181
+
182
+ llm = ChatDiffbot(client=Diffbot(token=os.environ["DIFFBOT_API_TOKEN"]))
183
+
184
+ for chunk in llm.stream([HumanMessage(content="What is the Diffbot Knowledge Graph?")]):
185
+ print(chunk.content, end="", flush=True)
186
+ ```
187
+
188
+ `_stream` / `_astream` are native — no thread-pool fallback. `.invoke()` aggregates the stream into a single message.
189
+
190
+ To let a tool-calling agent *consult* Diffbot's LLM (rather than use it as the primary model), hand it `DiffbotAskTool` instead — it answers a natural-language question from the Knowledge Graph + live web and returns a synthesized string:
191
+
192
+ ```python
193
+ import os
194
+ from diffbot import Diffbot
195
+ from langchain_diffbot import DiffbotAskTool
196
+
197
+ ask = DiffbotAskTool(client=Diffbot(token=os.environ["DIFFBOT_API_TOKEN"]))
198
+ print(ask.invoke({"question": "Who founded Diffbot, and when?"}))
199
+ ```
200
+
201
+ ## Agent tools
202
+
203
+ Every Diffbot API is also exposed as an agent-callable `BaseTool`. Hand a tool-calling agent only the tools you want — they all share whatever client you pass. `DiffbotExtractTool` and `DiffbotAskTool` are shown above; the rest:
204
+
205
+ ### DiffbotWebSearchTool
206
+
207
+ Runs a [Diffbot web search](https://docs.diffbot.com/reference/web-search-get) and returns the result list — each item with `title`, `pageUrl`, `score`, and `content`. New accounts include 100,000 free web searches per month. (Use `DiffbotWebSearchRetriever` when you want `Document` output instead.)
208
+
209
+ ```python
210
+ import os
211
+ from diffbot import Diffbot
212
+ from langchain_diffbot import DiffbotWebSearchTool
213
+
214
+ tool = DiffbotWebSearchTool(client=Diffbot(token=os.environ["DIFFBOT_API_TOKEN"]))
215
+ results = tool.invoke({"text": "diffbot knowledge graph"})
216
+ ```
217
+
218
+ ### DiffbotKnowledgeGraphTool
219
+
220
+ Runs a DQL query against the Knowledge Graph from within an agent and returns the raw response dict. (Use `DiffbotKnowledgeGraphRetriever` when you want `Document` output instead.)
221
+
222
+ ```python
223
+ import os
224
+ from diffbot import Diffbot
225
+ from langchain_diffbot import DiffbotKnowledgeGraphTool
226
+
227
+ tool = DiffbotKnowledgeGraphTool(client=Diffbot(token=os.environ["DIFFBOT_API_TOKEN"]))
228
+ body = tool.invoke({"query": 'type:Organization name:"Diffbot"'})
229
+ ```
230
+
231
+ ### DiffbotEntitiesTool
232
+
233
+ Identifies named entities and sentiment in text via Diffbot's NLP API. The returned entity IDs can be looked up in the Knowledge Graph (e.g. `id:or("E1","E2")`).
234
+
235
+ ```python
236
+ import os
237
+ from diffbot import Diffbot
238
+ from langchain_diffbot import DiffbotEntitiesTool
239
+
240
+ tool = DiffbotEntitiesTool(client=Diffbot(token=os.environ["DIFFBOT_API_TOKEN"]))
241
+ result = tool.invoke({"text": "Diffbot was founded in Menlo Park."})
242
+ ```
243
+
244
+ ### Authoring DQL on the fly: DiffbotOntologyTool + DiffbotDQLProbeTool
245
+
246
+ So an agent can build valid DQL instead of guessing field names, two tools wrap Diffbot's DQL-authoring helpers. The intended loop is **introspect (ontology) → probe → run (`DiffbotKnowledgeGraphTool`) → refine**.
247
+
248
+ `DiffbotOntologyTool` navigates the KG ontology — discover real entity types, field paths, taxonomy, and enum values before querying. The ontology is fetched once over HTTP and cached on the tool instance for the rest of its lifetime.
249
+
250
+ ```python
251
+ import os
252
+ from diffbot import Diffbot
253
+ from langchain_diffbot import DiffbotOntologyTool
254
+
255
+ tool = DiffbotOntologyTool(client=Diffbot(token=os.environ["DIFFBOT_API_TOKEN"]))
256
+ types = tool.invoke({"op": "types"})
257
+ ```
258
+
259
+ `DiffbotDQLProbeTool` probes query variants at `size=0` (hit counts only), so an agent can check selectivity — not zero, not millions — before committing to a full query.
260
+
261
+ ```python
262
+ import os
263
+ from diffbot import Diffbot
264
+ from langchain_diffbot import DiffbotDQLProbeTool
265
+
266
+ tool = DiffbotDQLProbeTool(client=Diffbot(token=os.environ["DIFFBOT_API_TOKEN"]))
267
+ counts = tool.invoke({"queries": ['type:Organization name:"Diffbot"', "type:Person"]})
268
+ ```
269
+
270
+ ## Using a retriever in a chain
271
+
272
+ The retrievers are standard `BaseRetriever`s, so they slot into LCEL like any other:
273
+
274
+ ```python
275
+ import os
276
+ from diffbot import Diffbot
277
+ from langchain_anthropic import ChatAnthropic
278
+ from langchain_core.output_parsers import StrOutputParser
279
+ from langchain_core.prompts import ChatPromptTemplate
280
+ from langchain_core.runnables import RunnablePassthrough
281
+ from langchain_diffbot import DiffbotKnowledgeGraphRetriever
282
+
283
+ retriever = DiffbotKnowledgeGraphRetriever(
284
+ client=Diffbot(token=os.environ["DIFFBOT_API_TOKEN"]),
285
+ k=5,
286
+ fields=["id", "name", "homepageUri", "nbEmployees", "industries"],
287
+ )
288
+
289
+ prompt = ChatPromptTemplate.from_template(
290
+ "Answer using only this Diffbot KG context:\n\n{context}\n\nQuestion: {question}"
291
+ )
292
+
293
+
294
+ def _format(docs):
295
+ return "\n---\n".join(
296
+ f"{d.metadata.get('name')} (id={d.metadata.get('id')}): {d.page_content}"
297
+ for d in docs
298
+ )
299
+
300
+
301
+ chain = (
302
+ {"context": retriever | _format, "question": RunnablePassthrough()}
303
+ | prompt
304
+ | ChatAnthropic(model="claude-sonnet-4-6")
305
+ | StrOutputParser()
306
+ )
307
+
308
+ chain.invoke('type:Organization location.city.name:"Boston" industries:"Biotech"')
309
+ ```
310
+
311
+ ## Sharing a client across components
312
+
313
+ Because every component takes a client, you configure the SDK once and hand the same client to as many components as you like — they share its connection pool, and there's no per-call pool churn. Build the tools/retrievers you actually want and add only those to your agent; the client is the shared resource, not a bundle.
314
+
315
+ ```python
316
+ import os
317
+ from diffbot import Diffbot
318
+ from langchain_diffbot import (
319
+ DiffbotKnowledgeGraphTool,
320
+ DiffbotAskTool,
321
+ DiffbotWebSearchRetriever,
322
+ )
323
+
324
+ # One client, configured once (timeout, transport, custom URLs, ...),
325
+ # shared across every component.
326
+ db = Diffbot(token=os.environ["DIFFBOT_API_TOKEN"], timeout=60.0)
327
+
328
+ kg = DiffbotKnowledgeGraphTool(client=db)
329
+ ask = DiffbotAskTool(client=db)
330
+ web = DiffbotWebSearchRetriever(client=db, k=5)
331
+
332
+ # `db.close()` when you're done — the components never close it for you.
333
+ ```
334
+
335
+ Anything the SDK supports (custom URLs, `transport=`, headers via a custom transport) is configured on the client you build — there's no second configuration surface to learn. For async, build a `diffbot.DiffbotAsync` and pass `async_client=` instead (or both, if a component is used both ways).
336
+
337
+ ## Components reference
338
+
339
+ | Class | Abstraction | Import path |
340
+ |-------|-------------|-------------|
341
+ | `ChatDiffbot` | Chat model | `from langchain_diffbot import ChatDiffbot` |
342
+ | `DiffbotKnowledgeGraphRetriever` | Retriever | `from langchain_diffbot import DiffbotKnowledgeGraphRetriever` |
343
+ | `DiffbotWebSearchRetriever` | Retriever | `from langchain_diffbot import DiffbotWebSearchRetriever` |
344
+ | `DiffbotExtractLoader` | Document loader | `from langchain_diffbot import DiffbotExtractLoader` |
345
+ | `DiffbotCrawlLoader` | Document loader | `from langchain_diffbot import DiffbotCrawlLoader` |
346
+ | `DiffbotExtractTool` | Tool | `from langchain_diffbot import DiffbotExtractTool` |
347
+ | `DiffbotWebSearchTool` | Tool | `from langchain_diffbot import DiffbotWebSearchTool` |
348
+ | `DiffbotKnowledgeGraphTool` | Tool | `from langchain_diffbot import DiffbotKnowledgeGraphTool` |
349
+ | `DiffbotEntitiesTool` | Tool | `from langchain_diffbot import DiffbotEntitiesTool` |
350
+ | `DiffbotAskTool` | Tool | `from langchain_diffbot import DiffbotAskTool` |
351
+ | `DiffbotOntologyTool` | Tool | `from langchain_diffbot import DiffbotOntologyTool` |
352
+ | `DiffbotDQLProbeTool` | Tool | `from langchain_diffbot import DiffbotDQLProbeTool` |
353
+
354
+ ## Examples
355
+
356
+ The [`examples/`](./examples) folder has runnable demos:
357
+
358
+ - [`examples/quickstart/`](./examples/quickstart) — full tour: every public class, output shaping, async, and a multi-tool research agent.
359
+ - [`examples/company_research/`](./examples/company_research) — the same multi-tool agent as a one-shot CLI: `cd examples && python -m company_research "your question"`. The agent combines KG search + web search + URL extract.
360
+
361
+ Both need `langchain` + `langchain-anthropic` on top of the base package — install the extra:
362
+
363
+ ```bash
364
+ pip install "langchain-diffbot[examples]"
365
+ ```
366
+
367
+ ## Development
368
+
369
+ ```bash
370
+ uv sync --all-groups
371
+ uv run pytest tests/unit_tests
372
+ ```
373
+
374
+ Integration tests hit the live Diffbot API and require `DIFFBOT_API_TOKEN`:
375
+
376
+ ```bash
377
+ uv run pytest tests/integration_tests
378
+ ```
379
+