embedflow 0.1.1__tar.gz → 0.3.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (133) hide show
  1. embedflow-0.3.0/CHANGELOG.md +30 -0
  2. {embedflow-0.1.1 → embedflow-0.3.0}/CITATION.cff +1 -1
  3. {embedflow-0.1.1 → embedflow-0.3.0}/CONTRIBUTING.md +1 -1
  4. {embedflow-0.1.1 → embedflow-0.3.0}/MANIFEST.in +2 -1
  5. {embedflow-0.1.1 → embedflow-0.3.0}/PKG-INFO +33 -11
  6. {embedflow-0.1.1 → embedflow-0.3.0}/README.md +13 -3
  7. {embedflow-0.1.1 → embedflow-0.3.0}/README_PYPI.md +21 -5
  8. {embedflow-0.1.1 → embedflow-0.3.0}/docs/api.md +8 -2
  9. {embedflow-0.1.1 → embedflow-0.3.0}/docs/cli.md +1 -1
  10. {embedflow-0.1.1 → embedflow-0.3.0}/docs/configuration.md +18 -1
  11. {embedflow-0.1.1 → embedflow-0.3.0}/docs/economics.md +1 -1
  12. {embedflow-0.1.1 → embedflow-0.3.0}/docs/installation.md +14 -0
  13. embedflow-0.3.0/docs/integrations/pgvector.md +140 -0
  14. embedflow-0.3.0/docs/integrations/pinecone.md +157 -0
  15. {embedflow-0.1.1 → embedflow-0.3.0}/docs/limitations.md +1 -1
  16. {embedflow-0.1.1 → embedflow-0.3.0}/docs/releasing.md +8 -8
  17. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow/__init__.py +2 -2
  18. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow/analysis.py +5 -1
  19. embedflow-0.3.0/embedflow/api.py +10 -0
  20. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow/cli.py +196 -39
  21. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow/config.py +85 -9
  22. embedflow-0.3.0/embedflow/indexes/__init__.py +11 -0
  23. embedflow-0.3.0/embedflow/indexes/pgvector_backend.py +685 -0
  24. embedflow-0.3.0/embedflow/indexes/pinecone_backend.py +764 -0
  25. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow/migration/facade.py +89 -13
  26. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow/migration/materializer.py +18 -6
  27. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow/runtime.py +61 -6
  28. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow/serving/api.py +13 -4
  29. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow/serving/engine.py +8 -2
  30. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow.egg-info/PKG-INFO +33 -11
  31. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow.egg-info/SOURCES.txt +16 -0
  32. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow.egg-info/requires.txt +12 -4
  33. embedflow-0.3.0/examples/pgvector/README.md +28 -0
  34. embedflow-0.3.0/examples/pgvector/build_index.py +57 -0
  35. embedflow-0.3.0/examples/pgvector/compose.yaml +20 -0
  36. embedflow-0.3.0/examples/pgvector/embedflow.yaml +48 -0
  37. embedflow-0.3.0/examples/pgvector/init.sql +9 -0
  38. embedflow-0.3.0/examples/pgvector/queries.jsonl +4 -0
  39. embedflow-0.3.0/examples/pgvector/run_demo.sh +11 -0
  40. embedflow-0.3.0/examples/pinecone/README.md +32 -0
  41. embedflow-0.3.0/examples/pinecone/run_smoke.sh +8 -0
  42. {embedflow-0.1.1 → embedflow-0.3.0}/pyproject.toml +9 -5
  43. embedflow-0.3.0/scripts/pinecone_smoke.py +54 -0
  44. {embedflow-0.1.1 → embedflow-0.3.0}/scripts/release_gate.py +26 -4
  45. embedflow-0.3.0/scripts/validate_pgvector_10k.py +812 -0
  46. embedflow-0.1.1/CHANGELOG.md +0 -14
  47. embedflow-0.1.1/embedflow/indexes/__init__.py +0 -5
  48. {embedflow-0.1.1 → embedflow-0.3.0}/LICENSE +0 -0
  49. {embedflow-0.1.1 → embedflow-0.3.0}/SECURITY.md +0 -0
  50. {embedflow-0.1.1 → embedflow-0.3.0}/docs/assets/README.md +0 -0
  51. {embedflow-0.1.1 → embedflow-0.3.0}/docs/assets/candidate-gap-example.svg +0 -0
  52. {embedflow-0.1.1 → embedflow-0.3.0}/docs/assets/dashboard-screenshot.md +0 -0
  53. {embedflow-0.1.1 → embedflow-0.3.0}/docs/assets/terminal-demo.txt +0 -0
  54. {embedflow-0.1.1 → embedflow-0.3.0}/docs/concepts.md +0 -0
  55. {embedflow-0.1.1 → embedflow-0.3.0}/docs/contributing-benchmarks.md +0 -0
  56. {embedflow-0.1.1 → embedflow-0.3.0}/docs/integrations/faiss.md +0 -0
  57. {embedflow-0.1.1 → embedflow-0.3.0}/docs/integrations/qdrant.md +0 -0
  58. {embedflow-0.1.1 → embedflow-0.3.0}/docs/methodology.md +0 -0
  59. {embedflow-0.1.1 → embedflow-0.3.0}/docs/quickstart.md +0 -0
  60. {embedflow-0.1.1 → embedflow-0.3.0}/docs/registry.md +0 -0
  61. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow/__main__.py +0 -0
  62. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow/cache/__init__.py +0 -0
  63. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow/cache/base.py +0 -0
  64. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow/cache/persistent_cache.py +0 -0
  65. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow/compatibility/__init__.py +0 -0
  66. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow/compatibility/candidate_gap.py +0 -0
  67. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow/compatibility/containment.py +0 -0
  68. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow/compatibility/evaluate.py +0 -0
  69. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow/compatibility/metrics.py +0 -0
  70. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow/compatibility/migration_depth.py +0 -0
  71. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow/compatibility/probe.py +0 -0
  72. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow/compatibility/report.py +0 -0
  73. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow/compatibility/t2.py +0 -0
  74. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow/data/__init__.py +0 -0
  75. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow/data/registry/__init__.py +0 -0
  76. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow/data/registry/benchmark_profiles.jsonl +0 -0
  77. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow/data/registry/checksums.sha256 +0 -0
  78. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow/data/registry/migrations.jsonl +0 -0
  79. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow/data/registry/registry_manifest.json +0 -0
  80. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow/data/registry/research_summaries.json +0 -0
  81. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow/data/registry/schema_version.json +0 -0
  82. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow/frozen/T2_V1_FROZEN_SPEC.md +0 -0
  83. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow/frozen/T2_V1_FROZEN_SPEC.sha256 +0 -0
  84. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow/indexes/base.py +0 -0
  85. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow/indexes/faiss_backend.py +0 -0
  86. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow/indexes/qdrant_backend.py +0 -0
  87. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow/metrics/__init__.py +0 -0
  88. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow/metrics/latency.py +0 -0
  89. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow/migration/__init__.py +0 -0
  90. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow/migration/compatibility.py +0 -0
  91. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow/migration/planner.py +0 -0
  92. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow/migration/state.py +0 -0
  93. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow/models/__init__.py +0 -0
  94. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow/models/base.py +0 -0
  95. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow/models/huggingface.py +0 -0
  96. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow/registry/__init__.py +0 -0
  97. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow/registry/loader.py +0 -0
  98. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow/registry/matcher.py +0 -0
  99. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow/registry/schema.py +0 -0
  100. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow/serving/__init__.py +0 -0
  101. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow/serving/factory.py +0 -0
  102. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow/serving/schemas.py +0 -0
  103. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow.egg-info/dependency_links.txt +0 -0
  104. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow.egg-info/entry_points.txt +0 -0
  105. {embedflow-0.1.1 → embedflow-0.3.0}/embedflow.egg-info/top_level.txt +0 -0
  106. {embedflow-0.1.1 → embedflow-0.3.0}/examples/faiss/README.md +0 -0
  107. {embedflow-0.1.1 → embedflow-0.3.0}/examples/faiss/documents.jsonl +0 -0
  108. {embedflow-0.1.1 → embedflow-0.3.0}/examples/faiss/embedflow.yaml +0 -0
  109. {embedflow-0.1.1 → embedflow-0.3.0}/examples/faiss/queries.jsonl +0 -0
  110. {embedflow-0.1.1 → embedflow-0.3.0}/examples/qdrant/README.md +0 -0
  111. {embedflow-0.1.1 → embedflow-0.3.0}/examples/qdrant/build_index.py +0 -0
  112. {embedflow-0.1.1 → embedflow-0.3.0}/examples/qdrant/documents.jsonl +0 -0
  113. {embedflow-0.1.1 → embedflow-0.3.0}/examples/qdrant/embedflow.yaml +0 -0
  114. {embedflow-0.1.1 → embedflow-0.3.0}/examples/qdrant/queries.jsonl +0 -0
  115. {embedflow-0.1.1 → embedflow-0.3.0}/examples/research_analysis/README.md +0 -0
  116. {embedflow-0.1.1 → embedflow-0.3.0}/examples/research_analysis/documents.jsonl +0 -0
  117. {embedflow-0.1.1 → embedflow-0.3.0}/examples/research_analysis/embedflow.yaml +0 -0
  118. {embedflow-0.1.1 → embedflow-0.3.0}/examples/research_analysis/qrels.json +0 -0
  119. {embedflow-0.1.1 → embedflow-0.3.0}/examples/research_analysis/queries.jsonl +0 -0
  120. {embedflow-0.1.1 → embedflow-0.3.0}/frozen/T2_V1_FROZEN_SPEC.md +0 -0
  121. {embedflow-0.1.1 → embedflow-0.3.0}/frozen/T2_V1_FROZEN_SPEC.sha256 +0 -0
  122. {embedflow-0.1.1 → embedflow-0.3.0}/requirements-dev.txt +0 -0
  123. {embedflow-0.1.1 → embedflow-0.3.0}/requirements.txt +0 -0
  124. {embedflow-0.1.1 → embedflow-0.3.0}/scripts/real_qdrant_smoke.py +0 -0
  125. {embedflow-0.1.1 → embedflow-0.3.0}/scripts/run_demo.sh +0 -0
  126. {embedflow-0.1.1 → embedflow-0.3.0}/scripts/run_tests.sh +0 -0
  127. {embedflow-0.1.1 → embedflow-0.3.0}/setup.cfg +0 -0
  128. {embedflow-0.1.1 → embedflow-0.3.0}/src/__init__.py +0 -0
  129. {embedflow-0.1.1 → embedflow-0.3.0}/src/embed.py +0 -0
  130. {embedflow-0.1.1 → embedflow-0.3.0}/src/probe_features.py +0 -0
  131. {embedflow-0.1.1 → embedflow-0.3.0}/src/storage.py +0 -0
  132. {embedflow-0.1.1 → embedflow-0.3.0}/src/t2_v1.py +0 -0
  133. {embedflow-0.1.1 → embedflow-0.3.0}/src/utils.py +0 -0
@@ -0,0 +1,30 @@
1
+ # Changelog
2
+
3
+ ## v0.3.0 — Pinecone backend
4
+
5
+ - Added a read-only Pinecone backend for existing dense indexes.
6
+ - Added host targeting, index-name resolution, namespace-aware retrieval, and
7
+ metadata or external-document text resolution.
8
+ - Added Pinecone status/audit integration, optional dependency packaging, and
9
+ unit/integration smoke fixtures.
10
+
11
+ ## v0.2.0 — pgvector backend
12
+
13
+ - Added a read-only pgvector backend for existing PostgreSQL vector tables.
14
+ - Added cosine, Euclidean/L2, and inner-product retrieval with safe SQL
15
+ identifier composition.
16
+ - Added pgvector table auditing, optional transaction-local HNSW/IVFFlat
17
+ settings, Docker example, and numerical/SQL-safety tests.
18
+
19
+ ## v0.1.0 — initial public release
20
+
21
+ - Candidate-compatibility analysis and Mode A evaluation workflows.
22
+ - Frozen T2-v1 finite-tail diagnostic with explicit empirical-warning language.
23
+ - FAISS and Qdrant index adapters.
24
+ - Persistent target-vector cache and progressive materialization queue.
25
+ - FastAPI service, search/status/prewarm endpoints, and dashboard.
26
+ - Latency telemetry, economics projections, report generation, and demos.
27
+ - Public packaging, examples, tests, CI configuration, and contributor docs.
28
+
29
+ This is an early open-source release. See the limitations in the README and
30
+ `docs/concepts.md` before using it for a production migration.
@@ -2,7 +2,7 @@ cff-version: 1.2.0
2
2
  title: "EmbedFlow: Upgrading Legacy Embeddings Without Full Upfront Re-Embedding"
3
3
  message: "If EmbedFlow contributes to your work, please cite this software release."
4
4
  type: software
5
- version: 0.1.1
5
+ version: 0.3.0
6
6
  date-released: 2026-09-06
7
7
  repository-code: "https://github.com/arnsri33/embedflow"
8
8
  url: "https://github.com/arnsri33/embedflow"
@@ -11,7 +11,7 @@ python -m pip install -e '.[dev]'
11
11
  ```
12
12
 
13
13
  Optional integrations can be installed with `.[faiss]`, `.[qdrant]`,
14
- `.[models]`, or `.[dashboard]`.
14
+ `.[pgvector]`, `.[pinecone]`, `.[models]`, or `.[dashboard]`.
15
15
 
16
16
  ## Checks before opening a pull request
17
17
 
@@ -1,10 +1,11 @@
1
1
  include README.md README_PYPI.md LICENSE CHANGELOG.md CONTRIBUTING.md SECURITY.md CITATION.cff requirements.txt requirements-dev.txt
2
2
  recursive-include docs *
3
- recursive-include examples *.md *.yaml *.json *.jsonl *.py
3
+ recursive-include examples *.md *.yaml *.json *.jsonl *.py *.sql *.sh
4
4
  recursive-include scripts *.sh *.py
5
5
  recursive-include frozen *.md *.sha256
6
6
  prune .github
7
7
  prune tests
8
+ prune examples/*/runtime
8
9
  global-exclude __pycache__
9
10
  global-exclude *.py[cod]
10
11
  global-exclude *.index *.npy *.npz *.sqlite *.sqlite3 *.db
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: embedflow
3
- Version: 0.1.1
3
+ Version: 0.3.0
4
4
  Summary: Progressive embedding-model migration over existing vector indexes.
5
5
  Author: Arnav Srivastav
6
6
  License-Expression: AGPL-3.0-only
@@ -8,7 +8,7 @@ Project-URL: Homepage, https://embedflow.org
8
8
  Project-URL: Repository, https://github.com/arnsri33/embedflow
9
9
  Project-URL: Documentation, https://github.com/arnsri33/embedflow#readme
10
10
  Project-URL: Issues, https://github.com/arnsri33/embedflow/issues
11
- Keywords: embeddings,vector-search,rag,information-retrieval,faiss,qdrant
11
+ Keywords: embeddings,vector-search,rag,information-retrieval,faiss,qdrant,pgvector,pinecone
12
12
  Classifier: Development Status :: 3 - Alpha
13
13
  Classifier: Intended Audience :: Developers
14
14
  Classifier: Intended Audience :: Science/Research
@@ -27,9 +27,13 @@ Provides-Extra: faiss
27
27
  Requires-Dist: faiss-cpu>=1.8.0; extra == "faiss"
28
28
  Provides-Extra: qdrant
29
29
  Requires-Dist: qdrant-client>=1.9; extra == "qdrant"
30
+ Provides-Extra: pgvector
31
+ Requires-Dist: psycopg[binary]>=3.2; extra == "pgvector"
32
+ Provides-Extra: pinecone
33
+ Requires-Dist: pinecone>=6.0; extra == "pinecone"
30
34
  Provides-Extra: models
31
- Requires-Dist: huggingface-hub>=0.20; extra == "models"
32
- Requires-Dist: transformers>=4.45; extra == "models"
35
+ Requires-Dist: huggingface-hub<1.0,>=0.34; extra == "models"
36
+ Requires-Dist: transformers<5.0,>=4.45; extra == "models"
33
37
  Requires-Dist: sentence-transformers>=3.0; extra == "models"
34
38
  Requires-Dist: torch>=2.2; extra == "models"
35
39
  Provides-Extra: dashboard
@@ -39,8 +43,10 @@ Requires-Dist: pydantic>=2.0; extra == "dashboard"
39
43
  Provides-Extra: all
40
44
  Requires-Dist: faiss-cpu>=1.8.0; extra == "all"
41
45
  Requires-Dist: qdrant-client>=1.9; extra == "all"
42
- Requires-Dist: huggingface-hub>=0.20; extra == "all"
43
- Requires-Dist: transformers>=4.45; extra == "all"
46
+ Requires-Dist: psycopg[binary]>=3.2; extra == "all"
47
+ Requires-Dist: pinecone>=6.0; extra == "all"
48
+ Requires-Dist: huggingface-hub<1.0,>=0.34; extra == "all"
49
+ Requires-Dist: transformers<5.0,>=4.45; extra == "all"
44
50
  Requires-Dist: sentence-transformers>=3.0; extra == "all"
45
51
  Requires-Dist: torch>=2.2; extra == "all"
46
52
  Requires-Dist: fastapi>=0.100; extra == "all"
@@ -59,7 +65,7 @@ Dynamic: license-file
59
65
  EmbedFlow lets a new embedding model serve over candidates from an existing
60
66
  vector index while target document vectors are materialized progressively. It
61
67
  supports migration analysis, persistent caching, background work, FAISS,
62
- Qdrant, a CLI, and FastAPI.
68
+ Qdrant, pgvector, Pinecone, a CLI, and FastAPI.
63
69
 
64
70
  The full project README and architecture diagram are on
65
71
  <https://github.com/arnsri33/embedflow>.
@@ -76,6 +82,18 @@ For FAISS and the dashboard:
76
82
  python -m pip install "embedflow[faiss,dashboard]"
77
83
  ```
78
84
 
85
+ For an existing PostgreSQL/pgvector table:
86
+
87
+ ```bash
88
+ python -m pip install "embedflow[pgvector]"
89
+ ```
90
+
91
+ For an existing Pinecone dense index:
92
+
93
+ ```bash
94
+ python -m pip install "embedflow[pinecone]"
95
+ ```
96
+
79
97
  Qdrant and model-runtime extras are documented in the
80
98
  [installation guide](https://github.com/arnsri33/embedflow/blob/main/docs/installation.md).
81
99
  For model-backed analysis, install `embedflow[faiss,models,dashboard]`.
@@ -176,9 +194,13 @@ for definitions and reproduction details.
176
194
  | --- | --- |
177
195
  | FAISS | Supported |
178
196
  | Qdrant | Supported |
197
+ | pgvector | Supported |
198
+ | Pinecone | Supported |
179
199
 
180
- See the [FAISS guide](https://github.com/arnsri33/embedflow/blob/main/docs/integrations/faiss.md)
181
- and [Qdrant guide](https://github.com/arnsri33/embedflow/blob/main/docs/integrations/qdrant.md).
200
+ See the [FAISS guide](https://github.com/arnsri33/embedflow/blob/main/docs/integrations/faiss.md),
201
+ [Qdrant guide](https://github.com/arnsri33/embedflow/blob/main/docs/integrations/qdrant.md),
202
+ and [pgvector guide](https://github.com/arnsri33/embedflow/blob/main/docs/integrations/pgvector.md),
203
+ and [Pinecone guide](https://github.com/arnsri33/embedflow/blob/main/docs/integrations/pinecone.md).
182
204
 
183
205
  ## CLI
184
206
 
@@ -188,7 +210,7 @@ embedflow analyze --help
188
210
  embedflow serve --config ./embedflow.yaml
189
211
  embedflow status --config ./embedflow.yaml
190
212
  embedflow registry list
191
- embedflow economics --corpus-size 1000000000 --docs-per-second 106.98 --gpu-price 3.29
213
+ embedflow economics --corpus-size 1000000000 --docs-per-second 100 --gpu-price 3.29
192
214
  embedflow doctor --config ./embedflow.yaml
193
215
  ```
194
216
 
@@ -198,7 +220,7 @@ cover the remaining commands and endpoints.
198
220
 
199
221
  ## Status
200
222
 
201
- EmbedFlow v0.1.0 is an alpha release for research and early real-world
223
+ EmbedFlow v0.3.0 is an alpha release for research and early real-world
202
224
  testing. T2-v1 is an empirical finite-tail diagnostic, partial rankings can
203
225
  differ from fully warm target reranking, and ANN fidelity needs a reference
204
226
  comparison to audit.
@@ -7,7 +7,7 @@
7
7
  EmbedFlow lets a new embedding model serve over candidates from an existing
8
8
  vector index while target document vectors are materialized progressively. It
9
9
  supports migration analysis, persistent caching, background work, and serving
10
- through FAISS, Qdrant, a CLI, and FastAPI.
10
+ through FAISS, Qdrant, pgvector, Pinecone, a CLI, and FastAPI.
11
11
 
12
12
  [Quickstart](#try-it) · [Documentation](#documentation) · [Research](#research)
13
13
 
@@ -52,6 +52,12 @@ For FAISS and the dashboard, add the optional integrations:
52
52
  python -m pip install "embedflow[faiss,dashboard]"
53
53
  ```
54
54
 
55
+ For an existing Pinecone dense index:
56
+
57
+ ```bash
58
+ python -m pip install "embedflow[pinecone]"
59
+ ```
60
+
55
61
  Qdrant and model-runtime extras are documented in
56
62
  [`docs/installation.md`](https://github.com/arnsri33/embedflow/blob/main/docs/installation.md).
57
63
  For model-backed analysis, install `embedflow[faiss,models,dashboard]`.
@@ -189,11 +195,15 @@ is measured separately and is `UNKNOWN` until an exact reference is supplied.
189
195
  | --- | --- |
190
196
  | FAISS | Supported |
191
197
  | Qdrant | Supported |
198
+ | pgvector | Supported |
199
+ | Pinecone | Supported |
192
200
 
193
201
  Backend-specific setup and examples:
194
202
 
195
203
  - [FAISS](https://github.com/arnsri33/embedflow/blob/main/docs/integrations/faiss.md)
196
204
  - [Qdrant](https://github.com/arnsri33/embedflow/blob/main/docs/integrations/qdrant.md)
205
+ - [pgvector](https://github.com/arnsri33/embedflow/blob/main/docs/integrations/pgvector.md)
206
+ - [Pinecone](https://github.com/arnsri33/embedflow/blob/main/docs/integrations/pinecone.md)
197
207
  - [Adding a backend](https://github.com/arnsri33/embedflow/blob/main/CONTRIBUTING.md)
198
208
 
199
209
  ## CLI
@@ -204,7 +214,7 @@ embedflow analyze --help
204
214
  embedflow serve --config ./embedflow.yaml
205
215
  embedflow status --config ./embedflow.yaml
206
216
  embedflow registry list
207
- embedflow economics --corpus-size 1000000000 --docs-per-second 106.98 --gpu-price 3.29
217
+ embedflow economics --corpus-size 1000000000 --docs-per-second 100 --gpu-price 3.29
208
218
  embedflow doctor --config ./embedflow.yaml
209
219
  ```
210
220
 
@@ -228,7 +238,7 @@ OpenAPI documentation; see
228
238
 
229
239
  ## Status
230
240
 
231
- EmbedFlow v0.1.0 is an alpha release for research and early real-world
241
+ EmbedFlow v0.3.0 is an alpha release for research and early real-world
232
242
  testing.
233
243
 
234
244
  - T2-v1 reports an empirical finite-tail diagnostic.
@@ -5,7 +5,7 @@
5
5
  EmbedFlow lets a new embedding model serve over candidates from an existing
6
6
  vector index while target document vectors are materialized progressively. It
7
7
  supports migration analysis, persistent caching, background work, FAISS,
8
- Qdrant, a CLI, and FastAPI.
8
+ Qdrant, pgvector, Pinecone, a CLI, and FastAPI.
9
9
 
10
10
  The full project README and architecture diagram are on
11
11
  <https://github.com/arnsri33/embedflow>.
@@ -22,6 +22,18 @@ For FAISS and the dashboard:
22
22
  python -m pip install "embedflow[faiss,dashboard]"
23
23
  ```
24
24
 
25
+ For an existing PostgreSQL/pgvector table:
26
+
27
+ ```bash
28
+ python -m pip install "embedflow[pgvector]"
29
+ ```
30
+
31
+ For an existing Pinecone dense index:
32
+
33
+ ```bash
34
+ python -m pip install "embedflow[pinecone]"
35
+ ```
36
+
25
37
  Qdrant and model-runtime extras are documented in the
26
38
  [installation guide](https://github.com/arnsri33/embedflow/blob/main/docs/installation.md).
27
39
  For model-backed analysis, install `embedflow[faiss,models,dashboard]`.
@@ -122,9 +134,13 @@ for definitions and reproduction details.
122
134
  | --- | --- |
123
135
  | FAISS | Supported |
124
136
  | Qdrant | Supported |
137
+ | pgvector | Supported |
138
+ | Pinecone | Supported |
125
139
 
126
- See the [FAISS guide](https://github.com/arnsri33/embedflow/blob/main/docs/integrations/faiss.md)
127
- and [Qdrant guide](https://github.com/arnsri33/embedflow/blob/main/docs/integrations/qdrant.md).
140
+ See the [FAISS guide](https://github.com/arnsri33/embedflow/blob/main/docs/integrations/faiss.md),
141
+ [Qdrant guide](https://github.com/arnsri33/embedflow/blob/main/docs/integrations/qdrant.md),
142
+ and [pgvector guide](https://github.com/arnsri33/embedflow/blob/main/docs/integrations/pgvector.md),
143
+ and [Pinecone guide](https://github.com/arnsri33/embedflow/blob/main/docs/integrations/pinecone.md).
128
144
 
129
145
  ## CLI
130
146
 
@@ -134,7 +150,7 @@ embedflow analyze --help
134
150
  embedflow serve --config ./embedflow.yaml
135
151
  embedflow status --config ./embedflow.yaml
136
152
  embedflow registry list
137
- embedflow economics --corpus-size 1000000000 --docs-per-second 106.98 --gpu-price 3.29
153
+ embedflow economics --corpus-size 1000000000 --docs-per-second 100 --gpu-price 3.29
138
154
  embedflow doctor --config ./embedflow.yaml
139
155
  ```
140
156
 
@@ -144,7 +160,7 @@ cover the remaining commands and endpoints.
144
160
 
145
161
  ## Status
146
162
 
147
- EmbedFlow v0.1.0 is an alpha release for research and early real-world
163
+ EmbedFlow v0.3.0 is an alpha release for research and early real-world
148
164
  testing. T2-v1 is an empirical finite-tail diagnostic, partial rankings can
149
165
  differ from fully warm target reranking, and ANN fidelity needs a reference
150
166
  comparison to audit.
@@ -13,8 +13,8 @@ Interactive OpenAPI documentation is available at
13
13
 
14
14
  | Method | Path | Purpose |
15
15
  | --- | --- | --- |
16
- | GET | `/health` | Process and backend health |
17
- | GET | `/status` | Cache, queue, model, and migration state |
16
+ | GET | `/health` | Lightweight process liveness |
17
+ | GET | `/status` | Cache, queue, model, migration, and backend metadata |
18
18
  | POST | `/search` | Source retrieval and target reranking |
19
19
  | POST | `/analyze` | Run or retrieve migration analysis |
20
20
  | POST | `/prewarm` | Queue document materialization |
@@ -48,6 +48,12 @@ The response includes the result list and migration fields such as:
48
48
  request. A partial response scores the available target vectors; it can differ
49
49
  from the fully warm ranking.
50
50
 
51
+ `/status` includes safe backend metadata. For pgvector this names the
52
+ schema/table and vector contract; for Pinecone it names the host/index,
53
+ namespace, dimension, metric, and safe vector counts. Neither backend returns
54
+ the DSN, API key, or credentials;
55
+ `/health` remains a compact liveness response for probes and load balancers.
56
+
51
57
  ## Errors
52
58
 
53
59
  Validation errors use FastAPI/Pydantic's normal JSON response. Backend,
@@ -52,7 +52,7 @@ Registry matching uses model contracts and corpus identity. See
52
52
  ```bash
53
53
  embedflow economics \
54
54
  --corpus-size 1000000000 \
55
- --docs-per-second 106.98 \
55
+ --docs-per-second 100 \
56
56
  --gpu-price 3.29
57
57
  embedflow doctor --config ./embedflow.yaml
58
58
  embedflow demo
@@ -14,12 +14,16 @@ target:
14
14
  device: cuda
15
15
 
16
16
  index:
17
- backend: faiss # faiss or qdrant
17
+ backend: faiss # faiss, qdrant, pgvector, or pinecone
18
18
  path: ./legacy.index
19
19
  ids: ./legacy.index.ids.json # FAISS sidecar
20
20
  metric: cosine
21
21
  nprobe: 64
22
22
  # Qdrant fields: url, collection, vector_name, api_key_env
23
+ # pgvector fields: dsn_env, schema, table, id_column, vector_column, text_column
24
+ # hnsw_ef_search, ivfflat_probes
25
+ # Pinecone fields: host (preferred) or index_name, api_key_env, namespace,
26
+ # text_metadata_field
23
27
 
24
28
  documents:
25
29
  path: ./documents.jsonl
@@ -78,6 +82,19 @@ EMBEDFLOW_INDEX_URL
78
82
  EMBEDFLOW_INDEX_COLLECTION
79
83
  EMBEDFLOW_INDEX_VECTOR_NAME
80
84
  EMBEDFLOW_QDRANT_API_KEY_ENV
85
+ EMBEDFLOW_PGVECTOR_DSN_ENV
86
+ EMBEDFLOW_PGVECTOR_SCHEMA
87
+ EMBEDFLOW_PGVECTOR_TABLE
88
+ EMBEDFLOW_PGVECTOR_ID_COLUMN
89
+ EMBEDFLOW_PGVECTOR_VECTOR_COLUMN
90
+ EMBEDFLOW_PGVECTOR_TEXT_COLUMN
91
+ EMBEDFLOW_PGVECTOR_HNSW_EF_SEARCH
92
+ EMBEDFLOW_PGVECTOR_IVFFLAT_PROBES
93
+ EMBEDFLOW_PINECONE_HOST
94
+ EMBEDFLOW_PINECONE_INDEX_NAME
95
+ EMBEDFLOW_PINECONE_NAMESPACE
96
+ EMBEDFLOW_PINECONE_TEXT_METADATA_FIELD
97
+ EMBEDFLOW_PINECONE_API_KEY_ENV
81
98
  EMBEDFLOW_INDEX_NPROBE
82
99
  EMBEDFLOW_DOCUMENTS_PATH
83
100
  EMBEDFLOW_CACHE_PATH
@@ -6,7 +6,7 @@ throughput and GPU price supplied by the user:
6
6
  ```bash
7
7
  embedflow economics \
8
8
  --corpus-size 1000000000 \
9
- --docs-per-second 106.98 \
9
+ --docs-per-second 100 \
10
10
  --gpu-price 3.29
11
11
  ```
12
12
 
@@ -30,6 +30,8 @@ python -m pip install -e .
30
30
  | --- | --- |
31
31
  | `faiss` | FAISS source-index adapter |
32
32
  | `qdrant` | Qdrant client and adapter |
33
+ | `pgvector` | Psycopg 3 binary driver and pgvector adapter |
34
+ | `pinecone` | Official Pinecone Python SDK and adapter |
33
35
  | `models` | PyTorch, Transformers, Sentence Transformers, and Hub client |
34
36
  | `dashboard` | FastAPI, Uvicorn, and Pydantic |
35
37
  | `dev` | Pytest, Ruff, and build tooling |
@@ -52,6 +54,18 @@ with `index.url`; a user-owned cloud or remote server can use an API key named
52
54
  by `index.api_key_env`. Keep the key in the environment. See
53
55
  [`integrations/qdrant.md`](integrations/qdrant.md).
54
56
 
57
+ ## PostgreSQL / pgvector
58
+
59
+ Install the optional adapter with `python -m pip install "embedflow[pgvector]"`.
60
+ The adapter connects to an existing table and reads the DSN from the
61
+ environment; see [`integrations/pgvector.md`](integrations/pgvector.md).
62
+
63
+ ## Pinecone
64
+
65
+ Install the optional adapter with `python -m pip install "embedflow[pinecone]"`.
66
+ Set `PINECONE_API_KEY` in the environment and configure an existing dense
67
+ index host; see [`integrations/pinecone.md`](integrations/pinecone.md).
68
+
55
69
  ## CPU and GPU
56
70
 
57
71
  The deterministic demo runs on CPU. Real model serving accepts `--device cpu`
@@ -0,0 +1,140 @@
1
+ # PostgreSQL / pgvector
2
+
3
+ EmbedFlow can read candidates from an existing PostgreSQL table that uses the
4
+ [pgvector](https://github.com/pgvector/pgvector) extension. The adapter uses
5
+ the shared `VectorIndex` interface, so analysis, serving, cache warming, and
6
+ the API follow the same path as FAISS and Qdrant.
7
+
8
+ ## Install
9
+
10
+ ```bash
11
+ python -m pip install "embedflow[pgvector]"
12
+ ```
13
+
14
+ The extra installs Psycopg 3 with its binary distribution. The base
15
+ `embedflow` install does not import or require PostgreSQL dependencies.
16
+
17
+ ## Connection and table layout
18
+
19
+ Keep the DSN in an environment variable:
20
+
21
+ ```bash
22
+ export EMBEDFLOW_PGVECTOR_DSN='postgresql://user:password@localhost:5432/app'
23
+ ```
24
+
25
+ Example configuration:
26
+
27
+ ```yaml
28
+ source:
29
+ model: sentence-transformers/all-MiniLM-L6-v2
30
+ device: cpu
31
+
32
+ target:
33
+ model: Qwen/Qwen3-Embedding-0.6B
34
+ device: cpu
35
+
36
+ index:
37
+ backend: pgvector
38
+ dsn_env: EMBEDFLOW_PGVECTOR_DSN
39
+ schema: public
40
+ table: documents
41
+ id_column: id
42
+ vector_column: embedding
43
+ text_column: content
44
+ metric: cosine
45
+ # hnsw_ef_search: 100
46
+ # ivfflat_probes: 20
47
+
48
+ documents:
49
+ # Omit this file when text_column is present in the table. A separate
50
+ # JSONL document store can be supplied when the table stores IDs only.
51
+ path: ./documents.jsonl
52
+ id_field: id
53
+ text_field: text
54
+ ```
55
+
56
+ When `documents.path` does not exist, EmbedFlow resolves candidate text from
57
+ `text_column` in the configured table. If a JSONL file exists, it remains the
58
+ document store and the pgvector table is used for vectors and IDs. The adapter
59
+ only reads the table during ordinary `analyze`, `serve`, `search`, `status`,
60
+ and `audit-index` operations; it does not create extensions, indexes, tables,
61
+ or modify rows.
62
+
63
+ The ID column may be an integer, bigint, UUID, or text type. IDs are normalized
64
+ to EmbedFlow's canonical string representation without numeric coercion.
65
+ The configured vector dimension is checked against `vector(n)` where available
66
+ and against a bounded non-null row for unbounded `vector` columns.
67
+
68
+ ## Metrics and score direction
69
+
70
+ Supported metrics are `cosine`, `l2`/`euclidean`, and
71
+ `inner_product`/`dot`. pgvector's distance operators are converted to the
72
+ EmbedFlow convention where larger scores rank first:
73
+
74
+ | EmbedFlow metric | pgvector operator | Internal score |
75
+ | --- | --- | --- |
76
+ | cosine | `<=>` | `1 - distance` |
77
+ | l2 / euclidean | `<->` | `-distance` |
78
+ | inner_product / dot | `<#>` | `-distance` |
79
+
80
+ The last conversion accounts for pgvector's negative inner-product operator.
81
+
82
+ ## ANN settings
83
+
84
+ EmbedFlow never creates or changes an ANN index. If configured, `hnsw_ef_search`
85
+ and `ivfflat_probes` are applied with transaction-local `set_config` calls for
86
+ the retrieval query. Omit them to use the database/index defaults. The
87
+ `metadata()` and `audit-index` output reports configured settings and detected
88
+ valid HNSW/IVFFlat indexes when PostgreSQL exposes them.
89
+
90
+ ## Commands
91
+
92
+ ```bash
93
+ embedflow analyze --config embedflow.yaml
94
+ embedflow serve --config embedflow.yaml
95
+ embedflow search --config embedflow.yaml "what causes auroras?"
96
+ embedflow status --config embedflow.yaml
97
+ embedflow audit-index --config embedflow.yaml
98
+ ```
99
+
100
+ `audit-index` checks table and column presence, pgvector type and dimension,
101
+ NULL vectors, duplicate canonical IDs, a retrieval probe, and available ANN
102
+ metadata. A successful connection is not an ANN recall audit; use an exact
103
+ reference comparison when fidelity matters.
104
+
105
+ ## Local example
106
+
107
+ The repository includes a small Docker example that creates a disposable
108
+ pgvector table and index:
109
+
110
+ ```bash
111
+ cd examples/pgvector
112
+ docker compose up -d
113
+ export EMBEDFLOW_PGVECTOR_DSN='postgresql://embedflow:embedflow@localhost:5432/embedflow'
114
+ ./run_demo.sh
115
+ ```
116
+
117
+ The example is for local testing only. It contains no credentials for a real
118
+ service and does not represent a production database configuration.
119
+
120
+ ## Troubleshooting
121
+
122
+ - `pgvector support requires psycopg`: install `embedflow[pgvector]` in the
123
+ active environment.
124
+ - `DSN not found`: export the variable named by `index.dsn_env`.
125
+ - `table ... was not found`: check the schema/table spelling and database.
126
+ - `is not a pgvector column`: install/enable the extension in the database and
127
+ point `vector_column` at a `vector` column.
128
+ - Dimension mismatch: set `source.dimension` to the encoder's dimension and
129
+ verify the table's vector type.
130
+ - Authentication errors are redacted in EmbedFlow output; inspect PostgreSQL
131
+ server logs without pasting passwords into issue reports.
132
+
133
+ ## Limitations
134
+
135
+ The adapter currently exposes the common table layout and no arbitrary SQL
136
+ filters. Metadata filtering will follow the shared backend interface when that
137
+ interface gains a portable filter contract. Connection operations on one
138
+ adapter are serialized; applications needing a larger pool can create several
139
+ sessions behind their own pool. ANN health is reported as `UNKNOWN` unless a
140
+ reference comparison is run.