embedflow 0.4.0__tar.gz → 0.6.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {embedflow-0.4.0 → embedflow-0.6.0}/CHANGELOG.md +19 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/CITATION.cff +2 -2
- {embedflow-0.4.0 → embedflow-0.6.0}/PKG-INFO +31 -4
- {embedflow-0.4.0 → embedflow-0.6.0}/README.md +25 -2
- {embedflow-0.4.0 → embedflow-0.6.0}/README_PYPI.md +26 -2
- {embedflow-0.4.0 → embedflow-0.6.0}/docs/api.md +5 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/docs/cli.md +10 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/docs/configuration.md +50 -1
- {embedflow-0.4.0 → embedflow-0.6.0}/docs/installation.md +7 -0
- embedflow-0.6.0/docs/integrations/weaviate.md +114 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/docs/limitations.md +1 -1
- embedflow-0.6.0/docs/planner.md +111 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/docs/releasing.md +7 -7
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/__init__.py +10 -4
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/cli.py +232 -50
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/compatibility/evaluate.py +3 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/config.py +199 -3
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/indexes/__init__.py +2 -0
- embedflow-0.6.0/embedflow/indexes/weaviate_backend.py +779 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/migration/facade.py +37 -10
- embedflow-0.6.0/embedflow/planner/__init__.py +22 -0
- embedflow-0.6.0/embedflow/planner/economics.py +153 -0
- embedflow-0.6.0/embedflow/planner/models.py +97 -0
- embedflow-0.6.0/embedflow/planner/planner.py +1390 -0
- embedflow-0.6.0/embedflow/planner/rendering.py +93 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/runtime.py +16 -3
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow.egg-info/PKG-INFO +31 -4
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow.egg-info/SOURCES.txt +17 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow.egg-info/requires.txt +4 -0
- embedflow-0.6.0/examples/planner/README.md +12 -0
- embedflow-0.6.0/examples/planner/run_demo.sh +14 -0
- embedflow-0.6.0/examples/weaviate/README.md +32 -0
- embedflow-0.6.0/examples/weaviate/compose.yaml +19 -0
- embedflow-0.6.0/examples/weaviate/embedflow.yaml.example +22 -0
- embedflow-0.6.0/examples/weaviate/run_demo.sh +7 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/pyproject.toml +4 -2
- {embedflow-0.4.0 → embedflow-0.6.0}/scripts/release_gate.py +20 -2
- embedflow-0.6.0/scripts/validate_weaviate.py +301 -0
- embedflow-0.6.0/scripts/weaviate_fixture.py +120 -0
- embedflow-0.6.0/scripts/weaviate_smoke.py +50 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/CONTRIBUTING.md +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/LICENSE +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/MANIFEST.in +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/SECURITY.md +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/docs/assets/README.md +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/docs/assets/candidate-gap-example.svg +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/docs/assets/dashboard-screenshot.md +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/docs/assets/terminal-demo.txt +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/docs/concepts.md +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/docs/contributing-benchmarks.md +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/docs/economics.md +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/docs/integrations/faiss.md +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/docs/integrations/milvus.md +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/docs/integrations/pgvector.md +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/docs/integrations/pinecone.md +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/docs/integrations/qdrant.md +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/docs/methodology.md +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/docs/quickstart.md +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/docs/registry.md +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/__main__.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/analysis.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/api.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/cache/__init__.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/cache/base.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/cache/persistent_cache.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/compatibility/__init__.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/compatibility/candidate_gap.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/compatibility/containment.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/compatibility/metrics.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/compatibility/migration_depth.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/compatibility/probe.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/compatibility/report.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/compatibility/t2.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/data/__init__.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/data/registry/__init__.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/data/registry/benchmark_profiles.jsonl +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/data/registry/checksums.sha256 +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/data/registry/migrations.jsonl +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/data/registry/registry_manifest.json +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/data/registry/research_summaries.json +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/data/registry/schema_version.json +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/frozen/T2_V1_FROZEN_SPEC.md +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/frozen/T2_V1_FROZEN_SPEC.sha256 +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/indexes/base.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/indexes/faiss_backend.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/indexes/milvus_backend.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/indexes/pgvector_backend.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/indexes/pinecone_backend.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/indexes/qdrant_backend.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/metrics/__init__.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/metrics/latency.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/migration/__init__.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/migration/compatibility.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/migration/materializer.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/migration/planner.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/migration/state.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/models/__init__.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/models/base.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/models/huggingface.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/registry/__init__.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/registry/loader.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/registry/matcher.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/registry/schema.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/serving/__init__.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/serving/api.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/serving/engine.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/serving/factory.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow/serving/schemas.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow.egg-info/dependency_links.txt +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow.egg-info/entry_points.txt +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/embedflow.egg-info/top_level.txt +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/examples/faiss/README.md +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/examples/faiss/documents.jsonl +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/examples/faiss/embedflow.yaml +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/examples/faiss/queries.jsonl +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/examples/milvus/README.md +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/examples/milvus/compose.yaml +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/examples/milvus/embedflow.yaml.example +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/examples/milvus/run_demo.sh +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/examples/pgvector/README.md +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/examples/pgvector/build_index.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/examples/pgvector/compose.yaml +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/examples/pgvector/embedflow.yaml +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/examples/pgvector/init.sql +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/examples/pgvector/queries.jsonl +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/examples/pgvector/run_demo.sh +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/examples/pinecone/README.md +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/examples/pinecone/embedflow.yaml.example +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/examples/pinecone/run_smoke.sh +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/examples/qdrant/README.md +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/examples/qdrant/build_index.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/examples/qdrant/documents.jsonl +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/examples/qdrant/embedflow.yaml +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/examples/qdrant/queries.jsonl +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/examples/research_analysis/README.md +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/examples/research_analysis/documents.jsonl +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/examples/research_analysis/embedflow.yaml +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/examples/research_analysis/qrels.json +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/examples/research_analysis/queries.jsonl +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/frozen/T2_V1_FROZEN_SPEC.md +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/frozen/T2_V1_FROZEN_SPEC.sha256 +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/requirements-dev.txt +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/requirements.txt +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/scripts/milvus_fixture.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/scripts/pinecone_smoke.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/scripts/real_qdrant_smoke.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/scripts/run_demo.sh +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/scripts/run_tests.sh +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/scripts/validate_milvus.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/scripts/validate_pgvector_10k.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/setup.cfg +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/src/__init__.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/src/embed.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/src/probe_features.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/src/storage.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/src/t2_v1.py +0 -0
- {embedflow-0.4.0 → embedflow-0.6.0}/src/utils.py +0 -0
|
@@ -1,5 +1,24 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## v0.6.0 — Migration planner
|
|
4
|
+
|
|
5
|
+
- Added an advisory `embedflow plan` command and Python API that combine
|
|
6
|
+
source-index preflight, registry evidence, representative probes, frozen
|
|
7
|
+
T2-v1 diagnostics, candidate-depth selection, cache planning, economics, and
|
|
8
|
+
staged rollout guidance without routing traffic or mutating the source.
|
|
9
|
+
- Added structured JSON/YAML plan artifacts with explicit warnings and
|
|
10
|
+
measured/user-supplied/modeled/unknown provenance.
|
|
11
|
+
|
|
12
|
+
## v0.5.0 — Weaviate backend
|
|
13
|
+
|
|
14
|
+
- Added a read-only Weaviate v4 backend for existing externally-vectorized
|
|
15
|
+
dense collections, including named-vector selection and UUID IDs.
|
|
16
|
+
- Added HTTP/gRPC connection settings, property-backed or external text
|
|
17
|
+
resolution, metric normalization, collection auditing, and safe client
|
|
18
|
+
shutdown.
|
|
19
|
+
- Added the optional `weaviate-client` extra, Docker fixture, example, and
|
|
20
|
+
live 10,000-object validation harness.
|
|
21
|
+
|
|
3
22
|
## v0.4.0 — Milvus backend
|
|
4
23
|
|
|
5
24
|
- Added a read-only Milvus backend for existing dense `FLOAT_VECTOR`
|
|
@@ -2,8 +2,8 @@ cff-version: 1.2.0
|
|
|
2
2
|
title: "EmbedFlow: Upgrading Legacy Embeddings Without Full Upfront Re-Embedding"
|
|
3
3
|
message: "If EmbedFlow contributes to your work, please cite this software release."
|
|
4
4
|
type: software
|
|
5
|
-
version: 0.
|
|
6
|
-
date-released: 2026-09-
|
|
5
|
+
version: 0.6.0
|
|
6
|
+
date-released: 2026-09-13
|
|
7
7
|
repository-code: "https://github.com/arnsri33/embedflow"
|
|
8
8
|
url: "https://github.com/arnsri33/embedflow"
|
|
9
9
|
license: AGPL-3.0-only
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.4
|
|
2
2
|
Name: embedflow
|
|
3
|
-
Version: 0.
|
|
3
|
+
Version: 0.6.0
|
|
4
4
|
Summary: Progressive embedding-model migration over existing vector indexes.
|
|
5
5
|
Author: Arnav Srivastav
|
|
6
6
|
License-Expression: AGPL-3.0-only
|
|
@@ -8,7 +8,7 @@ Project-URL: Homepage, https://embedflow.org
|
|
|
8
8
|
Project-URL: Repository, https://github.com/arnsri33/embedflow
|
|
9
9
|
Project-URL: Documentation, https://github.com/arnsri33/embedflow#readme
|
|
10
10
|
Project-URL: Issues, https://github.com/arnsri33/embedflow/issues
|
|
11
|
-
Keywords: embeddings,vector-search,rag,information-retrieval,faiss,qdrant,pgvector,pinecone,milvus
|
|
11
|
+
Keywords: embeddings,vector-search,rag,information-retrieval,faiss,qdrant,pgvector,pinecone,milvus,weaviate
|
|
12
12
|
Classifier: Development Status :: 3 - Alpha
|
|
13
13
|
Classifier: Intended Audience :: Developers
|
|
14
14
|
Classifier: Intended Audience :: Science/Research
|
|
@@ -33,6 +33,8 @@ Provides-Extra: pinecone
|
|
|
33
33
|
Requires-Dist: pinecone>=6.0; extra == "pinecone"
|
|
34
34
|
Provides-Extra: milvus
|
|
35
35
|
Requires-Dist: pymilvus<4,>=2.5.5; extra == "milvus"
|
|
36
|
+
Provides-Extra: weaviate
|
|
37
|
+
Requires-Dist: weaviate-client<5,>=4.16; extra == "weaviate"
|
|
36
38
|
Provides-Extra: models
|
|
37
39
|
Requires-Dist: huggingface-hub<1.0,>=0.34; extra == "models"
|
|
38
40
|
Requires-Dist: transformers<5.0,>=4.45; extra == "models"
|
|
@@ -48,6 +50,7 @@ Requires-Dist: qdrant-client>=1.9; extra == "all"
|
|
|
48
50
|
Requires-Dist: psycopg[binary]>=3.2; extra == "all"
|
|
49
51
|
Requires-Dist: pinecone>=6.0; extra == "all"
|
|
50
52
|
Requires-Dist: pymilvus<4,>=2.5.5; extra == "all"
|
|
53
|
+
Requires-Dist: weaviate-client<5,>=4.16; extra == "all"
|
|
51
54
|
Requires-Dist: huggingface-hub<1.0,>=0.34; extra == "all"
|
|
52
55
|
Requires-Dist: transformers<5.0,>=4.45; extra == "all"
|
|
53
56
|
Requires-Dist: sentence-transformers>=3.0; extra == "all"
|
|
@@ -68,7 +71,7 @@ Dynamic: license-file
|
|
|
68
71
|
EmbedFlow lets a new embedding model serve over candidates from an existing
|
|
69
72
|
vector index while target document vectors are materialized progressively. It
|
|
70
73
|
supports migration analysis, persistent caching, background work, FAISS,
|
|
71
|
-
Qdrant, pgvector, Pinecone, Milvus, a CLI, and FastAPI.
|
|
74
|
+
Qdrant, pgvector, Pinecone, Milvus, Weaviate, a CLI, and FastAPI.
|
|
72
75
|
|
|
73
76
|
The full project README and architecture diagram are on
|
|
74
77
|
<https://github.com/arnsri33/embedflow>.
|
|
@@ -103,6 +106,12 @@ For an existing Milvus collection:
|
|
|
103
106
|
python -m pip install "embedflow[milvus]"
|
|
104
107
|
```
|
|
105
108
|
|
|
109
|
+
For an existing Weaviate v4 collection:
|
|
110
|
+
|
|
111
|
+
```bash
|
|
112
|
+
python -m pip install "embedflow[weaviate]"
|
|
113
|
+
```
|
|
114
|
+
|
|
106
115
|
Qdrant and model-runtime extras are documented in the
|
|
107
116
|
[installation guide](https://github.com/arnsri33/embedflow/blob/main/docs/installation.md).
|
|
108
117
|
For model-backed analysis, install `embedflow[faiss,models,dashboard]`.
|
|
@@ -180,6 +189,21 @@ Selected core records (nDCG@10, `G(50)`):
|
|
|
180
189
|
See the [registry documentation](https://github.com/arnsri33/embedflow/blob/main/docs/registry.md)
|
|
181
190
|
for matching levels, contract fingerprints, and provenance.
|
|
182
191
|
|
|
192
|
+
## Advisory migration plans
|
|
193
|
+
|
|
194
|
+
Use representative query probes to produce a bounded migration recommendation
|
|
195
|
+
before serving:
|
|
196
|
+
|
|
197
|
+
```bash
|
|
198
|
+
embedflow plan --config ./embedflow.yaml --queries ./probe_queries.jsonl
|
|
199
|
+
embedflow plan --config ./embedflow.yaml --queries ./probe_queries.jsonl --format json --output migration-plan.json
|
|
200
|
+
```
|
|
201
|
+
|
|
202
|
+
The planner reports preflight, registry evidence, frozen T2-v1 finite-tail
|
|
203
|
+
behavior, candidate K, cache/economics projections, and staged rollout
|
|
204
|
+
guidance. It never routes traffic or mutates the source index; `SAFE` is not a
|
|
205
|
+
qrels-based retrieval-quality guarantee. See the [planner guide](https://github.com/arnsri33/embedflow/blob/main/docs/planner.md).
|
|
206
|
+
|
|
183
207
|
## Research
|
|
184
208
|
|
|
185
209
|
For candidate depth `K`, EmbedFlow measures:
|
|
@@ -206,18 +230,21 @@ for definitions and reproduction details.
|
|
|
206
230
|
| pgvector | Supported |
|
|
207
231
|
| Pinecone | Supported |
|
|
208
232
|
| Milvus | Supported |
|
|
233
|
+
| Weaviate | Supported |
|
|
209
234
|
|
|
210
235
|
See the [FAISS guide](https://github.com/arnsri33/embedflow/blob/main/docs/integrations/faiss.md),
|
|
211
236
|
[Qdrant guide](https://github.com/arnsri33/embedflow/blob/main/docs/integrations/qdrant.md),
|
|
212
237
|
and [pgvector guide](https://github.com/arnsri33/embedflow/blob/main/docs/integrations/pgvector.md),
|
|
213
238
|
and [Pinecone guide](https://github.com/arnsri33/embedflow/blob/main/docs/integrations/pinecone.md),
|
|
214
239
|
and [Milvus guide](https://github.com/arnsri33/embedflow/blob/main/docs/integrations/milvus.md).
|
|
240
|
+
See the [Weaviate guide](https://github.com/arnsri33/embedflow/blob/main/docs/integrations/weaviate.md).
|
|
215
241
|
|
|
216
242
|
## CLI
|
|
217
243
|
|
|
218
244
|
```bash
|
|
219
245
|
embedflow --help
|
|
220
246
|
embedflow analyze --help
|
|
247
|
+
embedflow plan --help
|
|
221
248
|
embedflow serve --config ./embedflow.yaml
|
|
222
249
|
embedflow status --config ./embedflow.yaml
|
|
223
250
|
embedflow registry list
|
|
@@ -231,7 +258,7 @@ cover the remaining commands and endpoints.
|
|
|
231
258
|
|
|
232
259
|
## Status
|
|
233
260
|
|
|
234
|
-
EmbedFlow v0.
|
|
261
|
+
EmbedFlow v0.6.0 is a pre-1.0 release for research and early real-world
|
|
235
262
|
testing. T2-v1 is an empirical finite-tail diagnostic, partial rankings can
|
|
236
263
|
differ from fully warm target reranking, and ANN fidelity needs a reference
|
|
237
264
|
comparison to audit.
|
|
@@ -7,7 +7,7 @@
|
|
|
7
7
|
EmbedFlow lets a new embedding model serve over candidates from an existing
|
|
8
8
|
vector index while target document vectors are materialized progressively. It
|
|
9
9
|
supports migration analysis, persistent caching, background work, and serving
|
|
10
|
-
through FAISS, Qdrant, pgvector, Pinecone, Milvus, a CLI, and FastAPI.
|
|
10
|
+
through FAISS, Qdrant, pgvector, Pinecone, Milvus, Weaviate, a CLI, and FastAPI.
|
|
11
11
|
|
|
12
12
|
[Quickstart](#try-it) · [Documentation](#documentation) · [Research](#research)
|
|
13
13
|
|
|
@@ -64,6 +64,12 @@ For an existing Milvus collection:
|
|
|
64
64
|
python -m pip install "embedflow[milvus]"
|
|
65
65
|
```
|
|
66
66
|
|
|
67
|
+
For an existing Weaviate v4 collection (HTTP and gRPC endpoints required):
|
|
68
|
+
|
|
69
|
+
```bash
|
|
70
|
+
python -m pip install "embedflow[weaviate]"
|
|
71
|
+
```
|
|
72
|
+
|
|
67
73
|
Qdrant and model-runtime extras are documented in
|
|
68
74
|
[`docs/installation.md`](https://github.com/arnsri33/embedflow/blob/main/docs/installation.md).
|
|
69
75
|
For model-backed analysis, install `embedflow[faiss,models,dashboard]`.
|
|
@@ -136,6 +142,20 @@ session = embedflow.migrate(
|
|
|
136
142
|
results = session.search("what causes auroras?", top_k=10)
|
|
137
143
|
```
|
|
138
144
|
|
|
145
|
+
## Plan a migration
|
|
146
|
+
|
|
147
|
+
Build a conservative, evidence-aware recommendation before serving. The
|
|
148
|
+
planner reuses backend preflight, registry matching, and frozen T2-v1; it is
|
|
149
|
+
advisory and never routes traffic or mutates the source index.
|
|
150
|
+
|
|
151
|
+
```bash
|
|
152
|
+
embedflow plan --config ./embedflow.yaml --queries ./probe_queries.jsonl
|
|
153
|
+
embedflow plan --config ./embedflow.yaml --queries ./probe_queries.jsonl --format json --output migration-plan.json
|
|
154
|
+
```
|
|
155
|
+
|
|
156
|
+
`SAFE` is an empirical finite-tail signal, not a retrieval-quality guarantee.
|
|
157
|
+
See [`docs/planner.md`](https://github.com/arnsri33/embedflow/blob/main/docs/planner.md).
|
|
158
|
+
|
|
139
159
|
Search responses expose `COLD`, `PARTIAL`, or `WARM`, cache hits and misses,
|
|
140
160
|
synchronous work, queued work, and stage timings. Once the candidate vectors
|
|
141
161
|
are warm, target scoring over that candidate set is deterministic.
|
|
@@ -204,6 +224,7 @@ is measured separately and is `UNKNOWN` until an exact reference is supplied.
|
|
|
204
224
|
| pgvector | Supported |
|
|
205
225
|
| Pinecone | Supported |
|
|
206
226
|
| Milvus | Supported |
|
|
227
|
+
| Weaviate | Supported |
|
|
207
228
|
|
|
208
229
|
Backend-specific setup and examples:
|
|
209
230
|
|
|
@@ -212,6 +233,7 @@ Backend-specific setup and examples:
|
|
|
212
233
|
- [pgvector](https://github.com/arnsri33/embedflow/blob/main/docs/integrations/pgvector.md)
|
|
213
234
|
- [Pinecone](https://github.com/arnsri33/embedflow/blob/main/docs/integrations/pinecone.md)
|
|
214
235
|
- [Milvus](https://github.com/arnsri33/embedflow/blob/main/docs/integrations/milvus.md)
|
|
236
|
+
- [Weaviate](https://github.com/arnsri33/embedflow/blob/main/docs/integrations/weaviate.md)
|
|
215
237
|
- [Adding a backend](https://github.com/arnsri33/embedflow/blob/main/CONTRIBUTING.md)
|
|
216
238
|
|
|
217
239
|
## CLI
|
|
@@ -219,6 +241,7 @@ Backend-specific setup and examples:
|
|
|
219
241
|
```bash
|
|
220
242
|
embedflow --help
|
|
221
243
|
embedflow analyze --help
|
|
244
|
+
embedflow plan --help
|
|
222
245
|
embedflow serve --config ./embedflow.yaml
|
|
223
246
|
embedflow status --config ./embedflow.yaml
|
|
224
247
|
embedflow registry list
|
|
@@ -246,7 +269,7 @@ OpenAPI documentation; see
|
|
|
246
269
|
|
|
247
270
|
## Status
|
|
248
271
|
|
|
249
|
-
EmbedFlow v0.
|
|
272
|
+
EmbedFlow v0.6.0 is a pre-1.0 release for research and early real-world
|
|
250
273
|
testing.
|
|
251
274
|
|
|
252
275
|
- T2-v1 reports an empirical finite-tail diagnostic.
|
|
@@ -5,7 +5,7 @@
|
|
|
5
5
|
EmbedFlow lets a new embedding model serve over candidates from an existing
|
|
6
6
|
vector index while target document vectors are materialized progressively. It
|
|
7
7
|
supports migration analysis, persistent caching, background work, FAISS,
|
|
8
|
-
Qdrant, pgvector, Pinecone, Milvus, a CLI, and FastAPI.
|
|
8
|
+
Qdrant, pgvector, Pinecone, Milvus, Weaviate, a CLI, and FastAPI.
|
|
9
9
|
|
|
10
10
|
The full project README and architecture diagram are on
|
|
11
11
|
<https://github.com/arnsri33/embedflow>.
|
|
@@ -40,6 +40,12 @@ For an existing Milvus collection:
|
|
|
40
40
|
python -m pip install "embedflow[milvus]"
|
|
41
41
|
```
|
|
42
42
|
|
|
43
|
+
For an existing Weaviate v4 collection:
|
|
44
|
+
|
|
45
|
+
```bash
|
|
46
|
+
python -m pip install "embedflow[weaviate]"
|
|
47
|
+
```
|
|
48
|
+
|
|
43
49
|
Qdrant and model-runtime extras are documented in the
|
|
44
50
|
[installation guide](https://github.com/arnsri33/embedflow/blob/main/docs/installation.md).
|
|
45
51
|
For model-backed analysis, install `embedflow[faiss,models,dashboard]`.
|
|
@@ -117,6 +123,21 @@ Selected core records (nDCG@10, `G(50)`):
|
|
|
117
123
|
See the [registry documentation](https://github.com/arnsri33/embedflow/blob/main/docs/registry.md)
|
|
118
124
|
for matching levels, contract fingerprints, and provenance.
|
|
119
125
|
|
|
126
|
+
## Advisory migration plans
|
|
127
|
+
|
|
128
|
+
Use representative query probes to produce a bounded migration recommendation
|
|
129
|
+
before serving:
|
|
130
|
+
|
|
131
|
+
```bash
|
|
132
|
+
embedflow plan --config ./embedflow.yaml --queries ./probe_queries.jsonl
|
|
133
|
+
embedflow plan --config ./embedflow.yaml --queries ./probe_queries.jsonl --format json --output migration-plan.json
|
|
134
|
+
```
|
|
135
|
+
|
|
136
|
+
The planner reports preflight, registry evidence, frozen T2-v1 finite-tail
|
|
137
|
+
behavior, candidate K, cache/economics projections, and staged rollout
|
|
138
|
+
guidance. It never routes traffic or mutates the source index; `SAFE` is not a
|
|
139
|
+
qrels-based retrieval-quality guarantee. See the [planner guide](https://github.com/arnsri33/embedflow/blob/main/docs/planner.md).
|
|
140
|
+
|
|
120
141
|
## Research
|
|
121
142
|
|
|
122
143
|
For candidate depth `K`, EmbedFlow measures:
|
|
@@ -143,18 +164,21 @@ for definitions and reproduction details.
|
|
|
143
164
|
| pgvector | Supported |
|
|
144
165
|
| Pinecone | Supported |
|
|
145
166
|
| Milvus | Supported |
|
|
167
|
+
| Weaviate | Supported |
|
|
146
168
|
|
|
147
169
|
See the [FAISS guide](https://github.com/arnsri33/embedflow/blob/main/docs/integrations/faiss.md),
|
|
148
170
|
[Qdrant guide](https://github.com/arnsri33/embedflow/blob/main/docs/integrations/qdrant.md),
|
|
149
171
|
and [pgvector guide](https://github.com/arnsri33/embedflow/blob/main/docs/integrations/pgvector.md),
|
|
150
172
|
and [Pinecone guide](https://github.com/arnsri33/embedflow/blob/main/docs/integrations/pinecone.md),
|
|
151
173
|
and [Milvus guide](https://github.com/arnsri33/embedflow/blob/main/docs/integrations/milvus.md).
|
|
174
|
+
See the [Weaviate guide](https://github.com/arnsri33/embedflow/blob/main/docs/integrations/weaviate.md).
|
|
152
175
|
|
|
153
176
|
## CLI
|
|
154
177
|
|
|
155
178
|
```bash
|
|
156
179
|
embedflow --help
|
|
157
180
|
embedflow analyze --help
|
|
181
|
+
embedflow plan --help
|
|
158
182
|
embedflow serve --config ./embedflow.yaml
|
|
159
183
|
embedflow status --config ./embedflow.yaml
|
|
160
184
|
embedflow registry list
|
|
@@ -168,7 +192,7 @@ cover the remaining commands and endpoints.
|
|
|
168
192
|
|
|
169
193
|
## Status
|
|
170
194
|
|
|
171
|
-
EmbedFlow v0.
|
|
195
|
+
EmbedFlow v0.6.0 is a pre-1.0 release for research and early real-world
|
|
172
196
|
testing. T2-v1 is an empirical finite-tail diagnostic, partial rankings can
|
|
173
197
|
differ from fully warm target reranking, and ANN fidelity needs a reference
|
|
174
198
|
comparison to audit.
|
|
@@ -48,6 +48,11 @@ The response includes the result list and migration fields such as:
|
|
|
48
48
|
request. A partial response scores the available target vectors; it can differ
|
|
49
49
|
from the fully warm ranking.
|
|
50
50
|
|
|
51
|
+
The advisory migration planner is exposed through the Python API and the
|
|
52
|
+
`embedflow plan` CLI. It is intentionally not a synchronous FastAPI endpoint:
|
|
53
|
+
probe analysis may load models and perform bounded candidate work, so operators
|
|
54
|
+
should generate a plan artifact out of band and publish/read it as needed.
|
|
55
|
+
|
|
51
56
|
`/status` includes safe backend metadata. For pgvector this names the
|
|
52
57
|
schema/table and vector contract; for Pinecone it names the host/index,
|
|
53
58
|
namespace, dimension, metric, and safe vector counts. Neither backend returns
|
|
@@ -9,6 +9,7 @@ list. The commands below are the main public entry points.
|
|
|
9
9
|
embedflow init --config ./embedflow.yaml
|
|
10
10
|
embedflow analyze --config ./embedflow.yaml --output-dir ./analysis
|
|
11
11
|
embedflow evaluate --config ./experiment.yaml --output-dir ./results
|
|
12
|
+
embedflow plan --config ./embedflow.yaml --queries ./probe_queries.jsonl
|
|
12
13
|
```
|
|
13
14
|
|
|
14
15
|
`analyze` is the no-target-index workflow. It uses probe queries and frozen
|
|
@@ -16,6 +17,15 @@ T2-v1 logic. `evaluate` is the labelled workflow; it computes source quality,
|
|
|
16
17
|
native target quality, target-within-source-candidates quality, candidate gap,
|
|
17
18
|
containment, and migration depth when the required inputs are available.
|
|
18
19
|
|
|
20
|
+
`plan` is the advisory migration planner. It runs source preflight and registry
|
|
21
|
+
matching, optionally samples a JSONL probe set, and reports a candidate-depth
|
|
22
|
+
recommendation, cache/economics projections, warnings, and staged rollout
|
|
23
|
+
guidance. Use `--format json` (or `yaml`) for a structured artifact. A plan
|
|
24
|
+
with no probes reports `T2-v1: NOT_RUN`; `SAFE` is a finite-tail diagnostic,
|
|
25
|
+
not a qrels-based retrieval-quality guarantee. A `DEFER`/`EXPAND_PROBE` result
|
|
26
|
+
is still a successful analytical command and exits zero; invalid configuration
|
|
27
|
+
or failed preflight exits nonzero.
|
|
28
|
+
|
|
19
29
|
## Serving and operations
|
|
20
30
|
|
|
21
31
|
```bash
|
|
@@ -14,7 +14,7 @@ target:
|
|
|
14
14
|
device: cuda
|
|
15
15
|
|
|
16
16
|
index:
|
|
17
|
-
backend: faiss # faiss, qdrant, pgvector, pinecone, or
|
|
17
|
+
backend: faiss # faiss, qdrant, pgvector, pinecone, milvus, or weaviate
|
|
18
18
|
path: ./legacy.index
|
|
19
19
|
ids: ./legacy.index.ids.json # FAISS sidecar
|
|
20
20
|
metric: cosine
|
|
@@ -26,6 +26,9 @@ index:
|
|
|
26
26
|
# text_metadata_field
|
|
27
27
|
# Milvus fields: uri, token_env, database, collection, id_field, vector_field,
|
|
28
28
|
# text_field, partition_names, search_params, auto_load
|
|
29
|
+
# Weaviate fields: uri (or http_host/http_port), grpc_host/grpc_port,
|
|
30
|
+
# secure, collection, vector_name, text_property,
|
|
31
|
+
# tenant, api_key_env
|
|
29
32
|
|
|
30
33
|
documents:
|
|
31
34
|
path: ./documents.jsonl
|
|
@@ -55,6 +58,25 @@ economics:
|
|
|
55
58
|
gpu_price_per_hour: null
|
|
56
59
|
target_docs_per_second: null
|
|
57
60
|
|
|
61
|
+
# Optional advisory planner settings. CLI flags override these values.
|
|
62
|
+
planner:
|
|
63
|
+
max_probes: 250
|
|
64
|
+
seed: 42
|
|
65
|
+
k_grid: [20, 50, 100, 200, 500]
|
|
66
|
+
max_candidates: null
|
|
67
|
+
max_target_encodes: null
|
|
68
|
+
max_sync_misses: null
|
|
69
|
+
background_batch_size: null
|
|
70
|
+
gpu_hourly_cost: null
|
|
71
|
+
target_docs_per_second: null
|
|
72
|
+
queries_per_second: null
|
|
73
|
+
daily_queries: null
|
|
74
|
+
cache_hit_rate: null
|
|
75
|
+
latency_budget_ms: null
|
|
76
|
+
access_trace: null
|
|
77
|
+
corpus_name: null
|
|
78
|
+
corpus_fingerprint: null
|
|
79
|
+
|
|
58
80
|
state_path: ./embedflow_state.json
|
|
59
81
|
```
|
|
60
82
|
|
|
@@ -106,6 +128,18 @@ EMBEDFLOW_MILVUS_VECTOR_FIELD
|
|
|
106
128
|
EMBEDFLOW_MILVUS_TEXT_FIELD
|
|
107
129
|
EMBEDFLOW_MILVUS_PARTITIONS
|
|
108
130
|
EMBEDFLOW_MILVUS_AUTO_LOAD
|
|
131
|
+
EMBEDFLOW_WEAVIATE_URI
|
|
132
|
+
EMBEDFLOW_WEAVIATE_HTTP_HOST
|
|
133
|
+
EMBEDFLOW_WEAVIATE_HTTP_PORT
|
|
134
|
+
EMBEDFLOW_WEAVIATE_GRPC_HOST
|
|
135
|
+
EMBEDFLOW_WEAVIATE_GRPC_PORT
|
|
136
|
+
EMBEDFLOW_WEAVIATE_SECURE
|
|
137
|
+
EMBEDFLOW_WEAVIATE_GRPC_SECURE
|
|
138
|
+
EMBEDFLOW_WEAVIATE_API_KEY_ENV
|
|
139
|
+
EMBEDFLOW_WEAVIATE_COLLECTION
|
|
140
|
+
EMBEDFLOW_WEAVIATE_VECTOR_NAME
|
|
141
|
+
EMBEDFLOW_WEAVIATE_TEXT_PROPERTY
|
|
142
|
+
EMBEDFLOW_WEAVIATE_TENANT
|
|
109
143
|
EMBEDFLOW_INDEX_NPROBE
|
|
110
144
|
EMBEDFLOW_DOCUMENTS_PATH
|
|
111
145
|
EMBEDFLOW_CACHE_PATH
|
|
@@ -115,6 +149,21 @@ EMBEDFLOW_CANDIDATE_DEPTH
|
|
|
115
149
|
EMBEDFLOW_MAX_SYNC_MISSES
|
|
116
150
|
EMBEDFLOW_BACKGROUND_BATCH_SIZE
|
|
117
151
|
EMBEDFLOW_PROBE_KMAX
|
|
152
|
+
EMBEDFLOW_PLANNER_MAX_PROBES
|
|
153
|
+
EMBEDFLOW_PLANNER_SEED
|
|
154
|
+
EMBEDFLOW_PLANNER_MAX_CANDIDATES
|
|
155
|
+
EMBEDFLOW_PLANNER_MAX_TARGET_ENCODINGS
|
|
156
|
+
EMBEDFLOW_PLANNER_MAX_SYNC_MISSES
|
|
157
|
+
EMBEDFLOW_PLANNER_BACKGROUND_BATCH_SIZE
|
|
158
|
+
EMBEDFLOW_PLANNER_GPU_HOURLY_COST
|
|
159
|
+
EMBEDFLOW_PLANNER_TARGET_DOCS_PER_SECOND
|
|
160
|
+
EMBEDFLOW_PLANNER_QPS
|
|
161
|
+
EMBEDFLOW_PLANNER_DAILY_QUERIES
|
|
162
|
+
EMBEDFLOW_PLANNER_CACHE_HIT_RATE
|
|
163
|
+
EMBEDFLOW_PLANNER_LATENCY_BUDGET_MS
|
|
164
|
+
EMBEDFLOW_PLANNER_ACCESS_TRACE
|
|
165
|
+
EMBEDFLOW_PLANNER_CORPUS_NAME
|
|
166
|
+
EMBEDFLOW_PLANNER_CORPUS_FINGERPRINT
|
|
118
167
|
```
|
|
119
168
|
|
|
120
169
|
## Input files
|
|
@@ -33,6 +33,7 @@ python -m pip install -e .
|
|
|
33
33
|
| `pgvector` | Psycopg 3 binary driver and pgvector adapter |
|
|
34
34
|
| `pinecone` | Official Pinecone Python SDK and adapter |
|
|
35
35
|
| `milvus` | Official pymilvus SDK and adapter |
|
|
36
|
+
| `weaviate` | Official Weaviate Python v4 SDK and adapter |
|
|
36
37
|
| `models` | PyTorch, Transformers, Sentence Transformers, and Hub client |
|
|
37
38
|
| `dashboard` | FastAPI, Uvicorn, and Pydantic |
|
|
38
39
|
| `dev` | Pytest, Ruff, and build tooling |
|
|
@@ -73,6 +74,12 @@ Install the optional adapter with `python -m pip install "embedflow[milvus]"`.
|
|
|
73
74
|
Configure an existing collection URI, database, and vector field; see
|
|
74
75
|
[`integrations/milvus.md`](integrations/milvus.md).
|
|
75
76
|
|
|
77
|
+
## Weaviate
|
|
78
|
+
|
|
79
|
+
Install the current v4 client with `python -m pip install "embedflow[weaviate]"`.
|
|
80
|
+
Configure an existing externally-vectorized collection, including both the
|
|
81
|
+
HTTP and gRPC endpoints; see [`integrations/weaviate.md`](integrations/weaviate.md).
|
|
82
|
+
|
|
76
83
|
## CPU and GPU
|
|
77
84
|
|
|
78
85
|
The deterministic demo runs on CPU. Real model serving accepts `--device cpu`
|
|
@@ -0,0 +1,114 @@
|
|
|
1
|
+
# Weaviate
|
|
2
|
+
|
|
3
|
+
EmbedFlow's Weaviate backend (tested with `weaviate-client` 4.23.1 and
|
|
4
|
+
Weaviate 1.30.6) reads an existing, externally vectorized dense collection.
|
|
5
|
+
It retrieves candidates with the v4 `near_vector` API, reranks them with the
|
|
6
|
+
target model, and progressively materializes target vectors in EmbedFlow's
|
|
7
|
+
local cache. It does not create, update, delete, batch-import, or re-vectorize
|
|
8
|
+
objects in the source collection.
|
|
9
|
+
|
|
10
|
+
## Install and connect
|
|
11
|
+
|
|
12
|
+
```bash
|
|
13
|
+
python -m pip install "embedflow[weaviate]"
|
|
14
|
+
```
|
|
15
|
+
|
|
16
|
+
The local v4 client needs both HTTP and gRPC. A standard Docker deployment
|
|
17
|
+
exposes HTTP 8080 and gRPC 50051:
|
|
18
|
+
|
|
19
|
+
```bash
|
|
20
|
+
cd examples/weaviate
|
|
21
|
+
docker compose up -d
|
|
22
|
+
```
|
|
23
|
+
|
|
24
|
+
Configuration uses a URI or explicit host/port fields:
|
|
25
|
+
|
|
26
|
+
```yaml
|
|
27
|
+
index:
|
|
28
|
+
backend: weaviate
|
|
29
|
+
uri: http://127.0.0.1:8080
|
|
30
|
+
http_port: 8080
|
|
31
|
+
grpc_host: 127.0.0.1
|
|
32
|
+
grpc_port: 50051
|
|
33
|
+
collection: Documents
|
|
34
|
+
vector_name: default
|
|
35
|
+
text_property: content
|
|
36
|
+
metric: cosine
|
|
37
|
+
api_key_env: WEAVIATE_API_KEY
|
|
38
|
+
```
|
|
39
|
+
|
|
40
|
+
For a custom or cloud-compatible endpoint set `http_host`, `http_port`,
|
|
41
|
+
`grpc_host`, `grpc_port`, `secure`, and `grpc_secure` as appropriate. Keep the
|
|
42
|
+
credential out of YAML:
|
|
43
|
+
|
|
44
|
+
```bash
|
|
45
|
+
export WEAVIATE_API_KEY='your-key'
|
|
46
|
+
embedflow doctor --config ./embedflow.yaml
|
|
47
|
+
```
|
|
48
|
+
|
|
49
|
+
Cloud connections use the same v4 URI/token conventions. This release has
|
|
50
|
+
validated local standalone Weaviate, not an independently operated Weaviate
|
|
51
|
+
Cloud cluster.
|
|
52
|
+
|
|
53
|
+
## Collection contract
|
|
54
|
+
|
|
55
|
+
The selected collection must contain a dense, externally supplied vector. Set
|
|
56
|
+
`vector_name` for named-vector collections; if more than one vector exists and
|
|
57
|
+
the name is omitted, EmbedFlow fails rather than guessing. `cosine`, `dot` (or
|
|
58
|
+
`inner_product`), and `l2`/`euclidean` are supported. Weaviate's distance is
|
|
59
|
+
converted to EmbedFlow's larger-is-better score convention (`-distance`), so
|
|
60
|
+
rank order remains the server's rank order.
|
|
61
|
+
|
|
62
|
+
Weaviate objects use UUIDs. EmbedFlow preserves their canonical string form;
|
|
63
|
+
it does not generate IDs during reads. The source model dimension must match
|
|
64
|
+
the existing vector index. Weaviate versions that do not expose dimensionality
|
|
65
|
+
in collection configuration report the value as unverified and the server
|
|
66
|
+
still validates each query vector.
|
|
67
|
+
|
|
68
|
+
For text, choose one of these modes:
|
|
69
|
+
|
|
70
|
+
* `text_property: content` requests only that property with each candidate and
|
|
71
|
+
validates that it is a string.
|
|
72
|
+
* Provide the normal EmbedFlow JSONL `documents` store. Candidate queries then
|
|
73
|
+
request no Weaviate properties and the shared document store resolves text.
|
|
74
|
+
|
|
75
|
+
The adapter requests no stored vector payload. `tenant` can bind a configured
|
|
76
|
+
multi-tenant collection to one tenant; an omitted tenant never selects one
|
|
77
|
+
implicitly.
|
|
78
|
+
|
|
79
|
+
## Audit and operations
|
|
80
|
+
|
|
81
|
+
```bash
|
|
82
|
+
embedflow audit-index --config ./embedflow.yaml
|
|
83
|
+
embedflow analyze --config ./embedflow.yaml
|
|
84
|
+
embedflow serve --config ./embedflow.yaml --device cuda
|
|
85
|
+
```
|
|
86
|
+
|
|
87
|
+
`doctor` and `audit-index` report readiness, collection/vector configuration,
|
|
88
|
+
metric, object count, tenant, and text resolution. Reachability is reported
|
|
89
|
+
separately from candidate compatibility and ANN fidelity; a healthy server is
|
|
90
|
+
not evidence that a model migration is scientifically safe. `client.close()`
|
|
91
|
+
is called when a one-shot command or serving session shuts down.
|
|
92
|
+
|
|
93
|
+
## Read-only and troubleshooting
|
|
94
|
+
|
|
95
|
+
The production adapter uses only collection lookup/config, `near_vector`,
|
|
96
|
+
batched property lookup, iteration, and aggregate counts. Fixture creation and
|
|
97
|
+
batch writes live only in `scripts/weaviate_fixture.py` and tests. Do not point
|
|
98
|
+
fixture setup at a user collection.
|
|
99
|
+
|
|
100
|
+
* `Weaviate support requires ...`: install the optional extra in the active
|
|
101
|
+
environment.
|
|
102
|
+
* HTTP works but startup says gRPC is unavailable: expose 50051 and set
|
|
103
|
+
`grpc_host`/`grpc_port` correctly; v4 search uses gRPC.
|
|
104
|
+
* `vector_name ... was not found` or multiple-vector errors: inspect the
|
|
105
|
+
collection schema and configure the exact name.
|
|
106
|
+
* Dimension or metric mismatch: use the source model contract and the
|
|
107
|
+
collection's configured distance metric.
|
|
108
|
+
* Missing/invalid text: fix `text_property` or use an external document store.
|
|
109
|
+
|
|
110
|
+
The adapter does not expose arbitrary Weaviate filters or index-tuning writes.
|
|
111
|
+
HNSW configuration is introspected read-only. Integrated Weaviate vectorizers
|
|
112
|
+
are not treated as the EmbedFlow source model unless their contract is
|
|
113
|
+
independently verified; this release focuses on externally vectorized dense
|
|
114
|
+
collections.
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Limitations and release scope
|
|
2
2
|
|
|
3
|
-
EmbedFlow v0.
|
|
3
|
+
EmbedFlow v0.6.0 is a pre-1.0 release for research and early real-world
|
|
4
4
|
testing. The serving path is designed to make migration experiments concrete;
|
|
5
5
|
production rollout still requires application-specific validation.
|
|
6
6
|
|
|
@@ -0,0 +1,111 @@
|
|
|
1
|
+
# Migration planner
|
|
2
|
+
|
|
3
|
+
`embedflow plan` is a bounded, advisory analysis for moving from an existing
|
|
4
|
+
source vector index to a target embedding model. It composes the existing
|
|
5
|
+
backend audit, model contracts, retained registry evidence, probe retrieval,
|
|
6
|
+
frozen T2-v1 diagnostic, cache behavior, and economics helpers. It never
|
|
7
|
+
routes traffic and does not create, update, or delete source-index data.
|
|
8
|
+
|
|
9
|
+
## Quickstart
|
|
10
|
+
|
|
11
|
+
Install EmbedFlow with the optional dependency for the configured backend, then
|
|
12
|
+
provide representative query probes as JSONL:
|
|
13
|
+
|
|
14
|
+
```json
|
|
15
|
+
{"id":"q1","query":"how do I reset my password?"}
|
|
16
|
+
{"id":"q2","query":"what is the vacation policy?"}
|
|
17
|
+
```
|
|
18
|
+
|
|
19
|
+
Run a terminal plan or save a machine-readable artifact:
|
|
20
|
+
|
|
21
|
+
```bash
|
|
22
|
+
embedflow plan --config embedflow.yaml --queries probes.jsonl
|
|
23
|
+
embedflow plan --config embedflow.yaml --queries probes.jsonl \
|
|
24
|
+
--k-grid 20,50,100,200,500 --format json --output migration-plan.json
|
|
25
|
+
```
|
|
26
|
+
|
|
27
|
+
For a deterministic, network-free walkthrough, run `examples/planner/run_demo.sh`.
|
|
28
|
+
The demo uses the same planner API and tiny hash encoders; it is not a quality
|
|
29
|
+
benchmark.
|
|
30
|
+
|
|
31
|
+
## What the result means
|
|
32
|
+
|
|
33
|
+
The result is a versioned (`schema_version: 1`) object with source/target
|
|
34
|
+
contracts, preflight checks, evidence, candidate-depth diagnostics, cache and
|
|
35
|
+
rollout guidance, economics, and structured warnings. The Python API is:
|
|
36
|
+
|
|
37
|
+
```python
|
|
38
|
+
from embedflow.planner import MigrationPlanner
|
|
39
|
+
|
|
40
|
+
result = MigrationPlanner("embedflow.yaml").plan("probes.jsonl")
|
|
41
|
+
artifact = result.to_dict()
|
|
42
|
+
```
|
|
43
|
+
|
|
44
|
+
Recommendation states are deliberately small:
|
|
45
|
+
|
|
46
|
+
- `PROCEED`: probe evidence and preflight are strong enough for an operator-controlled shadow/canary sequence.
|
|
47
|
+
- `PROCEED_WITH_CAUTION`: probes are acceptable, but evidence coverage or contract/ANN provenance is limited.
|
|
48
|
+
- `EXPAND_PROBE`: T2-v1 asks for a deeper pool or the evidence is too small.
|
|
49
|
+
- `DEFER`: finite-tail behavior is `UNSAFE_OR_UNCERTAIN`; do not canary based on this plan.
|
|
50
|
+
- `BLOCKED`: structural/backend preflight failed; no candidate K is certified.
|
|
51
|
+
|
|
52
|
+
These are recommendations, not automatic deployment actions. The suggested
|
|
53
|
+
rollout is Validate → Shadow → small canary → expanded canary → target-primary
|
|
54
|
+
path, with the source remaining authoritative until the operator approves each
|
|
55
|
+
step. Stop conditions include source/target errors, p95 latency, queue depth,
|
|
56
|
+
cache hit rate, and any native/qrels evaluation regressions.
|
|
57
|
+
|
|
58
|
+
## Evidence and T2-v1
|
|
59
|
+
|
|
60
|
+
Registry matching preserves the existing hierarchy: `EXACT REGISTRY MATCH`,
|
|
61
|
+
`PRIOR EVIDENCE AVAILABLE` (same contracts, different/unknown corpus),
|
|
62
|
+
`RELATED EVIDENCE ONLY`, and `NO REGISTRY MATCH`. Prior or related evidence can
|
|
63
|
+
seed context but cannot override current probe results or certify a new corpus.
|
|
64
|
+
|
|
65
|
+
The planner calls the existing frozen T2-v1 implementation unchanged. T2-v1 is
|
|
66
|
+
a no-qrels finite-tail diagnostic of candidate behavior. `SAFE` does **not**
|
|
67
|
+
prove a candidate gap, nDCG, recall, or zero quality loss. Without actual qrels
|
|
68
|
+
or native-target evaluation, `candidate_gap` is `null`/`UNKNOWN`. ANN fidelity
|
|
69
|
+
is also `UNKNOWN` unless an exact/reference comparison is supplied by the
|
|
70
|
+
backend audit workflow.
|
|
71
|
+
|
|
72
|
+
No probes is a valid preflight/economics mode, but T2 is `NOT_RUN`, confidence
|
|
73
|
+
is `INSUFFICIENT`, and no K is presented as safe. Probe JSONL rows may use
|
|
74
|
+
`query` (or the existing `text` alias) and an optional `id`; empty, malformed,
|
|
75
|
+
or duplicate IDs are rejected. `--max-probes` and `--seed` provide deterministic
|
|
76
|
+
bounded sampling. Candidate documents are deduplicated by canonical ID and
|
|
77
|
+
encoded in bounded batches; `--max-target-encodes` can impose a hard cap.
|
|
78
|
+
|
|
79
|
+
## Cache, performance, and economics
|
|
80
|
+
|
|
81
|
+
Planning uses an ephemeral target-cache/state directory when it opens a runtime,
|
|
82
|
+
so the normal serving cache is not polluted. It recommends conservative sync
|
|
83
|
+
miss and background batch settings and defaults to traffic-driven progressive
|
|
84
|
+
warming when no access trace is supplied. An access trace can be supplied with
|
|
85
|
+
`--access-trace` as rows such as `{"document_id":"abc","count":192}` to
|
|
86
|
+
model hot-document coverage.
|
|
87
|
+
|
|
88
|
+
Performance values retain provenance (`measured`, `user_supplied`, `modeled`,
|
|
89
|
+
`registry`, or `unknown`). `--profile` performs only a small local encode/search
|
|
90
|
+
sample with one unmeasured warmup pass excluded; it is diagnostic and is not a
|
|
91
|
+
formal benchmark. Economics accepts `--target-docs-per-second` and
|
|
92
|
+
`--gpu-hourly-cost`; absent inputs remain `UNKNOWN`. Raw vector storage is
|
|
93
|
+
`documents × target dimension × dtype bytes` and excludes ANN overhead,
|
|
94
|
+
metadata, replicas, backups, and database overhead. Progressive percentages
|
|
95
|
+
are materialization scenarios, not claims that a percentage is sufficient.
|
|
96
|
+
|
|
97
|
+
## Configuration and privacy
|
|
98
|
+
|
|
99
|
+
An optional `planner:` section mirrors the CLI options (`max_probes`, `seed`,
|
|
100
|
+
`k_grid`, work caps, cache hints, throughput/cost inputs, and corpus identity).
|
|
101
|
+
CLI values take precedence. `EMBEDFLOW_PLANNER_*` environment overrides follow
|
|
102
|
+
the normal config mechanism. Probe text is never copied into plan artifacts or
|
|
103
|
+
telemetry by default; only counts, IDs used for diagnostics, and aggregates are
|
|
104
|
+
reported. Backend credentials are handled by the existing backend adapters and
|
|
105
|
+
are redacted from errors and structured output.
|
|
106
|
+
|
|
107
|
+
The planner does not provide qrel evaluation, traffic routing, rollback
|
|
108
|
+
automation, distributed scheduling, ANN tuning, integrated-vectorizer contract
|
|
109
|
+
verification, or production cost guarantees. Use `embedflow evaluate` with
|
|
110
|
+
qrels/native target rankings when empirical retrieval-quality claims are
|
|
111
|
+
required.
|