embedflow 0.5.0__tar.gz → 0.7.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {embedflow-0.5.0 → embedflow-0.7.0}/CHANGELOG.md +21 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/CITATION.cff +1 -1
- {embedflow-0.5.0 → embedflow-0.7.0}/PKG-INFO +38 -2
- {embedflow-0.5.0 → embedflow-0.7.0}/README.md +35 -1
- {embedflow-0.5.0 → embedflow-0.7.0}/README_PYPI.md +37 -1
- {embedflow-0.5.0 → embedflow-0.7.0}/docs/api.md +13 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/docs/cli.md +17 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/docs/configuration.md +81 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/docs/limitations.md +1 -1
- embedflow-0.7.0/docs/planner.md +111 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/docs/releasing.md +8 -8
- embedflow-0.7.0/docs/shadow-mode.md +104 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/__init__.py +9 -3
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/cache/persistent_cache.py +23 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/cli.py +239 -11
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/config.py +297 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/migration/facade.py +50 -4
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/migration/materializer.py +65 -2
- embedflow-0.7.0/embedflow/planner/__init__.py +22 -0
- embedflow-0.7.0/embedflow/planner/economics.py +153 -0
- embedflow-0.7.0/embedflow/planner/models.py +97 -0
- embedflow-0.7.0/embedflow/planner/planner.py +1390 -0
- embedflow-0.7.0/embedflow/planner/rendering.py +93 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/runtime.py +60 -2
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/serving/api.py +26 -3
- embedflow-0.7.0/embedflow/serving/engine.py +630 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/serving/schemas.py +1 -0
- embedflow-0.7.0/embedflow/shadow/__init__.py +9 -0
- embedflow-0.7.0/embedflow/shadow/models.py +60 -0
- embedflow-0.7.0/embedflow/shadow/report.py +48 -0
- embedflow-0.7.0/embedflow/shadow/runner.py +328 -0
- embedflow-0.7.0/embedflow/shadow/telemetry.py +773 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow.egg-info/PKG-INFO +38 -2
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow.egg-info/SOURCES.txt +18 -0
- embedflow-0.7.0/examples/planner/README.md +12 -0
- embedflow-0.7.0/examples/planner/run_demo.sh +14 -0
- embedflow-0.7.0/examples/shadow/README.md +13 -0
- embedflow-0.7.0/examples/shadow/embedflow.yaml.example +25 -0
- embedflow-0.7.0/examples/shadow/run_demo.py +95 -0
- embedflow-0.7.0/examples/shadow/run_demo.sh +10 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/pyproject.toml +1 -1
- {embedflow-0.5.0 → embedflow-0.7.0}/scripts/release_gate.py +17 -4
- embedflow-0.5.0/embedflow/serving/engine.py +0 -228
- {embedflow-0.5.0 → embedflow-0.7.0}/CONTRIBUTING.md +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/LICENSE +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/MANIFEST.in +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/SECURITY.md +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/docs/assets/README.md +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/docs/assets/candidate-gap-example.svg +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/docs/assets/dashboard-screenshot.md +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/docs/assets/terminal-demo.txt +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/docs/concepts.md +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/docs/contributing-benchmarks.md +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/docs/economics.md +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/docs/installation.md +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/docs/integrations/faiss.md +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/docs/integrations/milvus.md +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/docs/integrations/pgvector.md +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/docs/integrations/pinecone.md +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/docs/integrations/qdrant.md +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/docs/integrations/weaviate.md +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/docs/methodology.md +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/docs/quickstart.md +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/docs/registry.md +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/__main__.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/analysis.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/api.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/cache/__init__.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/cache/base.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/compatibility/__init__.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/compatibility/candidate_gap.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/compatibility/containment.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/compatibility/evaluate.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/compatibility/metrics.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/compatibility/migration_depth.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/compatibility/probe.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/compatibility/report.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/compatibility/t2.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/data/__init__.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/data/registry/__init__.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/data/registry/benchmark_profiles.jsonl +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/data/registry/checksums.sha256 +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/data/registry/migrations.jsonl +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/data/registry/registry_manifest.json +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/data/registry/research_summaries.json +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/data/registry/schema_version.json +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/frozen/T2_V1_FROZEN_SPEC.md +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/frozen/T2_V1_FROZEN_SPEC.sha256 +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/indexes/__init__.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/indexes/base.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/indexes/faiss_backend.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/indexes/milvus_backend.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/indexes/pgvector_backend.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/indexes/pinecone_backend.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/indexes/qdrant_backend.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/indexes/weaviate_backend.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/metrics/__init__.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/metrics/latency.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/migration/__init__.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/migration/compatibility.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/migration/planner.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/migration/state.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/models/__init__.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/models/base.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/models/huggingface.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/registry/__init__.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/registry/loader.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/registry/matcher.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/registry/schema.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/serving/__init__.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/serving/factory.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow.egg-info/dependency_links.txt +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow.egg-info/entry_points.txt +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow.egg-info/requires.txt +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/embedflow.egg-info/top_level.txt +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/examples/faiss/README.md +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/examples/faiss/documents.jsonl +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/examples/faiss/embedflow.yaml +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/examples/faiss/queries.jsonl +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/examples/milvus/README.md +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/examples/milvus/compose.yaml +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/examples/milvus/embedflow.yaml.example +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/examples/milvus/run_demo.sh +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/examples/pgvector/README.md +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/examples/pgvector/build_index.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/examples/pgvector/compose.yaml +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/examples/pgvector/embedflow.yaml +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/examples/pgvector/init.sql +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/examples/pgvector/queries.jsonl +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/examples/pgvector/run_demo.sh +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/examples/pinecone/README.md +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/examples/pinecone/embedflow.yaml.example +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/examples/pinecone/run_smoke.sh +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/examples/qdrant/README.md +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/examples/qdrant/build_index.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/examples/qdrant/documents.jsonl +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/examples/qdrant/embedflow.yaml +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/examples/qdrant/queries.jsonl +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/examples/research_analysis/README.md +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/examples/research_analysis/documents.jsonl +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/examples/research_analysis/embedflow.yaml +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/examples/research_analysis/qrels.json +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/examples/research_analysis/queries.jsonl +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/examples/weaviate/README.md +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/examples/weaviate/compose.yaml +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/examples/weaviate/embedflow.yaml.example +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/examples/weaviate/run_demo.sh +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/frozen/T2_V1_FROZEN_SPEC.md +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/frozen/T2_V1_FROZEN_SPEC.sha256 +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/requirements-dev.txt +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/requirements.txt +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/scripts/milvus_fixture.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/scripts/pinecone_smoke.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/scripts/real_qdrant_smoke.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/scripts/run_demo.sh +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/scripts/run_tests.sh +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/scripts/validate_milvus.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/scripts/validate_pgvector_10k.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/scripts/validate_weaviate.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/scripts/weaviate_fixture.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/scripts/weaviate_smoke.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/setup.cfg +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/src/__init__.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/src/embed.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/src/probe_features.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/src/storage.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/src/t2_v1.py +0 -0
- {embedflow-0.5.0 → embedflow-0.7.0}/src/utils.py +0 -0
|
@@ -1,5 +1,26 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## v0.7.0 — Source-authoritative Shadow Mode
|
|
4
|
+
|
|
5
|
+
- Added bounded, deterministic, source-authoritative Shadow Mode for observing
|
|
6
|
+
target reranking on sampled traffic without delaying or changing source
|
|
7
|
+
responses.
|
|
8
|
+
- Added failure/timeout isolation, queue backpressure, optional asynchronous
|
|
9
|
+
target-cache materialization, target-coverage and ranking diagnostics, and
|
|
10
|
+
privacy-conscious SQLite telemetry.
|
|
11
|
+
- Added `embedflow shadow report`, API/status integration, documentation, and a
|
|
12
|
+
deterministic offline demonstration. Shadow reports remain operational
|
|
13
|
+
diagnostics and do not claim qrel-based retrieval quality.
|
|
14
|
+
|
|
15
|
+
## v0.6.0 — Migration planner
|
|
16
|
+
|
|
17
|
+
- Added an advisory `embedflow plan` command and Python API that combine
|
|
18
|
+
source-index preflight, registry evidence, representative probes, frozen
|
|
19
|
+
T2-v1 diagnostics, candidate-depth selection, cache planning, economics, and
|
|
20
|
+
staged rollout guidance without routing traffic or mutating the source.
|
|
21
|
+
- Added structured JSON/YAML plan artifacts with explicit warnings and
|
|
22
|
+
measured/user-supplied/modeled/unknown provenance.
|
|
23
|
+
|
|
3
24
|
## v0.5.0 — Weaviate backend
|
|
4
25
|
|
|
5
26
|
- Added a read-only Weaviate v4 backend for existing externally-vectorized
|
|
@@ -2,7 +2,7 @@ cff-version: 1.2.0
|
|
|
2
2
|
title: "EmbedFlow: Upgrading Legacy Embeddings Without Full Upfront Re-Embedding"
|
|
3
3
|
message: "If EmbedFlow contributes to your work, please cite this software release."
|
|
4
4
|
type: software
|
|
5
|
-
version: 0.
|
|
5
|
+
version: 0.7.0
|
|
6
6
|
date-released: 2026-09-13
|
|
7
7
|
repository-code: "https://github.com/arnsri33/embedflow"
|
|
8
8
|
url: "https://github.com/arnsri33/embedflow"
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.4
|
|
2
2
|
Name: embedflow
|
|
3
|
-
Version: 0.
|
|
3
|
+
Version: 0.7.0
|
|
4
4
|
Summary: Progressive embedding-model migration over existing vector indexes.
|
|
5
5
|
Author: Arnav Srivastav
|
|
6
6
|
License-Expression: AGPL-3.0-only
|
|
@@ -189,6 +189,41 @@ Selected core records (nDCG@10, `G(50)`):
|
|
|
189
189
|
See the [registry documentation](https://github.com/arnsri33/embedflow/blob/main/docs/registry.md)
|
|
190
190
|
for matching levels, contract fingerprints, and provenance.
|
|
191
191
|
|
|
192
|
+
## Advisory migration plans
|
|
193
|
+
|
|
194
|
+
Use representative query probes to produce a bounded migration recommendation
|
|
195
|
+
before serving:
|
|
196
|
+
|
|
197
|
+
```bash
|
|
198
|
+
embedflow plan --config ./embedflow.yaml --queries ./probe_queries.jsonl
|
|
199
|
+
embedflow plan --config ./embedflow.yaml --queries ./probe_queries.jsonl --format json --output migration-plan.json
|
|
200
|
+
```
|
|
201
|
+
|
|
202
|
+
The planner reports preflight, registry evidence, frozen T2-v1 finite-tail
|
|
203
|
+
behavior, candidate K, cache/economics projections, and staged rollout
|
|
204
|
+
guidance. It never routes traffic or mutates the source index; `SAFE` is not a
|
|
205
|
+
qrels-based retrieval-quality guarantee. See the [planner guide](https://github.com/arnsri33/embedflow/blob/main/docs/planner.md).
|
|
206
|
+
|
|
207
|
+
### Source-authoritative Shadow Mode
|
|
208
|
+
|
|
209
|
+
Run the target migration path beside sampled traffic while always returning the
|
|
210
|
+
source result:
|
|
211
|
+
|
|
212
|
+
```yaml
|
|
213
|
+
runtime: {mode: shadow}
|
|
214
|
+
shadow: {enabled: true, sample_rate: 0.10, candidate_k: 100, materialize: true}
|
|
215
|
+
```
|
|
216
|
+
|
|
217
|
+
```bash
|
|
218
|
+
embedflow serve --config ./embedflow.yaml
|
|
219
|
+
embedflow shadow report --config ./embedflow.yaml --since 24h --format json
|
|
220
|
+
```
|
|
221
|
+
|
|
222
|
+
Shadow work is asynchronous, bounded, privacy-conscious, and failure isolated.
|
|
223
|
+
Reports contain operational and ranking-disagreement diagnostics—not qrel-based
|
|
224
|
+
quality guarantees—and recommendations never route canary traffic. See the
|
|
225
|
+
[Shadow Mode guide](https://github.com/arnsri33/embedflow/blob/main/docs/shadow-mode.md).
|
|
226
|
+
|
|
192
227
|
## Research
|
|
193
228
|
|
|
194
229
|
For candidate depth `K`, EmbedFlow measures:
|
|
@@ -229,6 +264,7 @@ See the [Weaviate guide](https://github.com/arnsri33/embedflow/blob/main/docs/in
|
|
|
229
264
|
```bash
|
|
230
265
|
embedflow --help
|
|
231
266
|
embedflow analyze --help
|
|
267
|
+
embedflow plan --help
|
|
232
268
|
embedflow serve --config ./embedflow.yaml
|
|
233
269
|
embedflow status --config ./embedflow.yaml
|
|
234
270
|
embedflow registry list
|
|
@@ -242,7 +278,7 @@ cover the remaining commands and endpoints.
|
|
|
242
278
|
|
|
243
279
|
## Status
|
|
244
280
|
|
|
245
|
-
EmbedFlow v0.
|
|
281
|
+
EmbedFlow v0.7.0 is a pre-1.0 release for research and early real-world
|
|
246
282
|
testing. T2-v1 is an empirical finite-tail diagnostic, partial rankings can
|
|
247
283
|
differ from fully warm target reranking, and ANN fidelity needs a reference
|
|
248
284
|
comparison to audit.
|
|
@@ -142,6 +142,38 @@ session = embedflow.migrate(
|
|
|
142
142
|
results = session.search("what causes auroras?", top_k=10)
|
|
143
143
|
```
|
|
144
144
|
|
|
145
|
+
## Plan a migration
|
|
146
|
+
|
|
147
|
+
Build a conservative, evidence-aware recommendation before serving. The
|
|
148
|
+
planner reuses backend preflight, registry matching, and frozen T2-v1; it is
|
|
149
|
+
advisory and never routes traffic or mutates the source index.
|
|
150
|
+
|
|
151
|
+
```bash
|
|
152
|
+
embedflow plan --config ./embedflow.yaml --queries ./probe_queries.jsonl
|
|
153
|
+
embedflow plan --config ./embedflow.yaml --queries ./probe_queries.jsonl --format json --output migration-plan.json
|
|
154
|
+
```
|
|
155
|
+
|
|
156
|
+
`SAFE` is an empirical finite-tail signal, not a retrieval-quality guarantee.
|
|
157
|
+
See [`docs/planner.md`](https://github.com/arnsri33/embedflow/blob/main/docs/planner.md).
|
|
158
|
+
|
|
159
|
+
Observe the reviewed migration path on real traffic without changing the
|
|
160
|
+
source result:
|
|
161
|
+
|
|
162
|
+
```yaml
|
|
163
|
+
runtime: {mode: shadow}
|
|
164
|
+
shadow: {enabled: true, sample_rate: 0.10, candidate_k: 100, materialize: true}
|
|
165
|
+
```
|
|
166
|
+
|
|
167
|
+
```bash
|
|
168
|
+
embedflow serve --config ./embedflow.yaml
|
|
169
|
+
embedflow shadow report --config ./embedflow.yaml --since 24h
|
|
170
|
+
```
|
|
171
|
+
|
|
172
|
+
Shadow Mode is source-authoritative, bounded, and advisory. It records cache,
|
|
173
|
+
coverage, latency, and ranking-disagreement diagnostics; it does not claim
|
|
174
|
+
retrieval-quality preservation without qrels and never routes canary traffic.
|
|
175
|
+
See [`docs/shadow-mode.md`](https://github.com/arnsri33/embedflow/blob/main/docs/shadow-mode.md).
|
|
176
|
+
|
|
145
177
|
Search responses expose `COLD`, `PARTIAL`, or `WARM`, cache hits and misses,
|
|
146
178
|
synchronous work, queued work, and stage timings. Once the candidate vectors
|
|
147
179
|
are warm, target scoring over that candidate set is deterministic.
|
|
@@ -227,6 +259,7 @@ Backend-specific setup and examples:
|
|
|
227
259
|
```bash
|
|
228
260
|
embedflow --help
|
|
229
261
|
embedflow analyze --help
|
|
262
|
+
embedflow plan --help
|
|
230
263
|
embedflow serve --config ./embedflow.yaml
|
|
231
264
|
embedflow status --config ./embedflow.yaml
|
|
232
265
|
embedflow registry list
|
|
@@ -248,13 +281,14 @@ OpenAPI documentation; see
|
|
|
248
281
|
- [CLI reference](https://github.com/arnsri33/embedflow/blob/main/docs/cli.md)
|
|
249
282
|
- [API](https://github.com/arnsri33/embedflow/blob/main/docs/api.md)
|
|
250
283
|
- [Economics](https://github.com/arnsri33/embedflow/blob/main/docs/economics.md)
|
|
284
|
+
- [Shadow Mode](https://github.com/arnsri33/embedflow/blob/main/docs/shadow-mode.md)
|
|
251
285
|
- [Limitations](https://github.com/arnsri33/embedflow/blob/main/docs/limitations.md)
|
|
252
286
|
- [Contributing](https://github.com/arnsri33/embedflow/blob/main/CONTRIBUTING.md)
|
|
253
287
|
- [Security](https://github.com/arnsri33/embedflow/blob/main/SECURITY.md)
|
|
254
288
|
|
|
255
289
|
## Status
|
|
256
290
|
|
|
257
|
-
EmbedFlow v0.
|
|
291
|
+
EmbedFlow v0.7.0 is a pre-1.0 release for research and early real-world
|
|
258
292
|
testing.
|
|
259
293
|
|
|
260
294
|
- T2-v1 reports an empirical finite-tail diagnostic.
|
|
@@ -123,6 +123,41 @@ Selected core records (nDCG@10, `G(50)`):
|
|
|
123
123
|
See the [registry documentation](https://github.com/arnsri33/embedflow/blob/main/docs/registry.md)
|
|
124
124
|
for matching levels, contract fingerprints, and provenance.
|
|
125
125
|
|
|
126
|
+
## Advisory migration plans
|
|
127
|
+
|
|
128
|
+
Use representative query probes to produce a bounded migration recommendation
|
|
129
|
+
before serving:
|
|
130
|
+
|
|
131
|
+
```bash
|
|
132
|
+
embedflow plan --config ./embedflow.yaml --queries ./probe_queries.jsonl
|
|
133
|
+
embedflow plan --config ./embedflow.yaml --queries ./probe_queries.jsonl --format json --output migration-plan.json
|
|
134
|
+
```
|
|
135
|
+
|
|
136
|
+
The planner reports preflight, registry evidence, frozen T2-v1 finite-tail
|
|
137
|
+
behavior, candidate K, cache/economics projections, and staged rollout
|
|
138
|
+
guidance. It never routes traffic or mutates the source index; `SAFE` is not a
|
|
139
|
+
qrels-based retrieval-quality guarantee. See the [planner guide](https://github.com/arnsri33/embedflow/blob/main/docs/planner.md).
|
|
140
|
+
|
|
141
|
+
### Source-authoritative Shadow Mode
|
|
142
|
+
|
|
143
|
+
Run the target migration path beside sampled traffic while always returning the
|
|
144
|
+
source result:
|
|
145
|
+
|
|
146
|
+
```yaml
|
|
147
|
+
runtime: {mode: shadow}
|
|
148
|
+
shadow: {enabled: true, sample_rate: 0.10, candidate_k: 100, materialize: true}
|
|
149
|
+
```
|
|
150
|
+
|
|
151
|
+
```bash
|
|
152
|
+
embedflow serve --config ./embedflow.yaml
|
|
153
|
+
embedflow shadow report --config ./embedflow.yaml --since 24h --format json
|
|
154
|
+
```
|
|
155
|
+
|
|
156
|
+
Shadow work is asynchronous, bounded, privacy-conscious, and failure isolated.
|
|
157
|
+
Reports contain operational and ranking-disagreement diagnostics—not qrel-based
|
|
158
|
+
quality guarantees—and recommendations never route canary traffic. See the
|
|
159
|
+
[Shadow Mode guide](https://github.com/arnsri33/embedflow/blob/main/docs/shadow-mode.md).
|
|
160
|
+
|
|
126
161
|
## Research
|
|
127
162
|
|
|
128
163
|
For candidate depth `K`, EmbedFlow measures:
|
|
@@ -163,6 +198,7 @@ See the [Weaviate guide](https://github.com/arnsri33/embedflow/blob/main/docs/in
|
|
|
163
198
|
```bash
|
|
164
199
|
embedflow --help
|
|
165
200
|
embedflow analyze --help
|
|
201
|
+
embedflow plan --help
|
|
166
202
|
embedflow serve --config ./embedflow.yaml
|
|
167
203
|
embedflow status --config ./embedflow.yaml
|
|
168
204
|
embedflow registry list
|
|
@@ -176,7 +212,7 @@ cover the remaining commands and endpoints.
|
|
|
176
212
|
|
|
177
213
|
## Status
|
|
178
214
|
|
|
179
|
-
EmbedFlow v0.
|
|
215
|
+
EmbedFlow v0.7.0 is a pre-1.0 release for research and early real-world
|
|
180
216
|
testing. T2-v1 is an empirical finite-tail diagnostic, partial rankings can
|
|
181
217
|
differ from fully warm target reranking, and ANN fidelity needs a reference
|
|
182
218
|
comparison to audit.
|
|
@@ -21,6 +21,7 @@ Interactive OpenAPI documentation is available at
|
|
|
21
21
|
| GET | `/metrics` | Aggregated latency and queue metrics |
|
|
22
22
|
| GET | `/plan` | Current migration plan |
|
|
23
23
|
| GET | `/economics` | Configured economics estimate |
|
|
24
|
+
| GET | `/shadow/report` | Bounded Shadow Mode diagnostics |
|
|
24
25
|
|
|
25
26
|
## Search
|
|
26
27
|
|
|
@@ -48,6 +49,18 @@ The response includes the result list and migration fields such as:
|
|
|
48
49
|
request. A partial response scores the available target vectors; it can differ
|
|
49
50
|
from the fully warm ranking.
|
|
50
51
|
|
|
52
|
+
When `runtime.mode=shadow`, `/search` returns the source-only result and marks
|
|
53
|
+
`migration.source_authoritative=true`; target work is scheduled in the
|
|
54
|
+
background. `/shadow/report`, `/status`, and `/metrics` expose aggregate shadow
|
|
55
|
+
observations without raw query/document text or vectors. Shadow failures and
|
|
56
|
+
timeouts are isolated from the response. Ranking overlap is diagnostic and is
|
|
57
|
+
not a qrels-based quality claim.
|
|
58
|
+
|
|
59
|
+
The advisory migration planner is exposed through the Python API and the
|
|
60
|
+
`embedflow plan` CLI. It is intentionally not a synchronous FastAPI endpoint:
|
|
61
|
+
probe analysis may load models and perform bounded candidate work, so operators
|
|
62
|
+
should generate a plan artifact out of band and publish/read it as needed.
|
|
63
|
+
|
|
51
64
|
`/status` includes safe backend metadata. For pgvector this names the
|
|
52
65
|
schema/table and vector contract; for Pinecone it names the host/index,
|
|
53
66
|
namespace, dimension, metric, and safe vector counts. Neither backend returns
|
|
@@ -9,6 +9,8 @@ list. The commands below are the main public entry points.
|
|
|
9
9
|
embedflow init --config ./embedflow.yaml
|
|
10
10
|
embedflow analyze --config ./embedflow.yaml --output-dir ./analysis
|
|
11
11
|
embedflow evaluate --config ./experiment.yaml --output-dir ./results
|
|
12
|
+
embedflow plan --config ./embedflow.yaml --queries ./probe_queries.jsonl
|
|
13
|
+
embedflow shadow report --config ./embedflow.yaml --since 24h
|
|
12
14
|
```
|
|
13
15
|
|
|
14
16
|
`analyze` is the no-target-index workflow. It uses probe queries and frozen
|
|
@@ -16,6 +18,15 @@ T2-v1 logic. `evaluate` is the labelled workflow; it computes source quality,
|
|
|
16
18
|
native target quality, target-within-source-candidates quality, candidate gap,
|
|
17
19
|
containment, and migration depth when the required inputs are available.
|
|
18
20
|
|
|
21
|
+
`plan` is the advisory migration planner. It runs source preflight and registry
|
|
22
|
+
matching, optionally samples a JSONL probe set, and reports a candidate-depth
|
|
23
|
+
recommendation, cache/economics projections, warnings, and staged rollout
|
|
24
|
+
guidance. Use `--format json` (or `yaml`) for a structured artifact. A plan
|
|
25
|
+
with no probes reports `T2-v1: NOT_RUN`; `SAFE` is a finite-tail diagnostic,
|
|
26
|
+
not a qrels-based retrieval-quality guarantee. A `DEFER`/`EXPAND_PROBE` result
|
|
27
|
+
is still a successful analytical command and exits zero; invalid configuration
|
|
28
|
+
or failed preflight exits nonzero.
|
|
29
|
+
|
|
19
30
|
## Serving and operations
|
|
20
31
|
|
|
21
32
|
```bash
|
|
@@ -32,6 +43,12 @@ materialization throughput. `audit-index` checks the source index against an
|
|
|
32
43
|
exact/reference configuration where supported. `prewarm` schedules target
|
|
33
44
|
document work; it does not change source-index results.
|
|
34
45
|
|
|
46
|
+
Use `embedflow serve --mode shadow` to opt into source-authoritative Shadow
|
|
47
|
+
Mode for one process. `embedflow shadow report` reads the bounded telemetry
|
|
48
|
+
store and supports `--format text|json|yaml`, `--since`, `--output`, and
|
|
49
|
+
`--quiet`. A valid `DEFER`/`EXPAND_K` report is still an analytical success;
|
|
50
|
+
only invalid configuration or initialization returns a non-zero exit code.
|
|
51
|
+
|
|
35
52
|
## Registry and profiles
|
|
36
53
|
|
|
37
54
|
```bash
|
|
@@ -58,9 +58,57 @@ economics:
|
|
|
58
58
|
gpu_price_per_hour: null
|
|
59
59
|
target_docs_per_second: null
|
|
60
60
|
|
|
61
|
+
# Optional advisory planner settings. CLI flags override these values.
|
|
62
|
+
planner:
|
|
63
|
+
max_probes: 250
|
|
64
|
+
seed: 42
|
|
65
|
+
k_grid: [20, 50, 100, 200, 500]
|
|
66
|
+
max_candidates: null
|
|
67
|
+
max_target_encodes: null
|
|
68
|
+
max_sync_misses: null
|
|
69
|
+
background_batch_size: null
|
|
70
|
+
gpu_hourly_cost: null
|
|
71
|
+
target_docs_per_second: null
|
|
72
|
+
queries_per_second: null
|
|
73
|
+
daily_queries: null
|
|
74
|
+
cache_hit_rate: null
|
|
75
|
+
latency_budget_ms: null
|
|
76
|
+
access_trace: null
|
|
77
|
+
corpus_name: null
|
|
78
|
+
corpus_fingerprint: null
|
|
79
|
+
|
|
80
|
+
# Source-authoritative observation mode. ``runtime.mode: migration`` (the
|
|
81
|
+
# default) leaves Shadow Mode inactive. ``mode: shadow`` is an explicit opt-in.
|
|
82
|
+
runtime:
|
|
83
|
+
mode: migration # migration, normal, source, or shadow
|
|
84
|
+
|
|
85
|
+
shadow:
|
|
86
|
+
enabled: true
|
|
87
|
+
sample_rate: 0.10
|
|
88
|
+
sample_seed: 42
|
|
89
|
+
candidate_k: 100
|
|
90
|
+
materialize: true
|
|
91
|
+
max_inflight: 32
|
|
92
|
+
queue_capacity: 1000
|
|
93
|
+
timeout_ms: 10000
|
|
94
|
+
shutdown_grace_ms: 1000
|
|
95
|
+
telemetry:
|
|
96
|
+
enabled: true
|
|
97
|
+
path: ./.embedflow/shadow.sqlite3
|
|
98
|
+
retain_query_records: false
|
|
99
|
+
retain_query_text: false
|
|
100
|
+
max_records: 10000
|
|
101
|
+
retention_days: null
|
|
102
|
+
report_k: 10
|
|
103
|
+
min_target_coverage_for_ranking: 1.0
|
|
104
|
+
|
|
61
105
|
state_path: ./embedflow_state.json
|
|
62
106
|
```
|
|
63
107
|
|
|
108
|
+
Shadow Mode always returns the source-authoritative result before target work
|
|
109
|
+
finishes. See [`shadow-mode.md`](shadow-mode.md) for queue, timeout, privacy,
|
|
110
|
+
materialization, and report semantics.
|
|
111
|
+
|
|
64
112
|
## Model contracts
|
|
65
113
|
|
|
66
114
|
Model configuration can include `revision`, `dimension`, `max_length`,
|
|
@@ -130,6 +178,39 @@ EMBEDFLOW_CANDIDATE_DEPTH
|
|
|
130
178
|
EMBEDFLOW_MAX_SYNC_MISSES
|
|
131
179
|
EMBEDFLOW_BACKGROUND_BATCH_SIZE
|
|
132
180
|
EMBEDFLOW_PROBE_KMAX
|
|
181
|
+
EMBEDFLOW_PLANNER_MAX_PROBES
|
|
182
|
+
EMBEDFLOW_PLANNER_SEED
|
|
183
|
+
EMBEDFLOW_PLANNER_MAX_CANDIDATES
|
|
184
|
+
EMBEDFLOW_PLANNER_MAX_TARGET_ENCODINGS
|
|
185
|
+
EMBEDFLOW_PLANNER_MAX_SYNC_MISSES
|
|
186
|
+
EMBEDFLOW_PLANNER_BACKGROUND_BATCH_SIZE
|
|
187
|
+
EMBEDFLOW_PLANNER_GPU_HOURLY_COST
|
|
188
|
+
EMBEDFLOW_PLANNER_TARGET_DOCS_PER_SECOND
|
|
189
|
+
EMBEDFLOW_PLANNER_QPS
|
|
190
|
+
EMBEDFLOW_PLANNER_DAILY_QUERIES
|
|
191
|
+
EMBEDFLOW_PLANNER_CACHE_HIT_RATE
|
|
192
|
+
EMBEDFLOW_PLANNER_LATENCY_BUDGET_MS
|
|
193
|
+
EMBEDFLOW_PLANNER_ACCESS_TRACE
|
|
194
|
+
EMBEDFLOW_PLANNER_CORPUS_NAME
|
|
195
|
+
EMBEDFLOW_PLANNER_CORPUS_FINGERPRINT
|
|
196
|
+
EMBEDFLOW_RUNTIME_MODE
|
|
197
|
+
EMBEDFLOW_SHADOW_ENABLED
|
|
198
|
+
EMBEDFLOW_SHADOW_SAMPLE_RATE
|
|
199
|
+
EMBEDFLOW_SHADOW_SAMPLE_SEED
|
|
200
|
+
EMBEDFLOW_SHADOW_CANDIDATE_K
|
|
201
|
+
EMBEDFLOW_SHADOW_MATERIALIZE
|
|
202
|
+
EMBEDFLOW_SHADOW_MAX_INFLIGHT
|
|
203
|
+
EMBEDFLOW_SHADOW_QUEUE_CAPACITY
|
|
204
|
+
EMBEDFLOW_SHADOW_TIMEOUT_MS
|
|
205
|
+
EMBEDFLOW_SHADOW_SHUTDOWN_GRACE_MS
|
|
206
|
+
EMBEDFLOW_SHADOW_TELEMETRY_ENABLED
|
|
207
|
+
EMBEDFLOW_SHADOW_TELEMETRY_PATH
|
|
208
|
+
EMBEDFLOW_SHADOW_RETAIN_QUERY_RECORDS
|
|
209
|
+
EMBEDFLOW_SHADOW_RETAIN_QUERY_TEXT
|
|
210
|
+
EMBEDFLOW_SHADOW_MAX_RECORDS
|
|
211
|
+
EMBEDFLOW_SHADOW_REPORT_K
|
|
212
|
+
EMBEDFLOW_SHADOW_RETENTION_DAYS
|
|
213
|
+
EMBEDFLOW_SHADOW_MIN_TARGET_COVERAGE
|
|
133
214
|
```
|
|
134
215
|
|
|
135
216
|
## Input files
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Limitations and release scope
|
|
2
2
|
|
|
3
|
-
EmbedFlow v0.
|
|
3
|
+
EmbedFlow v0.7.0 is a pre-1.0 release for research and early real-world
|
|
4
4
|
testing. The serving path is designed to make migration experiments concrete;
|
|
5
5
|
production rollout still requires application-specific validation.
|
|
6
6
|
|
|
@@ -0,0 +1,111 @@
|
|
|
1
|
+
# Migration planner
|
|
2
|
+
|
|
3
|
+
`embedflow plan` is a bounded, advisory analysis for moving from an existing
|
|
4
|
+
source vector index to a target embedding model. It composes the existing
|
|
5
|
+
backend audit, model contracts, retained registry evidence, probe retrieval,
|
|
6
|
+
frozen T2-v1 diagnostic, cache behavior, and economics helpers. It never
|
|
7
|
+
routes traffic and does not create, update, or delete source-index data.
|
|
8
|
+
|
|
9
|
+
## Quickstart
|
|
10
|
+
|
|
11
|
+
Install EmbedFlow with the optional dependency for the configured backend, then
|
|
12
|
+
provide representative query probes as JSONL:
|
|
13
|
+
|
|
14
|
+
```json
|
|
15
|
+
{"id":"q1","query":"how do I reset my password?"}
|
|
16
|
+
{"id":"q2","query":"what is the vacation policy?"}
|
|
17
|
+
```
|
|
18
|
+
|
|
19
|
+
Run a terminal plan or save a machine-readable artifact:
|
|
20
|
+
|
|
21
|
+
```bash
|
|
22
|
+
embedflow plan --config embedflow.yaml --queries probes.jsonl
|
|
23
|
+
embedflow plan --config embedflow.yaml --queries probes.jsonl \
|
|
24
|
+
--k-grid 20,50,100,200,500 --format json --output migration-plan.json
|
|
25
|
+
```
|
|
26
|
+
|
|
27
|
+
For a deterministic, network-free walkthrough, run `examples/planner/run_demo.sh`.
|
|
28
|
+
The demo uses the same planner API and tiny hash encoders; it is not a quality
|
|
29
|
+
benchmark.
|
|
30
|
+
|
|
31
|
+
## What the result means
|
|
32
|
+
|
|
33
|
+
The result is a versioned (`schema_version: 1`) object with source/target
|
|
34
|
+
contracts, preflight checks, evidence, candidate-depth diagnostics, cache and
|
|
35
|
+
rollout guidance, economics, and structured warnings. The Python API is:
|
|
36
|
+
|
|
37
|
+
```python
|
|
38
|
+
from embedflow.planner import MigrationPlanner
|
|
39
|
+
|
|
40
|
+
result = MigrationPlanner("embedflow.yaml").plan("probes.jsonl")
|
|
41
|
+
artifact = result.to_dict()
|
|
42
|
+
```
|
|
43
|
+
|
|
44
|
+
Recommendation states are deliberately small:
|
|
45
|
+
|
|
46
|
+
- `PROCEED`: probe evidence and preflight are strong enough for an operator-controlled shadow/canary sequence.
|
|
47
|
+
- `PROCEED_WITH_CAUTION`: probes are acceptable, but evidence coverage or contract/ANN provenance is limited.
|
|
48
|
+
- `EXPAND_PROBE`: T2-v1 asks for a deeper pool or the evidence is too small.
|
|
49
|
+
- `DEFER`: finite-tail behavior is `UNSAFE_OR_UNCERTAIN`; do not canary based on this plan.
|
|
50
|
+
- `BLOCKED`: structural/backend preflight failed; no candidate K is certified.
|
|
51
|
+
|
|
52
|
+
These are recommendations, not automatic deployment actions. The suggested
|
|
53
|
+
rollout is Validate → Shadow → small canary → expanded canary → target-primary
|
|
54
|
+
path, with the source remaining authoritative until the operator approves each
|
|
55
|
+
step. Stop conditions include source/target errors, p95 latency, queue depth,
|
|
56
|
+
cache hit rate, and any native/qrels evaluation regressions.
|
|
57
|
+
|
|
58
|
+
## Evidence and T2-v1
|
|
59
|
+
|
|
60
|
+
Registry matching preserves the existing hierarchy: `EXACT REGISTRY MATCH`,
|
|
61
|
+
`PRIOR EVIDENCE AVAILABLE` (same contracts, different/unknown corpus),
|
|
62
|
+
`RELATED EVIDENCE ONLY`, and `NO REGISTRY MATCH`. Prior or related evidence can
|
|
63
|
+
seed context but cannot override current probe results or certify a new corpus.
|
|
64
|
+
|
|
65
|
+
The planner calls the existing frozen T2-v1 implementation unchanged. T2-v1 is
|
|
66
|
+
a no-qrels finite-tail diagnostic of candidate behavior. `SAFE` does **not**
|
|
67
|
+
prove a candidate gap, nDCG, recall, or zero quality loss. Without actual qrels
|
|
68
|
+
or native-target evaluation, `candidate_gap` is `null`/`UNKNOWN`. ANN fidelity
|
|
69
|
+
is also `UNKNOWN` unless an exact/reference comparison is supplied by the
|
|
70
|
+
backend audit workflow.
|
|
71
|
+
|
|
72
|
+
No probes is a valid preflight/economics mode, but T2 is `NOT_RUN`, confidence
|
|
73
|
+
is `INSUFFICIENT`, and no K is presented as safe. Probe JSONL rows may use
|
|
74
|
+
`query` (or the existing `text` alias) and an optional `id`; empty, malformed,
|
|
75
|
+
or duplicate IDs are rejected. `--max-probes` and `--seed` provide deterministic
|
|
76
|
+
bounded sampling. Candidate documents are deduplicated by canonical ID and
|
|
77
|
+
encoded in bounded batches; `--max-target-encodes` can impose a hard cap.
|
|
78
|
+
|
|
79
|
+
## Cache, performance, and economics
|
|
80
|
+
|
|
81
|
+
Planning uses an ephemeral target-cache/state directory when it opens a runtime,
|
|
82
|
+
so the normal serving cache is not polluted. It recommends conservative sync
|
|
83
|
+
miss and background batch settings and defaults to traffic-driven progressive
|
|
84
|
+
warming when no access trace is supplied. An access trace can be supplied with
|
|
85
|
+
`--access-trace` as rows such as `{"document_id":"abc","count":192}` to
|
|
86
|
+
model hot-document coverage.
|
|
87
|
+
|
|
88
|
+
Performance values retain provenance (`measured`, `user_supplied`, `modeled`,
|
|
89
|
+
`registry`, or `unknown`). `--profile` performs only a small local encode/search
|
|
90
|
+
sample with one unmeasured warmup pass excluded; it is diagnostic and is not a
|
|
91
|
+
formal benchmark. Economics accepts `--target-docs-per-second` and
|
|
92
|
+
`--gpu-hourly-cost`; absent inputs remain `UNKNOWN`. Raw vector storage is
|
|
93
|
+
`documents × target dimension × dtype bytes` and excludes ANN overhead,
|
|
94
|
+
metadata, replicas, backups, and database overhead. Progressive percentages
|
|
95
|
+
are materialization scenarios, not claims that a percentage is sufficient.
|
|
96
|
+
|
|
97
|
+
## Configuration and privacy
|
|
98
|
+
|
|
99
|
+
An optional `planner:` section mirrors the CLI options (`max_probes`, `seed`,
|
|
100
|
+
`k_grid`, work caps, cache hints, throughput/cost inputs, and corpus identity).
|
|
101
|
+
CLI values take precedence. `EMBEDFLOW_PLANNER_*` environment overrides follow
|
|
102
|
+
the normal config mechanism. Probe text is never copied into plan artifacts or
|
|
103
|
+
telemetry by default; only counts, IDs used for diagnostics, and aggregates are
|
|
104
|
+
reported. Backend credentials are handled by the existing backend adapters and
|
|
105
|
+
are redacted from errors and structured output.
|
|
106
|
+
|
|
107
|
+
The planner does not provide qrel evaluation, traffic routing, rollback
|
|
108
|
+
automation, distributed scheduling, ANN tuning, integrated-vectorizer contract
|
|
109
|
+
verification, or production cost guarantees. Use `embedflow evaluate` with
|
|
110
|
+
qrels/native target rankings when empirical retrieval-quality claims are
|
|
111
|
+
required.
|
|
@@ -18,8 +18,8 @@ python -m twine check dist/*
|
|
|
18
18
|
Inspect both archives before uploading:
|
|
19
19
|
|
|
20
20
|
```bash
|
|
21
|
-
unzip -l dist/embedflow-0.
|
|
22
|
-
tar -tzf dist/embedflow-0.
|
|
21
|
+
unzip -l dist/embedflow-0.7.0-py3-none-any.whl
|
|
22
|
+
tar -tzf dist/embedflow-0.7.0.tar.gz
|
|
23
23
|
sha256sum dist/*
|
|
24
24
|
```
|
|
25
25
|
|
|
@@ -32,7 +32,7 @@ Test the wheel outside the source tree:
|
|
|
32
32
|
```bash
|
|
33
33
|
python -m venv /tmp/embedflow-wheel-test
|
|
34
34
|
/tmp/embedflow-wheel-test/bin/python -m pip install --upgrade pip
|
|
35
|
-
/tmp/embedflow-wheel-test/bin/python -m pip install dist/embedflow-0.
|
|
35
|
+
/tmp/embedflow-wheel-test/bin/python -m pip install dist/embedflow-0.7.0-py3-none-any.whl
|
|
36
36
|
cd /tmp
|
|
37
37
|
/tmp/embedflow-wheel-test/bin/python -c "import embedflow; print(embedflow.__version__)"
|
|
38
38
|
/tmp/embedflow-wheel-test/bin/embedflow --help
|
|
@@ -65,7 +65,7 @@ python -m venv /tmp/embedflow-testpypi
|
|
|
65
65
|
/tmp/embedflow-testpypi/bin/python -m pip install \
|
|
66
66
|
--index-url https://test.pypi.org/simple/ \
|
|
67
67
|
--extra-index-url https://pypi.org/simple/ \
|
|
68
|
-
embedflow==0.
|
|
68
|
+
embedflow==0.7.0
|
|
69
69
|
cd /tmp
|
|
70
70
|
/tmp/embedflow-testpypi/bin/python -c "import embedflow; print(embedflow.__version__)"
|
|
71
71
|
/tmp/embedflow-testpypi/bin/embedflow --help
|
|
@@ -79,11 +79,11 @@ Test optional integrations in a second clean environment:
|
|
|
79
79
|
/tmp/embedflow-testpypi/bin/python -m pip install \
|
|
80
80
|
--index-url https://test.pypi.org/simple/ \
|
|
81
81
|
--extra-index-url https://pypi.org/simple/ \
|
|
82
|
-
"embedflow[faiss,dashboard,pinecone,milvus,weaviate]==0.
|
|
82
|
+
"embedflow[faiss,dashboard,pinecone,milvus,weaviate]==0.7.0"
|
|
83
83
|
```
|
|
84
84
|
|
|
85
|
-
If the same filename already exists on TestPyPI, use a pre-release
|
|
86
|
-
|
|
85
|
+
If the same filename already exists on TestPyPI, use a pre-release for the
|
|
86
|
+
TestPyPI-only trial. Never overwrite the production version.
|
|
87
87
|
|
|
88
88
|
## Trusted Publishing configuration
|
|
89
89
|
|
|
@@ -115,7 +115,7 @@ above keeps the test step explicit.
|
|
|
115
115
|
3. Run the final release gate and review the generated report.
|
|
116
116
|
4. Configure the PyPI pending publisher and protected `pypi` environment.
|
|
117
117
|
5. Create a Git tag and GitHub Release for the exact package version, for
|
|
118
|
-
example `v0.
|
|
118
|
+
example `v0.7.0`.
|
|
119
119
|
6. Approve the `pypi` environment when the release workflow is ready.
|
|
120
120
|
7. Verify the files and metadata on PyPI.
|
|
121
121
|
8. Install from production PyPI in a directory outside this checkout.
|
|
@@ -0,0 +1,104 @@
|
|
|
1
|
+
# Shadow Mode
|
|
2
|
+
|
|
3
|
+
Shadow Mode runs the configured target migration path beside real traffic while
|
|
4
|
+
the existing source result remains authoritative. It is an observation and
|
|
5
|
+
cache-warming tool, not a traffic router:
|
|
6
|
+
|
|
7
|
+
```yaml
|
|
8
|
+
runtime:
|
|
9
|
+
mode: shadow
|
|
10
|
+
|
|
11
|
+
shadow:
|
|
12
|
+
enabled: true
|
|
13
|
+
sample_rate: 0.10
|
|
14
|
+
sample_seed: 42
|
|
15
|
+
candidate_k: 100
|
|
16
|
+
materialize: true
|
|
17
|
+
max_inflight: 32
|
|
18
|
+
queue_capacity: 1000
|
|
19
|
+
timeout_ms: 10000
|
|
20
|
+
telemetry:
|
|
21
|
+
enabled: true
|
|
22
|
+
path: ./.embedflow/shadow.sqlite3
|
|
23
|
+
retain_query_records: false
|
|
24
|
+
retain_query_text: false
|
|
25
|
+
```
|
|
26
|
+
|
|
27
|
+
Install the optional backend/model extras required by the selected
|
|
28
|
+
configuration, then start the ordinary service:
|
|
29
|
+
|
|
30
|
+
```bash
|
|
31
|
+
pip install "embedflow[dashboard]"
|
|
32
|
+
embedflow serve --config embedflow.yaml
|
|
33
|
+
# or override the file for one process:
|
|
34
|
+
embedflow serve --config embedflow.yaml --mode shadow
|
|
35
|
+
```
|
|
36
|
+
|
|
37
|
+
The request path performs source encoding, source ANN retrieval, and document
|
|
38
|
+
resolution exactly as source-only serving. It returns that result without
|
|
39
|
+
waiting for target encoding, cache misses, reranking, telemetry, or the
|
|
40
|
+
materialization worker. Shadow jobs use a bounded queue and worker pool. A full
|
|
41
|
+
queue drops only shadow work; a target/model/cache/telemetry failure or timeout
|
|
42
|
+
is recorded and cannot change the primary response.
|
|
43
|
+
|
|
44
|
+
If the target model cannot be loaded during startup, source/shadow serving
|
|
45
|
+
still opens with target work marked unavailable; sampled requests record the
|
|
46
|
+
isolated target-encoding failure. Normal migration mode remains fail-fast for
|
|
47
|
+
the same startup error.
|
|
48
|
+
|
|
49
|
+
`candidate_k` is explicit. The planner may suggest a value, but Shadow Mode
|
|
50
|
+
does not silently select one. `materialize: false` reads already cached target
|
|
51
|
+
vectors without changing cache or queue state. With `materialize: true`, cold
|
|
52
|
+
candidate IDs are deduplicated in the existing persistent materialization queue
|
|
53
|
+
and warmed asynchronously; synchronous target misses remain zero in Shadow
|
|
54
|
+
Mode. A comparison with incomplete target vectors is reported as `partial`,
|
|
55
|
+
with target coverage shown separately.
|
|
56
|
+
|
|
57
|
+
## Reports
|
|
58
|
+
|
|
59
|
+
Reports use the privacy-preserving SQLite telemetry store (by default beside
|
|
60
|
+
the configured cache):
|
|
61
|
+
|
|
62
|
+
```bash
|
|
63
|
+
embedflow shadow report --config embedflow.yaml --since 24h
|
|
64
|
+
embedflow shadow report --config embedflow.yaml --since 24h --format json --output shadow-report.json
|
|
65
|
+
```
|
|
66
|
+
|
|
67
|
+
Use `/shadow/report` or the `shadow` section of `/status` and `/metrics` in the
|
|
68
|
+
FastAPI service. JSON/YAML stdout contains only the requested artifact; progress
|
|
69
|
+
and errors belong on stderr. `--since` accepts seconds or `s`, `m`, `h`, `d`,
|
|
70
|
+
and `w` suffixes.
|
|
71
|
+
|
|
72
|
+
Reports contain sampled/completed/partial/failed/timed-out/dropped counts,
|
|
73
|
+
cache and materialization accounting, source and shadow latency distributions,
|
|
74
|
+
target coverage, top-1 agreement, and top-k overlap. Shadow latency is not
|
|
75
|
+
user-facing latency because it is off the primary critical path. Ranking
|
|
76
|
+
disagreement and overlap are diagnostics, not recall, nDCG, or quality loss;
|
|
77
|
+
without qrels/native-target evaluation no quality guarantee is made.
|
|
78
|
+
|
|
79
|
+
T2-v1 remains the frozen window-level diagnostic. EmbedFlow does not label an
|
|
80
|
+
individual production query `T2 SAFE`. A report recommendation is operational
|
|
81
|
+
guidance only:
|
|
82
|
+
|
|
83
|
+
- `CONTINUE_SHADOW` means evidence or coverage is still limited.
|
|
84
|
+
- `EXPAND_K` means the accumulated T2 window requests a larger candidate pool.
|
|
85
|
+
- `INVESTIGATE` means failures, timeouts, or uncertain T2 behavior need review.
|
|
86
|
+
- `READY_FOR_CANARY_EVALUATION` means an operator may consider a separately
|
|
87
|
+
designed canary; it never routes traffic automatically.
|
|
88
|
+
|
|
89
|
+
Raw query text, candidate text, source vectors, target vectors, and credentials
|
|
90
|
+
are not persisted in telemetry. Query IDs may be retained only when explicitly
|
|
91
|
+
enabled, and retention is bounded by `max_records`. The telemetry database is
|
|
92
|
+
segmented by source/target contract and candidate-K fingerprint so unrelated
|
|
93
|
+
migrations are not mixed.
|
|
94
|
+
|
|
95
|
+
On shutdown EmbedFlow stops accepting new shadow work, drops queued jobs when
|
|
96
|
+
necessary, and waits only the configured bounded grace period. A corrupted or
|
|
97
|
+
unwritable telemetry file degrades observability; it does not take source
|
|
98
|
+
serving down. Source indexes are read-only during Shadow Mode. The planner and
|
|
99
|
+
Shadow Mode are separate: generate a plan first, then copy its recommended K
|
|
100
|
+
explicitly into a reviewed Shadow configuration.
|
|
101
|
+
|
|
102
|
+
Shadow Mode is advisory and experimental operational instrumentation. It does
|
|
103
|
+
not provide autonomous rollout, rollback, qrel evaluation, or ANN-fidelity
|
|
104
|
+
proof.
|