embedflow 0.5.0__tar.gz → 0.7.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (168) hide show
  1. {embedflow-0.5.0 → embedflow-0.7.0}/CHANGELOG.md +21 -0
  2. {embedflow-0.5.0 → embedflow-0.7.0}/CITATION.cff +1 -1
  3. {embedflow-0.5.0 → embedflow-0.7.0}/PKG-INFO +38 -2
  4. {embedflow-0.5.0 → embedflow-0.7.0}/README.md +35 -1
  5. {embedflow-0.5.0 → embedflow-0.7.0}/README_PYPI.md +37 -1
  6. {embedflow-0.5.0 → embedflow-0.7.0}/docs/api.md +13 -0
  7. {embedflow-0.5.0 → embedflow-0.7.0}/docs/cli.md +17 -0
  8. {embedflow-0.5.0 → embedflow-0.7.0}/docs/configuration.md +81 -0
  9. {embedflow-0.5.0 → embedflow-0.7.0}/docs/limitations.md +1 -1
  10. embedflow-0.7.0/docs/planner.md +111 -0
  11. {embedflow-0.5.0 → embedflow-0.7.0}/docs/releasing.md +8 -8
  12. embedflow-0.7.0/docs/shadow-mode.md +104 -0
  13. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/__init__.py +9 -3
  14. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/cache/persistent_cache.py +23 -0
  15. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/cli.py +239 -11
  16. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/config.py +297 -0
  17. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/migration/facade.py +50 -4
  18. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/migration/materializer.py +65 -2
  19. embedflow-0.7.0/embedflow/planner/__init__.py +22 -0
  20. embedflow-0.7.0/embedflow/planner/economics.py +153 -0
  21. embedflow-0.7.0/embedflow/planner/models.py +97 -0
  22. embedflow-0.7.0/embedflow/planner/planner.py +1390 -0
  23. embedflow-0.7.0/embedflow/planner/rendering.py +93 -0
  24. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/runtime.py +60 -2
  25. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/serving/api.py +26 -3
  26. embedflow-0.7.0/embedflow/serving/engine.py +630 -0
  27. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/serving/schemas.py +1 -0
  28. embedflow-0.7.0/embedflow/shadow/__init__.py +9 -0
  29. embedflow-0.7.0/embedflow/shadow/models.py +60 -0
  30. embedflow-0.7.0/embedflow/shadow/report.py +48 -0
  31. embedflow-0.7.0/embedflow/shadow/runner.py +328 -0
  32. embedflow-0.7.0/embedflow/shadow/telemetry.py +773 -0
  33. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow.egg-info/PKG-INFO +38 -2
  34. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow.egg-info/SOURCES.txt +18 -0
  35. embedflow-0.7.0/examples/planner/README.md +12 -0
  36. embedflow-0.7.0/examples/planner/run_demo.sh +14 -0
  37. embedflow-0.7.0/examples/shadow/README.md +13 -0
  38. embedflow-0.7.0/examples/shadow/embedflow.yaml.example +25 -0
  39. embedflow-0.7.0/examples/shadow/run_demo.py +95 -0
  40. embedflow-0.7.0/examples/shadow/run_demo.sh +10 -0
  41. {embedflow-0.5.0 → embedflow-0.7.0}/pyproject.toml +1 -1
  42. {embedflow-0.5.0 → embedflow-0.7.0}/scripts/release_gate.py +17 -4
  43. embedflow-0.5.0/embedflow/serving/engine.py +0 -228
  44. {embedflow-0.5.0 → embedflow-0.7.0}/CONTRIBUTING.md +0 -0
  45. {embedflow-0.5.0 → embedflow-0.7.0}/LICENSE +0 -0
  46. {embedflow-0.5.0 → embedflow-0.7.0}/MANIFEST.in +0 -0
  47. {embedflow-0.5.0 → embedflow-0.7.0}/SECURITY.md +0 -0
  48. {embedflow-0.5.0 → embedflow-0.7.0}/docs/assets/README.md +0 -0
  49. {embedflow-0.5.0 → embedflow-0.7.0}/docs/assets/candidate-gap-example.svg +0 -0
  50. {embedflow-0.5.0 → embedflow-0.7.0}/docs/assets/dashboard-screenshot.md +0 -0
  51. {embedflow-0.5.0 → embedflow-0.7.0}/docs/assets/terminal-demo.txt +0 -0
  52. {embedflow-0.5.0 → embedflow-0.7.0}/docs/concepts.md +0 -0
  53. {embedflow-0.5.0 → embedflow-0.7.0}/docs/contributing-benchmarks.md +0 -0
  54. {embedflow-0.5.0 → embedflow-0.7.0}/docs/economics.md +0 -0
  55. {embedflow-0.5.0 → embedflow-0.7.0}/docs/installation.md +0 -0
  56. {embedflow-0.5.0 → embedflow-0.7.0}/docs/integrations/faiss.md +0 -0
  57. {embedflow-0.5.0 → embedflow-0.7.0}/docs/integrations/milvus.md +0 -0
  58. {embedflow-0.5.0 → embedflow-0.7.0}/docs/integrations/pgvector.md +0 -0
  59. {embedflow-0.5.0 → embedflow-0.7.0}/docs/integrations/pinecone.md +0 -0
  60. {embedflow-0.5.0 → embedflow-0.7.0}/docs/integrations/qdrant.md +0 -0
  61. {embedflow-0.5.0 → embedflow-0.7.0}/docs/integrations/weaviate.md +0 -0
  62. {embedflow-0.5.0 → embedflow-0.7.0}/docs/methodology.md +0 -0
  63. {embedflow-0.5.0 → embedflow-0.7.0}/docs/quickstart.md +0 -0
  64. {embedflow-0.5.0 → embedflow-0.7.0}/docs/registry.md +0 -0
  65. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/__main__.py +0 -0
  66. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/analysis.py +0 -0
  67. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/api.py +0 -0
  68. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/cache/__init__.py +0 -0
  69. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/cache/base.py +0 -0
  70. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/compatibility/__init__.py +0 -0
  71. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/compatibility/candidate_gap.py +0 -0
  72. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/compatibility/containment.py +0 -0
  73. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/compatibility/evaluate.py +0 -0
  74. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/compatibility/metrics.py +0 -0
  75. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/compatibility/migration_depth.py +0 -0
  76. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/compatibility/probe.py +0 -0
  77. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/compatibility/report.py +0 -0
  78. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/compatibility/t2.py +0 -0
  79. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/data/__init__.py +0 -0
  80. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/data/registry/__init__.py +0 -0
  81. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/data/registry/benchmark_profiles.jsonl +0 -0
  82. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/data/registry/checksums.sha256 +0 -0
  83. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/data/registry/migrations.jsonl +0 -0
  84. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/data/registry/registry_manifest.json +0 -0
  85. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/data/registry/research_summaries.json +0 -0
  86. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/data/registry/schema_version.json +0 -0
  87. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/frozen/T2_V1_FROZEN_SPEC.md +0 -0
  88. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/frozen/T2_V1_FROZEN_SPEC.sha256 +0 -0
  89. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/indexes/__init__.py +0 -0
  90. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/indexes/base.py +0 -0
  91. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/indexes/faiss_backend.py +0 -0
  92. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/indexes/milvus_backend.py +0 -0
  93. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/indexes/pgvector_backend.py +0 -0
  94. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/indexes/pinecone_backend.py +0 -0
  95. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/indexes/qdrant_backend.py +0 -0
  96. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/indexes/weaviate_backend.py +0 -0
  97. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/metrics/__init__.py +0 -0
  98. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/metrics/latency.py +0 -0
  99. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/migration/__init__.py +0 -0
  100. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/migration/compatibility.py +0 -0
  101. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/migration/planner.py +0 -0
  102. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/migration/state.py +0 -0
  103. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/models/__init__.py +0 -0
  104. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/models/base.py +0 -0
  105. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/models/huggingface.py +0 -0
  106. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/registry/__init__.py +0 -0
  107. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/registry/loader.py +0 -0
  108. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/registry/matcher.py +0 -0
  109. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/registry/schema.py +0 -0
  110. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/serving/__init__.py +0 -0
  111. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow/serving/factory.py +0 -0
  112. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow.egg-info/dependency_links.txt +0 -0
  113. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow.egg-info/entry_points.txt +0 -0
  114. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow.egg-info/requires.txt +0 -0
  115. {embedflow-0.5.0 → embedflow-0.7.0}/embedflow.egg-info/top_level.txt +0 -0
  116. {embedflow-0.5.0 → embedflow-0.7.0}/examples/faiss/README.md +0 -0
  117. {embedflow-0.5.0 → embedflow-0.7.0}/examples/faiss/documents.jsonl +0 -0
  118. {embedflow-0.5.0 → embedflow-0.7.0}/examples/faiss/embedflow.yaml +0 -0
  119. {embedflow-0.5.0 → embedflow-0.7.0}/examples/faiss/queries.jsonl +0 -0
  120. {embedflow-0.5.0 → embedflow-0.7.0}/examples/milvus/README.md +0 -0
  121. {embedflow-0.5.0 → embedflow-0.7.0}/examples/milvus/compose.yaml +0 -0
  122. {embedflow-0.5.0 → embedflow-0.7.0}/examples/milvus/embedflow.yaml.example +0 -0
  123. {embedflow-0.5.0 → embedflow-0.7.0}/examples/milvus/run_demo.sh +0 -0
  124. {embedflow-0.5.0 → embedflow-0.7.0}/examples/pgvector/README.md +0 -0
  125. {embedflow-0.5.0 → embedflow-0.7.0}/examples/pgvector/build_index.py +0 -0
  126. {embedflow-0.5.0 → embedflow-0.7.0}/examples/pgvector/compose.yaml +0 -0
  127. {embedflow-0.5.0 → embedflow-0.7.0}/examples/pgvector/embedflow.yaml +0 -0
  128. {embedflow-0.5.0 → embedflow-0.7.0}/examples/pgvector/init.sql +0 -0
  129. {embedflow-0.5.0 → embedflow-0.7.0}/examples/pgvector/queries.jsonl +0 -0
  130. {embedflow-0.5.0 → embedflow-0.7.0}/examples/pgvector/run_demo.sh +0 -0
  131. {embedflow-0.5.0 → embedflow-0.7.0}/examples/pinecone/README.md +0 -0
  132. {embedflow-0.5.0 → embedflow-0.7.0}/examples/pinecone/embedflow.yaml.example +0 -0
  133. {embedflow-0.5.0 → embedflow-0.7.0}/examples/pinecone/run_smoke.sh +0 -0
  134. {embedflow-0.5.0 → embedflow-0.7.0}/examples/qdrant/README.md +0 -0
  135. {embedflow-0.5.0 → embedflow-0.7.0}/examples/qdrant/build_index.py +0 -0
  136. {embedflow-0.5.0 → embedflow-0.7.0}/examples/qdrant/documents.jsonl +0 -0
  137. {embedflow-0.5.0 → embedflow-0.7.0}/examples/qdrant/embedflow.yaml +0 -0
  138. {embedflow-0.5.0 → embedflow-0.7.0}/examples/qdrant/queries.jsonl +0 -0
  139. {embedflow-0.5.0 → embedflow-0.7.0}/examples/research_analysis/README.md +0 -0
  140. {embedflow-0.5.0 → embedflow-0.7.0}/examples/research_analysis/documents.jsonl +0 -0
  141. {embedflow-0.5.0 → embedflow-0.7.0}/examples/research_analysis/embedflow.yaml +0 -0
  142. {embedflow-0.5.0 → embedflow-0.7.0}/examples/research_analysis/qrels.json +0 -0
  143. {embedflow-0.5.0 → embedflow-0.7.0}/examples/research_analysis/queries.jsonl +0 -0
  144. {embedflow-0.5.0 → embedflow-0.7.0}/examples/weaviate/README.md +0 -0
  145. {embedflow-0.5.0 → embedflow-0.7.0}/examples/weaviate/compose.yaml +0 -0
  146. {embedflow-0.5.0 → embedflow-0.7.0}/examples/weaviate/embedflow.yaml.example +0 -0
  147. {embedflow-0.5.0 → embedflow-0.7.0}/examples/weaviate/run_demo.sh +0 -0
  148. {embedflow-0.5.0 → embedflow-0.7.0}/frozen/T2_V1_FROZEN_SPEC.md +0 -0
  149. {embedflow-0.5.0 → embedflow-0.7.0}/frozen/T2_V1_FROZEN_SPEC.sha256 +0 -0
  150. {embedflow-0.5.0 → embedflow-0.7.0}/requirements-dev.txt +0 -0
  151. {embedflow-0.5.0 → embedflow-0.7.0}/requirements.txt +0 -0
  152. {embedflow-0.5.0 → embedflow-0.7.0}/scripts/milvus_fixture.py +0 -0
  153. {embedflow-0.5.0 → embedflow-0.7.0}/scripts/pinecone_smoke.py +0 -0
  154. {embedflow-0.5.0 → embedflow-0.7.0}/scripts/real_qdrant_smoke.py +0 -0
  155. {embedflow-0.5.0 → embedflow-0.7.0}/scripts/run_demo.sh +0 -0
  156. {embedflow-0.5.0 → embedflow-0.7.0}/scripts/run_tests.sh +0 -0
  157. {embedflow-0.5.0 → embedflow-0.7.0}/scripts/validate_milvus.py +0 -0
  158. {embedflow-0.5.0 → embedflow-0.7.0}/scripts/validate_pgvector_10k.py +0 -0
  159. {embedflow-0.5.0 → embedflow-0.7.0}/scripts/validate_weaviate.py +0 -0
  160. {embedflow-0.5.0 → embedflow-0.7.0}/scripts/weaviate_fixture.py +0 -0
  161. {embedflow-0.5.0 → embedflow-0.7.0}/scripts/weaviate_smoke.py +0 -0
  162. {embedflow-0.5.0 → embedflow-0.7.0}/setup.cfg +0 -0
  163. {embedflow-0.5.0 → embedflow-0.7.0}/src/__init__.py +0 -0
  164. {embedflow-0.5.0 → embedflow-0.7.0}/src/embed.py +0 -0
  165. {embedflow-0.5.0 → embedflow-0.7.0}/src/probe_features.py +0 -0
  166. {embedflow-0.5.0 → embedflow-0.7.0}/src/storage.py +0 -0
  167. {embedflow-0.5.0 → embedflow-0.7.0}/src/t2_v1.py +0 -0
  168. {embedflow-0.5.0 → embedflow-0.7.0}/src/utils.py +0 -0
@@ -1,5 +1,26 @@
1
1
  # Changelog
2
2
 
3
+ ## v0.7.0 — Source-authoritative Shadow Mode
4
+
5
+ - Added bounded, deterministic, source-authoritative Shadow Mode for observing
6
+ target reranking on sampled traffic without delaying or changing source
7
+ responses.
8
+ - Added failure/timeout isolation, queue backpressure, optional asynchronous
9
+ target-cache materialization, target-coverage and ranking diagnostics, and
10
+ privacy-conscious SQLite telemetry.
11
+ - Added `embedflow shadow report`, API/status integration, documentation, and a
12
+ deterministic offline demonstration. Shadow reports remain operational
13
+ diagnostics and do not claim qrel-based retrieval quality.
14
+
15
+ ## v0.6.0 — Migration planner
16
+
17
+ - Added an advisory `embedflow plan` command and Python API that combine
18
+ source-index preflight, registry evidence, representative probes, frozen
19
+ T2-v1 diagnostics, candidate-depth selection, cache planning, economics, and
20
+ staged rollout guidance without routing traffic or mutating the source.
21
+ - Added structured JSON/YAML plan artifacts with explicit warnings and
22
+ measured/user-supplied/modeled/unknown provenance.
23
+
3
24
  ## v0.5.0 — Weaviate backend
4
25
 
5
26
  - Added a read-only Weaviate v4 backend for existing externally-vectorized
@@ -2,7 +2,7 @@ cff-version: 1.2.0
2
2
  title: "EmbedFlow: Upgrading Legacy Embeddings Without Full Upfront Re-Embedding"
3
3
  message: "If EmbedFlow contributes to your work, please cite this software release."
4
4
  type: software
5
- version: 0.5.0
5
+ version: 0.7.0
6
6
  date-released: 2026-09-13
7
7
  repository-code: "https://github.com/arnsri33/embedflow"
8
8
  url: "https://github.com/arnsri33/embedflow"
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: embedflow
3
- Version: 0.5.0
3
+ Version: 0.7.0
4
4
  Summary: Progressive embedding-model migration over existing vector indexes.
5
5
  Author: Arnav Srivastav
6
6
  License-Expression: AGPL-3.0-only
@@ -189,6 +189,41 @@ Selected core records (nDCG@10, `G(50)`):
189
189
  See the [registry documentation](https://github.com/arnsri33/embedflow/blob/main/docs/registry.md)
190
190
  for matching levels, contract fingerprints, and provenance.
191
191
 
192
+ ## Advisory migration plans
193
+
194
+ Use representative query probes to produce a bounded migration recommendation
195
+ before serving:
196
+
197
+ ```bash
198
+ embedflow plan --config ./embedflow.yaml --queries ./probe_queries.jsonl
199
+ embedflow plan --config ./embedflow.yaml --queries ./probe_queries.jsonl --format json --output migration-plan.json
200
+ ```
201
+
202
+ The planner reports preflight, registry evidence, frozen T2-v1 finite-tail
203
+ behavior, candidate K, cache/economics projections, and staged rollout
204
+ guidance. It never routes traffic or mutates the source index; `SAFE` is not a
205
+ qrels-based retrieval-quality guarantee. See the [planner guide](https://github.com/arnsri33/embedflow/blob/main/docs/planner.md).
206
+
207
+ ### Source-authoritative Shadow Mode
208
+
209
+ Run the target migration path beside sampled traffic while always returning the
210
+ source result:
211
+
212
+ ```yaml
213
+ runtime: {mode: shadow}
214
+ shadow: {enabled: true, sample_rate: 0.10, candidate_k: 100, materialize: true}
215
+ ```
216
+
217
+ ```bash
218
+ embedflow serve --config ./embedflow.yaml
219
+ embedflow shadow report --config ./embedflow.yaml --since 24h --format json
220
+ ```
221
+
222
+ Shadow work is asynchronous, bounded, privacy-conscious, and failure isolated.
223
+ Reports contain operational and ranking-disagreement diagnostics—not qrel-based
224
+ quality guarantees—and recommendations never route canary traffic. See the
225
+ [Shadow Mode guide](https://github.com/arnsri33/embedflow/blob/main/docs/shadow-mode.md).
226
+
192
227
  ## Research
193
228
 
194
229
  For candidate depth `K`, EmbedFlow measures:
@@ -229,6 +264,7 @@ See the [Weaviate guide](https://github.com/arnsri33/embedflow/blob/main/docs/in
229
264
  ```bash
230
265
  embedflow --help
231
266
  embedflow analyze --help
267
+ embedflow plan --help
232
268
  embedflow serve --config ./embedflow.yaml
233
269
  embedflow status --config ./embedflow.yaml
234
270
  embedflow registry list
@@ -242,7 +278,7 @@ cover the remaining commands and endpoints.
242
278
 
243
279
  ## Status
244
280
 
245
- EmbedFlow v0.5.0 is an alpha release for research and early real-world
281
+ EmbedFlow v0.7.0 is a pre-1.0 release for research and early real-world
246
282
  testing. T2-v1 is an empirical finite-tail diagnostic, partial rankings can
247
283
  differ from fully warm target reranking, and ANN fidelity needs a reference
248
284
  comparison to audit.
@@ -142,6 +142,38 @@ session = embedflow.migrate(
142
142
  results = session.search("what causes auroras?", top_k=10)
143
143
  ```
144
144
 
145
+ ## Plan a migration
146
+
147
+ Build a conservative, evidence-aware recommendation before serving. The
148
+ planner reuses backend preflight, registry matching, and frozen T2-v1; it is
149
+ advisory and never routes traffic or mutates the source index.
150
+
151
+ ```bash
152
+ embedflow plan --config ./embedflow.yaml --queries ./probe_queries.jsonl
153
+ embedflow plan --config ./embedflow.yaml --queries ./probe_queries.jsonl --format json --output migration-plan.json
154
+ ```
155
+
156
+ `SAFE` is an empirical finite-tail signal, not a retrieval-quality guarantee.
157
+ See [`docs/planner.md`](https://github.com/arnsri33/embedflow/blob/main/docs/planner.md).
158
+
159
+ Observe the reviewed migration path on real traffic without changing the
160
+ source result:
161
+
162
+ ```yaml
163
+ runtime: {mode: shadow}
164
+ shadow: {enabled: true, sample_rate: 0.10, candidate_k: 100, materialize: true}
165
+ ```
166
+
167
+ ```bash
168
+ embedflow serve --config ./embedflow.yaml
169
+ embedflow shadow report --config ./embedflow.yaml --since 24h
170
+ ```
171
+
172
+ Shadow Mode is source-authoritative, bounded, and advisory. It records cache,
173
+ coverage, latency, and ranking-disagreement diagnostics; it does not claim
174
+ retrieval-quality preservation without qrels and never routes canary traffic.
175
+ See [`docs/shadow-mode.md`](https://github.com/arnsri33/embedflow/blob/main/docs/shadow-mode.md).
176
+
145
177
  Search responses expose `COLD`, `PARTIAL`, or `WARM`, cache hits and misses,
146
178
  synchronous work, queued work, and stage timings. Once the candidate vectors
147
179
  are warm, target scoring over that candidate set is deterministic.
@@ -227,6 +259,7 @@ Backend-specific setup and examples:
227
259
  ```bash
228
260
  embedflow --help
229
261
  embedflow analyze --help
262
+ embedflow plan --help
230
263
  embedflow serve --config ./embedflow.yaml
231
264
  embedflow status --config ./embedflow.yaml
232
265
  embedflow registry list
@@ -248,13 +281,14 @@ OpenAPI documentation; see
248
281
  - [CLI reference](https://github.com/arnsri33/embedflow/blob/main/docs/cli.md)
249
282
  - [API](https://github.com/arnsri33/embedflow/blob/main/docs/api.md)
250
283
  - [Economics](https://github.com/arnsri33/embedflow/blob/main/docs/economics.md)
284
+ - [Shadow Mode](https://github.com/arnsri33/embedflow/blob/main/docs/shadow-mode.md)
251
285
  - [Limitations](https://github.com/arnsri33/embedflow/blob/main/docs/limitations.md)
252
286
  - [Contributing](https://github.com/arnsri33/embedflow/blob/main/CONTRIBUTING.md)
253
287
  - [Security](https://github.com/arnsri33/embedflow/blob/main/SECURITY.md)
254
288
 
255
289
  ## Status
256
290
 
257
- EmbedFlow v0.5.0 is an alpha release for research and early real-world
291
+ EmbedFlow v0.7.0 is a pre-1.0 release for research and early real-world
258
292
  testing.
259
293
 
260
294
  - T2-v1 reports an empirical finite-tail diagnostic.
@@ -123,6 +123,41 @@ Selected core records (nDCG@10, `G(50)`):
123
123
  See the [registry documentation](https://github.com/arnsri33/embedflow/blob/main/docs/registry.md)
124
124
  for matching levels, contract fingerprints, and provenance.
125
125
 
126
+ ## Advisory migration plans
127
+
128
+ Use representative query probes to produce a bounded migration recommendation
129
+ before serving:
130
+
131
+ ```bash
132
+ embedflow plan --config ./embedflow.yaml --queries ./probe_queries.jsonl
133
+ embedflow plan --config ./embedflow.yaml --queries ./probe_queries.jsonl --format json --output migration-plan.json
134
+ ```
135
+
136
+ The planner reports preflight, registry evidence, frozen T2-v1 finite-tail
137
+ behavior, candidate K, cache/economics projections, and staged rollout
138
+ guidance. It never routes traffic or mutates the source index; `SAFE` is not a
139
+ qrels-based retrieval-quality guarantee. See the [planner guide](https://github.com/arnsri33/embedflow/blob/main/docs/planner.md).
140
+
141
+ ### Source-authoritative Shadow Mode
142
+
143
+ Run the target migration path beside sampled traffic while always returning the
144
+ source result:
145
+
146
+ ```yaml
147
+ runtime: {mode: shadow}
148
+ shadow: {enabled: true, sample_rate: 0.10, candidate_k: 100, materialize: true}
149
+ ```
150
+
151
+ ```bash
152
+ embedflow serve --config ./embedflow.yaml
153
+ embedflow shadow report --config ./embedflow.yaml --since 24h --format json
154
+ ```
155
+
156
+ Shadow work is asynchronous, bounded, privacy-conscious, and failure isolated.
157
+ Reports contain operational and ranking-disagreement diagnostics—not qrel-based
158
+ quality guarantees—and recommendations never route canary traffic. See the
159
+ [Shadow Mode guide](https://github.com/arnsri33/embedflow/blob/main/docs/shadow-mode.md).
160
+
126
161
  ## Research
127
162
 
128
163
  For candidate depth `K`, EmbedFlow measures:
@@ -163,6 +198,7 @@ See the [Weaviate guide](https://github.com/arnsri33/embedflow/blob/main/docs/in
163
198
  ```bash
164
199
  embedflow --help
165
200
  embedflow analyze --help
201
+ embedflow plan --help
166
202
  embedflow serve --config ./embedflow.yaml
167
203
  embedflow status --config ./embedflow.yaml
168
204
  embedflow registry list
@@ -176,7 +212,7 @@ cover the remaining commands and endpoints.
176
212
 
177
213
  ## Status
178
214
 
179
- EmbedFlow v0.5.0 is an alpha release for research and early real-world
215
+ EmbedFlow v0.7.0 is a pre-1.0 release for research and early real-world
180
216
  testing. T2-v1 is an empirical finite-tail diagnostic, partial rankings can
181
217
  differ from fully warm target reranking, and ANN fidelity needs a reference
182
218
  comparison to audit.
@@ -21,6 +21,7 @@ Interactive OpenAPI documentation is available at
21
21
  | GET | `/metrics` | Aggregated latency and queue metrics |
22
22
  | GET | `/plan` | Current migration plan |
23
23
  | GET | `/economics` | Configured economics estimate |
24
+ | GET | `/shadow/report` | Bounded Shadow Mode diagnostics |
24
25
 
25
26
  ## Search
26
27
 
@@ -48,6 +49,18 @@ The response includes the result list and migration fields such as:
48
49
  request. A partial response scores the available target vectors; it can differ
49
50
  from the fully warm ranking.
50
51
 
52
+ When `runtime.mode=shadow`, `/search` returns the source-only result and marks
53
+ `migration.source_authoritative=true`; target work is scheduled in the
54
+ background. `/shadow/report`, `/status`, and `/metrics` expose aggregate shadow
55
+ observations without raw query/document text or vectors. Shadow failures and
56
+ timeouts are isolated from the response. Ranking overlap is diagnostic and is
57
+ not a qrels-based quality claim.
58
+
59
+ The advisory migration planner is exposed through the Python API and the
60
+ `embedflow plan` CLI. It is intentionally not a synchronous FastAPI endpoint:
61
+ probe analysis may load models and perform bounded candidate work, so operators
62
+ should generate a plan artifact out of band and publish/read it as needed.
63
+
51
64
  `/status` includes safe backend metadata. For pgvector this names the
52
65
  schema/table and vector contract; for Pinecone it names the host/index,
53
66
  namespace, dimension, metric, and safe vector counts. Neither backend returns
@@ -9,6 +9,8 @@ list. The commands below are the main public entry points.
9
9
  embedflow init --config ./embedflow.yaml
10
10
  embedflow analyze --config ./embedflow.yaml --output-dir ./analysis
11
11
  embedflow evaluate --config ./experiment.yaml --output-dir ./results
12
+ embedflow plan --config ./embedflow.yaml --queries ./probe_queries.jsonl
13
+ embedflow shadow report --config ./embedflow.yaml --since 24h
12
14
  ```
13
15
 
14
16
  `analyze` is the no-target-index workflow. It uses probe queries and frozen
@@ -16,6 +18,15 @@ T2-v1 logic. `evaluate` is the labelled workflow; it computes source quality,
16
18
  native target quality, target-within-source-candidates quality, candidate gap,
17
19
  containment, and migration depth when the required inputs are available.
18
20
 
21
+ `plan` is the advisory migration planner. It runs source preflight and registry
22
+ matching, optionally samples a JSONL probe set, and reports a candidate-depth
23
+ recommendation, cache/economics projections, warnings, and staged rollout
24
+ guidance. Use `--format json` (or `yaml`) for a structured artifact. A plan
25
+ with no probes reports `T2-v1: NOT_RUN`; `SAFE` is a finite-tail diagnostic,
26
+ not a qrels-based retrieval-quality guarantee. A `DEFER`/`EXPAND_PROBE` result
27
+ is still a successful analytical command and exits zero; invalid configuration
28
+ or failed preflight exits nonzero.
29
+
19
30
  ## Serving and operations
20
31
 
21
32
  ```bash
@@ -32,6 +43,12 @@ materialization throughput. `audit-index` checks the source index against an
32
43
  exact/reference configuration where supported. `prewarm` schedules target
33
44
  document work; it does not change source-index results.
34
45
 
46
+ Use `embedflow serve --mode shadow` to opt into source-authoritative Shadow
47
+ Mode for one process. `embedflow shadow report` reads the bounded telemetry
48
+ store and supports `--format text|json|yaml`, `--since`, `--output`, and
49
+ `--quiet`. A valid `DEFER`/`EXPAND_K` report is still an analytical success;
50
+ only invalid configuration or initialization returns a non-zero exit code.
51
+
35
52
  ## Registry and profiles
36
53
 
37
54
  ```bash
@@ -58,9 +58,57 @@ economics:
58
58
  gpu_price_per_hour: null
59
59
  target_docs_per_second: null
60
60
 
61
+ # Optional advisory planner settings. CLI flags override these values.
62
+ planner:
63
+ max_probes: 250
64
+ seed: 42
65
+ k_grid: [20, 50, 100, 200, 500]
66
+ max_candidates: null
67
+ max_target_encodes: null
68
+ max_sync_misses: null
69
+ background_batch_size: null
70
+ gpu_hourly_cost: null
71
+ target_docs_per_second: null
72
+ queries_per_second: null
73
+ daily_queries: null
74
+ cache_hit_rate: null
75
+ latency_budget_ms: null
76
+ access_trace: null
77
+ corpus_name: null
78
+ corpus_fingerprint: null
79
+
80
+ # Source-authoritative observation mode. ``runtime.mode: migration`` (the
81
+ # default) leaves Shadow Mode inactive. ``mode: shadow`` is an explicit opt-in.
82
+ runtime:
83
+ mode: migration # migration, normal, source, or shadow
84
+
85
+ shadow:
86
+ enabled: true
87
+ sample_rate: 0.10
88
+ sample_seed: 42
89
+ candidate_k: 100
90
+ materialize: true
91
+ max_inflight: 32
92
+ queue_capacity: 1000
93
+ timeout_ms: 10000
94
+ shutdown_grace_ms: 1000
95
+ telemetry:
96
+ enabled: true
97
+ path: ./.embedflow/shadow.sqlite3
98
+ retain_query_records: false
99
+ retain_query_text: false
100
+ max_records: 10000
101
+ retention_days: null
102
+ report_k: 10
103
+ min_target_coverage_for_ranking: 1.0
104
+
61
105
  state_path: ./embedflow_state.json
62
106
  ```
63
107
 
108
+ Shadow Mode always returns the source-authoritative result before target work
109
+ finishes. See [`shadow-mode.md`](shadow-mode.md) for queue, timeout, privacy,
110
+ materialization, and report semantics.
111
+
64
112
  ## Model contracts
65
113
 
66
114
  Model configuration can include `revision`, `dimension`, `max_length`,
@@ -130,6 +178,39 @@ EMBEDFLOW_CANDIDATE_DEPTH
130
178
  EMBEDFLOW_MAX_SYNC_MISSES
131
179
  EMBEDFLOW_BACKGROUND_BATCH_SIZE
132
180
  EMBEDFLOW_PROBE_KMAX
181
+ EMBEDFLOW_PLANNER_MAX_PROBES
182
+ EMBEDFLOW_PLANNER_SEED
183
+ EMBEDFLOW_PLANNER_MAX_CANDIDATES
184
+ EMBEDFLOW_PLANNER_MAX_TARGET_ENCODINGS
185
+ EMBEDFLOW_PLANNER_MAX_SYNC_MISSES
186
+ EMBEDFLOW_PLANNER_BACKGROUND_BATCH_SIZE
187
+ EMBEDFLOW_PLANNER_GPU_HOURLY_COST
188
+ EMBEDFLOW_PLANNER_TARGET_DOCS_PER_SECOND
189
+ EMBEDFLOW_PLANNER_QPS
190
+ EMBEDFLOW_PLANNER_DAILY_QUERIES
191
+ EMBEDFLOW_PLANNER_CACHE_HIT_RATE
192
+ EMBEDFLOW_PLANNER_LATENCY_BUDGET_MS
193
+ EMBEDFLOW_PLANNER_ACCESS_TRACE
194
+ EMBEDFLOW_PLANNER_CORPUS_NAME
195
+ EMBEDFLOW_PLANNER_CORPUS_FINGERPRINT
196
+ EMBEDFLOW_RUNTIME_MODE
197
+ EMBEDFLOW_SHADOW_ENABLED
198
+ EMBEDFLOW_SHADOW_SAMPLE_RATE
199
+ EMBEDFLOW_SHADOW_SAMPLE_SEED
200
+ EMBEDFLOW_SHADOW_CANDIDATE_K
201
+ EMBEDFLOW_SHADOW_MATERIALIZE
202
+ EMBEDFLOW_SHADOW_MAX_INFLIGHT
203
+ EMBEDFLOW_SHADOW_QUEUE_CAPACITY
204
+ EMBEDFLOW_SHADOW_TIMEOUT_MS
205
+ EMBEDFLOW_SHADOW_SHUTDOWN_GRACE_MS
206
+ EMBEDFLOW_SHADOW_TELEMETRY_ENABLED
207
+ EMBEDFLOW_SHADOW_TELEMETRY_PATH
208
+ EMBEDFLOW_SHADOW_RETAIN_QUERY_RECORDS
209
+ EMBEDFLOW_SHADOW_RETAIN_QUERY_TEXT
210
+ EMBEDFLOW_SHADOW_MAX_RECORDS
211
+ EMBEDFLOW_SHADOW_REPORT_K
212
+ EMBEDFLOW_SHADOW_RETENTION_DAYS
213
+ EMBEDFLOW_SHADOW_MIN_TARGET_COVERAGE
133
214
  ```
134
215
 
135
216
  ## Input files
@@ -1,6 +1,6 @@
1
1
  # Limitations and release scope
2
2
 
3
- EmbedFlow v0.5.0 is an alpha release for research and early real-world
3
+ EmbedFlow v0.7.0 is a pre-1.0 release for research and early real-world
4
4
  testing. The serving path is designed to make migration experiments concrete;
5
5
  production rollout still requires application-specific validation.
6
6
 
@@ -0,0 +1,111 @@
1
+ # Migration planner
2
+
3
+ `embedflow plan` is a bounded, advisory analysis for moving from an existing
4
+ source vector index to a target embedding model. It composes the existing
5
+ backend audit, model contracts, retained registry evidence, probe retrieval,
6
+ frozen T2-v1 diagnostic, cache behavior, and economics helpers. It never
7
+ routes traffic and does not create, update, or delete source-index data.
8
+
9
+ ## Quickstart
10
+
11
+ Install EmbedFlow with the optional dependency for the configured backend, then
12
+ provide representative query probes as JSONL:
13
+
14
+ ```json
15
+ {"id":"q1","query":"how do I reset my password?"}
16
+ {"id":"q2","query":"what is the vacation policy?"}
17
+ ```
18
+
19
+ Run a terminal plan or save a machine-readable artifact:
20
+
21
+ ```bash
22
+ embedflow plan --config embedflow.yaml --queries probes.jsonl
23
+ embedflow plan --config embedflow.yaml --queries probes.jsonl \
24
+ --k-grid 20,50,100,200,500 --format json --output migration-plan.json
25
+ ```
26
+
27
+ For a deterministic, network-free walkthrough, run `examples/planner/run_demo.sh`.
28
+ The demo uses the same planner API and tiny hash encoders; it is not a quality
29
+ benchmark.
30
+
31
+ ## What the result means
32
+
33
+ The result is a versioned (`schema_version: 1`) object with source/target
34
+ contracts, preflight checks, evidence, candidate-depth diagnostics, cache and
35
+ rollout guidance, economics, and structured warnings. The Python API is:
36
+
37
+ ```python
38
+ from embedflow.planner import MigrationPlanner
39
+
40
+ result = MigrationPlanner("embedflow.yaml").plan("probes.jsonl")
41
+ artifact = result.to_dict()
42
+ ```
43
+
44
+ Recommendation states are deliberately small:
45
+
46
+ - `PROCEED`: probe evidence and preflight are strong enough for an operator-controlled shadow/canary sequence.
47
+ - `PROCEED_WITH_CAUTION`: probes are acceptable, but evidence coverage or contract/ANN provenance is limited.
48
+ - `EXPAND_PROBE`: T2-v1 asks for a deeper pool or the evidence is too small.
49
+ - `DEFER`: finite-tail behavior is `UNSAFE_OR_UNCERTAIN`; do not canary based on this plan.
50
+ - `BLOCKED`: structural/backend preflight failed; no candidate K is certified.
51
+
52
+ These are recommendations, not automatic deployment actions. The suggested
53
+ rollout is Validate → Shadow → small canary → expanded canary → target-primary
54
+ path, with the source remaining authoritative until the operator approves each
55
+ step. Stop conditions include source/target errors, p95 latency, queue depth,
56
+ cache hit rate, and any native/qrels evaluation regressions.
57
+
58
+ ## Evidence and T2-v1
59
+
60
+ Registry matching preserves the existing hierarchy: `EXACT REGISTRY MATCH`,
61
+ `PRIOR EVIDENCE AVAILABLE` (same contracts, different/unknown corpus),
62
+ `RELATED EVIDENCE ONLY`, and `NO REGISTRY MATCH`. Prior or related evidence can
63
+ seed context but cannot override current probe results or certify a new corpus.
64
+
65
+ The planner calls the existing frozen T2-v1 implementation unchanged. T2-v1 is
66
+ a no-qrels finite-tail diagnostic of candidate behavior. `SAFE` does **not**
67
+ prove a candidate gap, nDCG, recall, or zero quality loss. Without actual qrels
68
+ or native-target evaluation, `candidate_gap` is `null`/`UNKNOWN`. ANN fidelity
69
+ is also `UNKNOWN` unless an exact/reference comparison is supplied by the
70
+ backend audit workflow.
71
+
72
+ No probes is a valid preflight/economics mode, but T2 is `NOT_RUN`, confidence
73
+ is `INSUFFICIENT`, and no K is presented as safe. Probe JSONL rows may use
74
+ `query` (or the existing `text` alias) and an optional `id`; empty, malformed,
75
+ or duplicate IDs are rejected. `--max-probes` and `--seed` provide deterministic
76
+ bounded sampling. Candidate documents are deduplicated by canonical ID and
77
+ encoded in bounded batches; `--max-target-encodes` can impose a hard cap.
78
+
79
+ ## Cache, performance, and economics
80
+
81
+ Planning uses an ephemeral target-cache/state directory when it opens a runtime,
82
+ so the normal serving cache is not polluted. It recommends conservative sync
83
+ miss and background batch settings and defaults to traffic-driven progressive
84
+ warming when no access trace is supplied. An access trace can be supplied with
85
+ `--access-trace` as rows such as `{"document_id":"abc","count":192}` to
86
+ model hot-document coverage.
87
+
88
+ Performance values retain provenance (`measured`, `user_supplied`, `modeled`,
89
+ `registry`, or `unknown`). `--profile` performs only a small local encode/search
90
+ sample with one unmeasured warmup pass excluded; it is diagnostic and is not a
91
+ formal benchmark. Economics accepts `--target-docs-per-second` and
92
+ `--gpu-hourly-cost`; absent inputs remain `UNKNOWN`. Raw vector storage is
93
+ `documents × target dimension × dtype bytes` and excludes ANN overhead,
94
+ metadata, replicas, backups, and database overhead. Progressive percentages
95
+ are materialization scenarios, not claims that a percentage is sufficient.
96
+
97
+ ## Configuration and privacy
98
+
99
+ An optional `planner:` section mirrors the CLI options (`max_probes`, `seed`,
100
+ `k_grid`, work caps, cache hints, throughput/cost inputs, and corpus identity).
101
+ CLI values take precedence. `EMBEDFLOW_PLANNER_*` environment overrides follow
102
+ the normal config mechanism. Probe text is never copied into plan artifacts or
103
+ telemetry by default; only counts, IDs used for diagnostics, and aggregates are
104
+ reported. Backend credentials are handled by the existing backend adapters and
105
+ are redacted from errors and structured output.
106
+
107
+ The planner does not provide qrel evaluation, traffic routing, rollback
108
+ automation, distributed scheduling, ANN tuning, integrated-vectorizer contract
109
+ verification, or production cost guarantees. Use `embedflow evaluate` with
110
+ qrels/native target rankings when empirical retrieval-quality claims are
111
+ required.
@@ -18,8 +18,8 @@ python -m twine check dist/*
18
18
  Inspect both archives before uploading:
19
19
 
20
20
  ```bash
21
- unzip -l dist/embedflow-0.5.0-py3-none-any.whl
22
- tar -tzf dist/embedflow-0.5.0.tar.gz
21
+ unzip -l dist/embedflow-0.7.0-py3-none-any.whl
22
+ tar -tzf dist/embedflow-0.7.0.tar.gz
23
23
  sha256sum dist/*
24
24
  ```
25
25
 
@@ -32,7 +32,7 @@ Test the wheel outside the source tree:
32
32
  ```bash
33
33
  python -m venv /tmp/embedflow-wheel-test
34
34
  /tmp/embedflow-wheel-test/bin/python -m pip install --upgrade pip
35
- /tmp/embedflow-wheel-test/bin/python -m pip install dist/embedflow-0.5.0-py3-none-any.whl
35
+ /tmp/embedflow-wheel-test/bin/python -m pip install dist/embedflow-0.7.0-py3-none-any.whl
36
36
  cd /tmp
37
37
  /tmp/embedflow-wheel-test/bin/python -c "import embedflow; print(embedflow.__version__)"
38
38
  /tmp/embedflow-wheel-test/bin/embedflow --help
@@ -65,7 +65,7 @@ python -m venv /tmp/embedflow-testpypi
65
65
  /tmp/embedflow-testpypi/bin/python -m pip install \
66
66
  --index-url https://test.pypi.org/simple/ \
67
67
  --extra-index-url https://pypi.org/simple/ \
68
- embedflow==0.5.0
68
+ embedflow==0.7.0
69
69
  cd /tmp
70
70
  /tmp/embedflow-testpypi/bin/python -c "import embedflow; print(embedflow.__version__)"
71
71
  /tmp/embedflow-testpypi/bin/embedflow --help
@@ -79,11 +79,11 @@ Test optional integrations in a second clean environment:
79
79
  /tmp/embedflow-testpypi/bin/python -m pip install \
80
80
  --index-url https://test.pypi.org/simple/ \
81
81
  --extra-index-url https://pypi.org/simple/ \
82
- "embedflow[faiss,dashboard,pinecone,milvus,weaviate]==0.5.0"
82
+ "embedflow[faiss,dashboard,pinecone,milvus,weaviate]==0.7.0"
83
83
  ```
84
84
 
85
- If the same filename already exists on TestPyPI, use a pre-release such as
86
- `0.5.0rc1` for the TestPyPI-only trial. Keep production `0.5.0` unchanged.
85
+ If the same filename already exists on TestPyPI, use a pre-release for the
86
+ TestPyPI-only trial. Never overwrite the production version.
87
87
 
88
88
  ## Trusted Publishing configuration
89
89
 
@@ -115,7 +115,7 @@ above keeps the test step explicit.
115
115
  3. Run the final release gate and review the generated report.
116
116
  4. Configure the PyPI pending publisher and protected `pypi` environment.
117
117
  5. Create a Git tag and GitHub Release for the exact package version, for
118
- example `v0.5.0`.
118
+ example `v0.7.0`.
119
119
  6. Approve the `pypi` environment when the release workflow is ready.
120
120
  7. Verify the files and metadata on PyPI.
121
121
  8. Install from production PyPI in a directory outside this checkout.
@@ -0,0 +1,104 @@
1
+ # Shadow Mode
2
+
3
+ Shadow Mode runs the configured target migration path beside real traffic while
4
+ the existing source result remains authoritative. It is an observation and
5
+ cache-warming tool, not a traffic router:
6
+
7
+ ```yaml
8
+ runtime:
9
+ mode: shadow
10
+
11
+ shadow:
12
+ enabled: true
13
+ sample_rate: 0.10
14
+ sample_seed: 42
15
+ candidate_k: 100
16
+ materialize: true
17
+ max_inflight: 32
18
+ queue_capacity: 1000
19
+ timeout_ms: 10000
20
+ telemetry:
21
+ enabled: true
22
+ path: ./.embedflow/shadow.sqlite3
23
+ retain_query_records: false
24
+ retain_query_text: false
25
+ ```
26
+
27
+ Install the optional backend/model extras required by the selected
28
+ configuration, then start the ordinary service:
29
+
30
+ ```bash
31
+ pip install "embedflow[dashboard]"
32
+ embedflow serve --config embedflow.yaml
33
+ # or override the file for one process:
34
+ embedflow serve --config embedflow.yaml --mode shadow
35
+ ```
36
+
37
+ The request path performs source encoding, source ANN retrieval, and document
38
+ resolution exactly as source-only serving. It returns that result without
39
+ waiting for target encoding, cache misses, reranking, telemetry, or the
40
+ materialization worker. Shadow jobs use a bounded queue and worker pool. A full
41
+ queue drops only shadow work; a target/model/cache/telemetry failure or timeout
42
+ is recorded and cannot change the primary response.
43
+
44
+ If the target model cannot be loaded during startup, source/shadow serving
45
+ still opens with target work marked unavailable; sampled requests record the
46
+ isolated target-encoding failure. Normal migration mode remains fail-fast for
47
+ the same startup error.
48
+
49
+ `candidate_k` is explicit. The planner may suggest a value, but Shadow Mode
50
+ does not silently select one. `materialize: false` reads already cached target
51
+ vectors without changing cache or queue state. With `materialize: true`, cold
52
+ candidate IDs are deduplicated in the existing persistent materialization queue
53
+ and warmed asynchronously; synchronous target misses remain zero in Shadow
54
+ Mode. A comparison with incomplete target vectors is reported as `partial`,
55
+ with target coverage shown separately.
56
+
57
+ ## Reports
58
+
59
+ Reports use the privacy-preserving SQLite telemetry store (by default beside
60
+ the configured cache):
61
+
62
+ ```bash
63
+ embedflow shadow report --config embedflow.yaml --since 24h
64
+ embedflow shadow report --config embedflow.yaml --since 24h --format json --output shadow-report.json
65
+ ```
66
+
67
+ Use `/shadow/report` or the `shadow` section of `/status` and `/metrics` in the
68
+ FastAPI service. JSON/YAML stdout contains only the requested artifact; progress
69
+ and errors belong on stderr. `--since` accepts seconds or `s`, `m`, `h`, `d`,
70
+ and `w` suffixes.
71
+
72
+ Reports contain sampled/completed/partial/failed/timed-out/dropped counts,
73
+ cache and materialization accounting, source and shadow latency distributions,
74
+ target coverage, top-1 agreement, and top-k overlap. Shadow latency is not
75
+ user-facing latency because it is off the primary critical path. Ranking
76
+ disagreement and overlap are diagnostics, not recall, nDCG, or quality loss;
77
+ without qrels/native-target evaluation no quality guarantee is made.
78
+
79
+ T2-v1 remains the frozen window-level diagnostic. EmbedFlow does not label an
80
+ individual production query `T2 SAFE`. A report recommendation is operational
81
+ guidance only:
82
+
83
+ - `CONTINUE_SHADOW` means evidence or coverage is still limited.
84
+ - `EXPAND_K` means the accumulated T2 window requests a larger candidate pool.
85
+ - `INVESTIGATE` means failures, timeouts, or uncertain T2 behavior need review.
86
+ - `READY_FOR_CANARY_EVALUATION` means an operator may consider a separately
87
+ designed canary; it never routes traffic automatically.
88
+
89
+ Raw query text, candidate text, source vectors, target vectors, and credentials
90
+ are not persisted in telemetry. Query IDs may be retained only when explicitly
91
+ enabled, and retention is bounded by `max_records`. The telemetry database is
92
+ segmented by source/target contract and candidate-K fingerprint so unrelated
93
+ migrations are not mixed.
94
+
95
+ On shutdown EmbedFlow stops accepting new shadow work, drops queued jobs when
96
+ necessary, and waits only the configured bounded grace period. A corrupted or
97
+ unwritable telemetry file degrades observability; it does not take source
98
+ serving down. Source indexes are read-only during Shadow Mode. The planner and
99
+ Shadow Mode are separate: generate a plan first, then copy its recommended K
100
+ explicitly into a reviewed Shadow configuration.
101
+
102
+ Shadow Mode is advisory and experimental operational instrumentation. It does
103
+ not provide autonomous rollout, rollback, qrel evaluation, or ANN-fidelity
104
+ proof.