rag-connector 0.2.0a1__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (145) hide show
  1. rag_connector-0.2.0a1/AGENTS.md +63 -0
  2. rag_connector-0.2.0a1/CHANGELOG.md +27 -0
  3. rag_connector-0.2.0a1/CONTRIBUTING.md +20 -0
  4. rag_connector-0.2.0a1/LICENSE +21 -0
  5. rag_connector-0.2.0a1/MANIFEST.in +35 -0
  6. rag_connector-0.2.0a1/PKG-INFO +272 -0
  7. rag_connector-0.2.0a1/README.md +235 -0
  8. rag_connector-0.2.0a1/docs/connector-author-guide.md +280 -0
  9. rag_connector-0.2.0a1/docs/contract.md +540 -0
  10. rag_connector-0.2.0a1/examples/connector_template.py +193 -0
  11. rag_connector-0.2.0a1/pyproject.toml +132 -0
  12. rag_connector-0.2.0a1/setup.cfg +4 -0
  13. rag_connector-0.2.0a1/src/rag_connector/CONNECTOR-INFO.md +45 -0
  14. rag_connector-0.2.0a1/src/rag_connector/__init__.py +162 -0
  15. rag_connector-0.2.0a1/src/rag_connector/base.py +625 -0
  16. rag_connector-0.2.0a1/src/rag_connector/capabilities.py +231 -0
  17. rag_connector-0.2.0a1/src/rag_connector/cli.py +104 -0
  18. rag_connector-0.2.0a1/src/rag_connector/corpus/README.md +34 -0
  19. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/account_companion_benefits.txt +11 -0
  20. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/account_corporate_travel.txt +11 -0
  21. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/account_payment_methods.txt +13 -0
  22. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/age_requirements.md +30 -0
  23. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/baggage_cameras_and_heritage.txt +11 -0
  24. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/baggage_equipment_rentals.txt +13 -0
  25. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/baggage_mass_allowance.md +25 -0
  26. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/baggage_personal_medication.txt +9 -0
  27. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/baggage_prohibited_items.txt +15 -0
  28. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/booking_cancellation_policy.md +178 -0
  29. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/booking_children_and_families.txt +11 -0
  30. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/booking_deposits_and_payment.txt +11 -0
  31. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/booking_group_reservations.txt +11 -0
  32. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/booking_how_it_works.txt +9 -0
  33. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/booking_upgrades_and_changes.txt +11 -0
  34. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/cabin_classes_comparison.md +26 -0
  35. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/cabin_classes_overview.txt +9 -0
  36. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/cancellation_group_bookings.txt +11 -0
  37. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/cancellation_launch_scrub.txt +11 -0
  38. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/cancellation_policy_standard.txt +9 -0
  39. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/conditions_disqualifying.txt +21 -0
  40. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/departure_scheduling_guide.md +106 -0
  41. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/dest_ceres_leap.txt +9 -0
  42. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/dest_ceres_occator_crater.txt +9 -0
  43. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/dest_ceres_overview.txt +9 -0
  44. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/dest_ceres_surface_walk.txt +7 -0
  45. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/dest_enceladus_geyser.txt +9 -0
  46. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/dest_guide_ceres.md +114 -0
  47. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/dest_guide_mars.md +119 -0
  48. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/dest_guide_moon.md +102 -0
  49. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/dest_guide_outer_system.md +92 -0
  50. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/dest_guide_venus.md +124 -0
  51. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/dest_mars_colony1_overview.txt +9 -0
  52. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/dest_mars_cydonia_formations.txt +9 -0
  53. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/dest_mars_hellas_basin.txt +9 -0
  54. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/dest_mars_olympus_mons.txt +7 -0
  55. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/dest_mars_terraforming_tour.txt +9 -0
  56. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/dest_moon_base7_overview.txt +9 -0
  57. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/dest_moon_golf_links.txt +9 -0
  58. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/dest_moon_rover_tour.txt +9 -0
  59. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/dest_moon_tranquility_heritage.txt +9 -0
  60. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/dest_saturn_ring_transit.txt +9 -0
  61. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/dest_titan_ice_flats.txt +9 -0
  62. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/dest_venus_aphrodite_overview.txt +9 -0
  63. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/dest_venus_glass_floor_tour.txt +9 -0
  64. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/dest_venus_gondola_descent.txt +9 -0
  65. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/dest_venus_pool_and_sun_deck.txt +9 -0
  66. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/dest_venus_science_centre.txt +7 -0
  67. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/fitness_training_requirements.txt +7 -0
  68. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/flyer_glass_floor_venus.md +25 -0
  69. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/flyer_gondola_descent_venus.md +30 -0
  70. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/flyer_olympus_mons_summit.md +25 -0
  71. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/flyer_the_leap_ceres.md +27 -0
  72. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/flyer_tranquility_heritage_walk.md +34 -0
  73. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/g_force_tolerance.txt +9 -0
  74. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/insurance_exclusions.txt +13 -0
  75. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/insurance_extraction_coverage.txt +11 -0
  76. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/insurance_what_is_covered.txt +13 -0
  77. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/itinerary_grand_solar_tour.md +39 -0
  78. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/itinerary_gst_guide.md +247 -0
  79. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/itinerary_inner_planets.md +31 -0
  80. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/itinerary_ips_guide.md +152 -0
  81. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/itinerary_ort_guide.md +179 -0
  82. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/itinerary_outer_reaches.md +40 -0
  83. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/loyalty_orbital_points.txt +9 -0
  84. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/loyalty_programme_guide.md +132 -0
  85. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/loyalty_redemption.txt +11 -0
  86. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/loyalty_tiers.md +31 -0
  87. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/medical_clearance_by_destination.md +28 -0
  88. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/medical_clearance_overview.txt +9 -0
  89. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/medical_events_in_transit.txt +11 -0
  90. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/medical_fitness_guide.md +237 -0
  91. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/medications_and_spaceflight.txt +9 -0
  92. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/onboard_communications.txt +13 -0
  93. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/onboard_dining_and_nutrition.txt +13 -0
  94. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/onboard_entertainment.txt +11 -0
  95. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/onboard_eva_suit_orientation.txt +9 -0
  96. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/onboard_medical_bay.txt +9 -0
  97. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/onboard_transit_times.md +31 -0
  98. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/onboard_zero_g_adaptation.txt +9 -0
  99. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/packing_guide.md +121 -0
  100. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/passenger_handbook.md +189 -0
  101. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/pre_departure_health_check.txt +9 -0
  102. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/pregnancy_policy.txt +9 -0
  103. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/rebooking_after_scrub.txt +11 -0
  104. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/refund_medical_disqualification.txt +9 -0
  105. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/refund_timing_by_method.md +19 -0
  106. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/scheduling_check_in_requirements.md +25 -0
  107. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/scheduling_departure_windows.txt +9 -0
  108. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/scheduling_mining_transport.txt +9 -0
  109. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/scheduling_missed_departure.txt +13 -0
  110. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space/welcome_aboard_guide.md +150 -0
  111. rag_connector-0.2.0a1/src/rag_connector/corpus/pelorus_space.chunks.jsonl +605 -0
  112. rag_connector-0.2.0a1/src/rag_connector/document_loader.py +14 -0
  113. rag_connector-0.2.0a1/src/rag_connector/embedding.py +107 -0
  114. rag_connector-0.2.0a1/src/rag_connector/errors.py +32 -0
  115. rag_connector-0.2.0a1/src/rag_connector/fingerprints.py +102 -0
  116. rag_connector-0.2.0a1/src/rag_connector/ingest.py +518 -0
  117. rag_connector-0.2.0a1/src/rag_connector/models.py +164 -0
  118. rag_connector-0.2.0a1/src/rag_connector/prompts.py +934 -0
  119. rag_connector-0.2.0a1/src/rag_connector/py.typed +1 -0
  120. rag_connector-0.2.0a1/src/rag_connector/reference.py +1073 -0
  121. rag_connector-0.2.0a1/src/rag_connector/reference_provision.py +336 -0
  122. rag_connector-0.2.0a1/src/rag_connector/registry.py +192 -0
  123. rag_connector-0.2.0a1/src/rag_connector/run_snapshot.py +78 -0
  124. rag_connector-0.2.0a1/src/rag_connector/validate.py +1589 -0
  125. rag_connector-0.2.0a1/src/rag_connector.egg-info/PKG-INFO +272 -0
  126. rag_connector-0.2.0a1/src/rag_connector.egg-info/SOURCES.txt +143 -0
  127. rag_connector-0.2.0a1/src/rag_connector.egg-info/dependency_links.txt +1 -0
  128. rag_connector-0.2.0a1/src/rag_connector.egg-info/entry_points.txt +2 -0
  129. rag_connector-0.2.0a1/src/rag_connector.egg-info/requires.txt +12 -0
  130. rag_connector-0.2.0a1/src/rag_connector.egg-info/top_level.txt +1 -0
  131. rag_connector-0.2.0a1/tests/conftest.py +65 -0
  132. rag_connector-0.2.0a1/tests/test_base_scores.py +51 -0
  133. rag_connector-0.2.0a1/tests/test_cli.py +120 -0
  134. rag_connector-0.2.0a1/tests/test_contract.py +314 -0
  135. rag_connector-0.2.0a1/tests/test_dataset_listing.py +202 -0
  136. rag_connector-0.2.0a1/tests/test_document_loader.py +144 -0
  137. rag_connector-0.2.0a1/tests/test_embedding.py +65 -0
  138. rag_connector-0.2.0a1/tests/test_ingest.py +503 -0
  139. rag_connector-0.2.0a1/tests/test_prompt_templates.py +800 -0
  140. rag_connector-0.2.0a1/tests/test_reference.py +616 -0
  141. rag_connector-0.2.0a1/tests/test_reference_durability.py +439 -0
  142. rag_connector-0.2.0a1/tests/test_registry.py +272 -0
  143. rag_connector-0.2.0a1/tests/test_retrieval_mode.py +78 -0
  144. rag_connector-0.2.0a1/tests/test_run_snapshot.py +36 -0
  145. rag_connector-0.2.0a1/tests/test_validate_connector.py +663 -0
@@ -0,0 +1,63 @@
1
+ # Agent guidance
2
+
3
+ RAG Connector is the product-neutral boundary for treating a RAG pipeline as a
4
+ black box. Do not add RAGauge evaluation policy or Pelorus curation behavior to
5
+ the core contract.
6
+
7
+ ## Contract priorities
8
+
9
+ 1. Evidence integrity outranks convenience. Every declaration a connector
10
+ makes is read as truth by measurement code downstream, so a value that
11
+ is guessed, synthesized, or quietly defaulted is worse than an absent
12
+ one.
13
+ 2. Stable chunk IDs must agree across retrieval and corpus-reading paths.
14
+ 3. A backend error is not an empty retrieval result.
15
+ 4. Canonical scores are higher-is-better; preserve backend-native scores.
16
+ 5. Retrieval mode controls whether scores and `top_k` have metric meaning.
17
+ 6. Fingerprints identify separate domains: connector configuration, corpus,
18
+ embedding space, and retrieval policy.
19
+ 7. Optional capabilities must remain independently optional.
20
+ 8. Persisted connection dictionaries must be JSON-serializable and secret-free.
21
+
22
+ ## Verification
23
+
24
+ ```bash
25
+ python -m pytest
26
+ python -m ruff check .
27
+ ```
28
+
29
+ Keep the core dependency-light. Connector-specific SDKs belong in separately
30
+ installed connector packages.
31
+
32
+ ## Packaging: purge `egg-info` before building a release
33
+
34
+ ```bash
35
+ rm -rf src/*.egg-info build dist && python -m build
36
+ ```
37
+
38
+ **A green suite says nothing about what ships.** Two things are only observable
39
+ from a built artifact, and every consumer editable-installs this package — an
40
+ editable install resolves to `src/` on disk, so it always sees `py.typed` and
41
+ never builds an sdist.
42
+
43
+ - `py.typed` must be in the **wheel**, or consumers silently lose PEP 561 type
44
+ information while the classifiers still advertise `Typing :: Typed`. It ships
45
+ because `[tool.setuptools.package-data]` declares it explicitly, not by
46
+ default.
47
+ - **Development notes must not be in the sdist.** Status write-ups, plans, and
48
+ roadmaps are written for whoever picks the work up next, not for someone
49
+ installing the package. Keep them untracked (there is a gitignored path for
50
+ exactly this), because a public repo publishes every tracked file whatever
51
+ `MANIFEST.in` says. What stays under `docs/` is reference a reader of the
52
+ package needs. Note what this cuts both ways on: `tests/` **does** ship, by an
53
+ explicit `MANIFEST.in` decision, so a comment or docstring in the suite is
54
+ published prose — hold it to the same standard as the docs.
55
+
56
+ ⚠ **`MANIFEST.in` changes are masked by a stale `SOURCES.txt`.** setuptools
57
+ reuses `src/rag_connector.egg-info/SOURCES.txt` when building an sdist, so
58
+ removing an entry from `MANIFEST.in` has no effect until that file is deleted.
59
+ This is not hypothetical: it hid the removal of an internal status file on the
60
+ first verified build (2026-07-28), and `egg-info/` is gitignored, so a dirty
61
+ working tree carries the stale list silently into a release. Verify a removal
62
+ by listing the built artifact, never by reading `MANIFEST.in`.
63
+
@@ -0,0 +1,27 @@
1
+ # Changelog
2
+
3
+ Notable changes to rag-connector. Dates are the day the work landed.
4
+
5
+ ## 0.2.0a1 — first public release
6
+
7
+ The first release published to PyPI, and the first from this public repository.
8
+ Development before it happened in a private repository whose history is not
9
+ carried here; its `v0.1.0a0` and `v0.2.0a0` tags were never published, and this
10
+ version is numbered after them so no version goes backwards.
11
+
12
+ What it contains:
13
+
14
+ - The black-box retrieval contract (`RagPipeline`, `ChunkRecord`,
15
+ `RetrievedChunk`), retrieval-mode and score semantics, and stable chunk
16
+ identity.
17
+ - Optional capabilities: corpus reading, query embedding, indexed chunk vectors,
18
+ generation, dataset listing, run snapshots, publication and telemetry.
19
+ - The connector registry (`rag_connector.connectors` entry points) with
20
+ reconstructable, secret-free connection documents.
21
+ - The conformance validator and reusable connector test kit
22
+ (`rag-connector validate`).
23
+ - The Reference RAG, a local FastEmbed + Chroma implementation (`reference`
24
+ extra).
25
+ - The ingest kit (`rag_connector.ingest`) and the bundled Pelorus Space demo
26
+ corpus: 92 documents and the frozen 605-chunk split that RAGauge and Pelorus
27
+ Query both cite.
@@ -0,0 +1,20 @@
1
+ # Contributing
2
+
3
+ RAG Connector is a small, product-neutral library that defines one thing: the
4
+ contract for treating a RAG pipeline as a black box. Its value is that the
5
+ contract stays narrow and that a connector written against it keeps working, so
6
+ changes are weighed against compatibility first — a redesign that is cleaner but
7
+ moves the surface is usually the wrong trade here.
8
+
9
+ Before opening a change:
10
+
11
+ ```bash
12
+ python -m pip install -e ".[dev,reference]"
13
+ python -m pytest
14
+ python -m ruff check .
15
+ ```
16
+
17
+ Connector-specific SDK dependencies do not belong in the core package. Add
18
+ concrete connectors through optional extras or independently installed packages.
19
+
20
+ All contributions are made under the MIT License.
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Stacey Farias
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,35 @@
1
+ # AGENTS.md and CONTRIBUTING.md ship: an sdist is a source distribution, and
2
+ # someone building or contributing from it wants the guidance.
3
+ #
4
+ # Development notes deliberately do NOT ship — status, plans, roadmaps, anything
5
+ # written for whoever picks the work up next rather than for someone installing
6
+ # the package. They are not tracked at all (see .gitignore), because a tracked
7
+ # file in a public repo is a published file whatever this file says. What
8
+ # remains under docs/ is reference a reader of the package needs, and it is
9
+ # included wholesale below on that basis.
10
+ include AGENTS.md
11
+ include CONTRIBUTING.md
12
+ include CHANGELOG.md
13
+ recursive-include docs *.md
14
+ recursive-include examples *.py
15
+ # The suite ships, in full and on purpose. Someone building from an sdist --
16
+ # a distro packager, a security reviewer, anyone who does not trust a wheel --
17
+ # runs the tests, and this package's whole claim is that a connector's behavior
18
+ # is checkable. Declaring it explicitly also repairs a half-shipped state:
19
+ # setuptools' own sdist default quietly picks up `tests/test*.py` and nothing
20
+ # else, so `conftest.py` was being left behind. That file is the autouse
21
+ # fixture that hides installed third-party connector entry points, so without
22
+ # it the shipped suite's result depends on whatever else is in site-packages --
23
+ # green here, red on someone else's machine, for reasons in neither tree.
24
+ recursive-include tests *.py
25
+ # The bundled corpora are package DATA read at runtime from inside the installed
26
+ # package, so they must be in the WHEEL too -- that is declared in pyproject's
27
+ # [tool.setuptools.package-data]. They are listed here as well so the sdist
28
+ # carries them; without a corpus, `pip install "rag-connector[reference]"`
29
+ # installs a Reference RAG with nothing to point at.
30
+ recursive-include src/rag_connector/corpus *.txt *.md *.pdf *.docx *.rtf
31
+ # The frozen pre-chunked form of each bundled corpus, read at runtime by
32
+ # rag_connector.ingest.load_bundled_chunks. It is the substrate Pelorus's
33
+ # pre-canned extracts and RAGauge's pre-canned dataset both cite, so an
34
+ # install without it can serve the documents but not the ids anyone stored.
35
+ recursive-include src/rag_connector/corpus *.chunks.jsonl
@@ -0,0 +1,272 @@
1
+ Metadata-Version: 2.4
2
+ Name: rag-connector
3
+ Version: 0.2.0a1
4
+ Summary: A product-neutral contract and conformance kit for treating RAG pipelines as black boxes.
5
+ Author: Stacey Farias
6
+ License-Expression: MIT
7
+ Project-URL: Homepage, https://github.com/staceyfarias/rag-connector
8
+ Project-URL: Repository, https://github.com/staceyfarias/rag-connector
9
+ Project-URL: Issues, https://github.com/staceyfarias/rag-connector/issues
10
+ Project-URL: Documentation, https://github.com/staceyfarias/rag-connector/blob/main/docs/contract.md
11
+ Keywords: rag,retrieval,retrieval-augmented-generation,connector,evaluation,conformance,chunking
12
+ Classifier: Development Status :: 3 - Alpha
13
+ Classifier: Intended Audience :: Developers
14
+ Classifier: Programming Language :: Python :: 3
15
+ Classifier: Programming Language :: Python :: 3.10
16
+ Classifier: Programming Language :: Python :: 3.11
17
+ Classifier: Programming Language :: Python :: 3.12
18
+ Classifier: Programming Language :: Python :: 3.13
19
+ Classifier: Operating System :: OS Independent
20
+ Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
21
+ Classifier: Topic :: Software Development :: Libraries :: Python Modules
22
+ Classifier: Typing :: Typed
23
+ Requires-Python: >=3.10
24
+ Description-Content-Type: text/markdown
25
+ License-File: LICENSE
26
+ Provides-Extra: reference
27
+ Requires-Dist: chromadb>=0.4.22; extra == "reference"
28
+ Requires-Dist: fastembed>=0.3.0; extra == "reference"
29
+ Requires-Dist: pypdf>=4; extra == "reference"
30
+ Requires-Dist: python-docx>=1.1; extra == "reference"
31
+ Requires-Dist: striprtf>=0.0.26; extra == "reference"
32
+ Provides-Extra: dev
33
+ Requires-Dist: pytest>=8; extra == "dev"
34
+ Requires-Dist: python-dotenv>=1; extra == "dev"
35
+ Requires-Dist: ruff>=0.9; extra == "dev"
36
+ Dynamic: license-file
37
+
38
+ # RAG Connector
39
+
40
+ RAG Connector is the shared contract that lets **RAGauge** and **Pelorus Query**
41
+ work with the same Retrieval-Augmented Generation pipeline as a black box —
42
+ without either of them knowing how that pipeline is built.
43
+
44
+ - **RAGauge** evaluates a RAG system: whether it retrieves the right evidence
45
+ and whether its answers are grounded in it, measured against a frozen corpus
46
+ and test set.
47
+ - **Pelorus Query** sits on top of a RAG system and serves curated,
48
+ source-grounded answers (Evidence Extracts), improving them as queries repeat.
49
+
50
+ Both need the same things from the pipeline underneath: ask it a question and
51
+ get ranked chunks back, read its corpus, know which embedding space and
52
+ retrieval mode produced a score, and cite a chunk by an id that means the same
53
+ thing to both. RAG Connector defines that once. A pipeline that implements the
54
+ contract — or is wrapped by a connector that does — can be measured by RAGauge
55
+ and curated by Pelorus Query, and a chunk one of them cites resolves to the
56
+ same text in the other. The bundled demo corpus and its frozen chunk set exist
57
+ for exactly that: both products are built against the same 605 chunk ids.
58
+
59
+ Interoperability between those two products is why this library exists, and it
60
+ is why the contract is narrow. Nothing here evaluates or curates: evaluation
61
+ policy belongs to RAGauge, curation behavior to Pelorus Query. Any other tool
62
+ that needs to treat a RAG pipeline as a black box can build on the same
63
+ contract. (RAGauge and Pelorus Query are not public yet.)
64
+
65
+ It defines:
66
+
67
+ - the retrieval contract and portable result models;
68
+ - stable chunk identity, score, and retrieval-mode semantics;
69
+ - optional capabilities such as corpus reading, query embedding, indexed chunk
70
+ vectors, generation, dataset listing, run snapshots, publication, and telemetry;
71
+ - reconstructable, secret-free connector registration;
72
+ - a conformance validator and reusable connector test kit;
73
+ - an optional FastEmbed + Chroma Reference RAG that can be provisioned from a
74
+ document folder, or from the 92-document demo corpus bundled with the package;
75
+ - an ingest kit (`rag_connector.ingest`) for hosts that build their own corpus —
76
+ deliberately outside the read-only connector contract.
77
+
78
+ The bundled Reference RAG is a known-good development and learning
79
+ implementation, not a production service.
80
+
81
+ ## Status
82
+
83
+ Pre-1.0, and used in production by its own authors. The retrieval contract,
84
+ chunk identity, the registry, the validator, document loading, fingerprinting
85
+ and the Reference RAG are all implemented and covered by the test suite; they
86
+ are what this package is for, and they are stable in practice.
87
+
88
+ What that means for depending on it: pin a version. The core contract —
89
+ `RagPipeline`, `ChunkRecord`, `RetrievedChunk`, retrieval modes, the registry
90
+ entry point — is the part least likely to move, and a change to it would break
91
+ the authors' own connectors first. The newer optional capabilities are younger
92
+ and may still gain fields. Anything listed under
93
+ [Reserved surface](https://github.com/staceyfarias/rag-connector/blob/main/docs/contract.md#reserved-surface) is exported but unwired:
94
+ nothing produces it, the validator does not check it, and it is not something
95
+ to build against yet.
96
+
97
+ ## Reference RAG
98
+
99
+ Install the optional implementation, provision a corpus, and validate the
100
+ result through the same black-box contract an external connector is held to:
101
+
102
+ ```bash
103
+ python -m pip install "rag-connector[reference]"
104
+ rag-connector reference provision \
105
+ --collection example \
106
+ --persist-dir ./rag_connector_data/reference
107
+ rag-connector validate \
108
+ --connector-type reference \
109
+ --params "{\"collection_name\":\"example\",\"persist_dir\":\"./rag_connector_data/reference\"}" \
110
+ --query "a question your documents can answer"
111
+ ```
112
+
113
+ With no folder named, `provision` uses the bundled **Pelorus Space** corpus — 92
114
+ short documents about a fictional interplanetary travel operator, MIT-licensed
115
+ like the rest of the package. Your own documents go in its place, as the first
116
+ positional argument:
117
+
118
+ ```bash
119
+ rag-connector reference provision ./my-documents \
120
+ --collection example \
121
+ --persist-dir ./rag_connector_data/reference
122
+ ```
123
+
124
+ Add `--normalized true|false` to declare whether the embedding space's vectors
125
+ are L2-normalized. Omit it and the space reports "not stated", which a host
126
+ reads as unverifiable — deliberately distinct from an asserted `false`. See
127
+ [the embedding space descriptor](https://github.com/staceyfarias/rag-connector/blob/main/docs/contract.md#the-embedding-space-descriptor).
128
+
129
+ The provision command prints the secret-free connection document needed to
130
+ reopen the same instance. Supported source formats are `.txt`, `.md`, `.pdf`,
131
+ `.docx`, and `.rtf`.
132
+
133
+ Two things to expect on a first run. The embedding model
134
+ (`BAAI/bge-small-en-v1.5`) is **downloaded on first use** into fastembed's
135
+ cache, so the first provision needs network and later ones do not. And **on
136
+ Windows**, install the `reference` extra into a virtualenv on a short path, or
137
+ enable long-path support first: onnxruntime, which fastembed depends on, ships
138
+ files whose full paths exceed the 260-character `MAX_PATH` limit, and pip
139
+ cannot unpack them under a deep directory.
140
+
141
+ ## Ingest kit
142
+
143
+ Loading a folder of documents, chunking it, and fingerprinting the result is
144
+ **not** part of the connector contract — a connector is a read interface onto a
145
+ system someone else already built, and nothing in that contract writes a corpus.
146
+ That work lives in `rag_connector.ingest`, which is declared API for the other
147
+ audience: a host building a corpus it owns.
148
+
149
+ ```python
150
+ from rag_connector.ingest import chunk_loaded_documents, load_documents
151
+
152
+ documents = load_documents("./my-documents")
153
+ chunks = chunk_loaded_documents(documents, chunk_size=1000, chunk_overlap=200)
154
+ ```
155
+
156
+ Each chunk's `char_start`/`char_end` are an invariant, not a hint:
157
+ `document.content[char_start:char_end] == chunk.text` exactly, so a stored
158
+ citation resolves back to its source. They index that **normalized text** — the
159
+ document as loaded, line endings read as `\n` — and not the bytes of the file,
160
+ so a CRLF source resolves by the same offsets a LF one does. A chunk id is
161
+ `<relative path>:chunk-<n>` — `booking_cancellation_policy.md:chunk-0` — and
162
+ subfolders are walked while dot-directories (`.git`, and any tool's state
163
+ folder) are not. The kit imports with no optional extras installed; a parser is
164
+ imported only when a file needs one.
165
+
166
+ ### The frozen chunk set
167
+
168
+ The bundled corpus also ships **pre-chunked**: 605 records at 1000/200, with
169
+ the ids anything downstream cites. Read it instead of re-chunking whenever a
170
+ stored artifact refers to a chunk — re-chunking is a re-derivation, and a
171
+ re-derivation that lands one character differently regrounds every citation.
172
+
173
+ ```python
174
+ from rag_connector.ingest import load_bundled_chunks
175
+ from rag_connector.reference_provision import provision_reference_rag_from_chunks
176
+
177
+ chunks = load_bundled_chunks() # 605 ChunkRecords
178
+ connector = provision_reference_rag_from_chunks(collection_name="frozen")
179
+ ```
180
+
181
+ The file is one JSON object per line in the shape RAGauge already writes, so
182
+ both products load the same corpus, the same split, and the same ids with no
183
+ adapter between them. Regenerating it is one library call —
184
+ `chunks_to_jsonl(build_bundled_chunks())`, written to
185
+ `rag_connector/corpus/<name>.chunks.jsonl` — and the suite fails if the
186
+ committed file and the kit ever disagree. (The repository keeps that call as
187
+ `tools/freeze_bundled_chunks.py`, which is development tooling and is not part
188
+ of the installed package.)
189
+
190
+ #### Reproducing it
191
+
192
+ Below is everything `build_bundled_chunks()` applies to the bundled documents.
193
+ That call **re-derives** the split; `load_bundled_chunks()` reads the frozen
194
+ file instead of re-deriving it, and is what anything citing a chunk id should
195
+ use. The settings are written down because a chunk set is only trustworthy if
196
+ someone else can arrive at the same one.
197
+
198
+ | Setting | Value |
199
+ |---|---|
200
+ | `chunk_size` | `1000` characters |
201
+ | `chunk_overlap` | `200` characters |
202
+ | Boundary search | last 20% of the window, preferring `
203
+
204
+ ` → `
205
+ ` → `. ` → `? ` → `! ` → space |
206
+ | Trailing chunk | dropped when its span lies wholly inside its predecessor |
207
+ | Whitespace | each chunk is stripped, and `char_start`/`char_end` record the **stripped** span |
208
+ | Source identity | the corpus-relative path, forward slashes, casefolded |
209
+ | Chunk id | `<source identity>:chunk-<n>`, `<n>` dense over emitted chunks |
210
+ | File discovery | recursive, sorted by relative POSIX path, dot-directories skipped, `.txt .md .pdf .docx .rtf` |
211
+ | Corpus layout | the 92 bundled documents, flat |
212
+
213
+ Verify a reproduction byte-for-byte. This re-derives the split from the
214
+ bundled documents and digests it, and needs no optional extras:
215
+
216
+ ```bash
217
+ python -c "import hashlib; from rag_connector.ingest import build_bundled_chunks, chunks_to_jsonl; b = chunks_to_jsonl(build_bundled_chunks()).encode('utf-8'); print(hashlib.sha256(b).hexdigest(), len(b), b.count(b'\n'))"
218
+ ```
219
+
220
+ ```
221
+ 290ca86076d9b878e54bdd2d5c1a1c870c4ab20c57edd1a013983e8782ba11fc 857219 605
222
+ ```
223
+
224
+ That is the sha256, the byte count and the line count of the shipped file
225
+ itself (LF; it is pinned to LF in `.gitattributes`), so digesting the file
226
+ rather than a re-derivation gives the same three numbers.
227
+
228
+ Two things that are **not** in the table, deliberately:
229
+
230
+ - **The embedding model does not affect the chunk set.** It decides the vectors,
231
+ not the split. The bundled datasets happen to use `BAAI/bge-small-en-v1.5`
232
+ (384-dim, cosine), and a different model over these same 605 chunks is a
233
+ different index of the same substrate.
234
+ - **Everything above is version-bound.** The boundary rule and the trailing-chunk
235
+ rule are behaviour, not configuration, so reproducing the hash needs the same
236
+ `rag-connector` release as well as the same parameters. A change to either
237
+ moves the frozen set, which is why the drift test exists.
238
+
239
+ One caveat if you reproduce this recipe on **your own** corpus rather than the
240
+ bundled one: text and Markdown are deterministic, but PDF, DOCX and RTF go
241
+ through third-party parsers whose output can change between their own releases.
242
+ A corpus of `.txt`/`.md` reproduces exactly; one full of PDFs reproduces only
243
+ against a pinned parser.
244
+
245
+ ## Installed connectors
246
+
247
+ List every connector registered in the current environment — the bundled
248
+ Reference RAG plus any independently installed connector package:
249
+
250
+ ```bash
251
+ rag-connector list
252
+ ```
253
+
254
+ A connector that does not appear here is not installed in this environment, or
255
+ its package publishes no `rag_connector.connectors` entry point.
256
+
257
+ ## Development
258
+
259
+ ```bash
260
+ python -m pip install -e ".[dev,reference]"
261
+ python -m pytest
262
+ python -m ruff check .
263
+ ```
264
+
265
+ `reference` is in there because the Reference RAG's own tests need chromadb and
266
+ fastembed. Without it the suite still runs clean — those tests skip, naming the
267
+ install that turns them back on.
268
+
269
+ See [Contract](https://github.com/staceyfarias/rag-connector/blob/main/docs/contract.md) and the
270
+ [Connector author guide](https://github.com/staceyfarias/rag-connector/blob/main/docs/connector-author-guide.md).
271
+
272
+ RAG Connector is licensed under the [MIT License](https://github.com/staceyfarias/rag-connector/blob/main/LICENSE).
@@ -0,0 +1,235 @@
1
+ # RAG Connector
2
+
3
+ RAG Connector is the shared contract that lets **RAGauge** and **Pelorus Query**
4
+ work with the same Retrieval-Augmented Generation pipeline as a black box —
5
+ without either of them knowing how that pipeline is built.
6
+
7
+ - **RAGauge** evaluates a RAG system: whether it retrieves the right evidence
8
+ and whether its answers are grounded in it, measured against a frozen corpus
9
+ and test set.
10
+ - **Pelorus Query** sits on top of a RAG system and serves curated,
11
+ source-grounded answers (Evidence Extracts), improving them as queries repeat.
12
+
13
+ Both need the same things from the pipeline underneath: ask it a question and
14
+ get ranked chunks back, read its corpus, know which embedding space and
15
+ retrieval mode produced a score, and cite a chunk by an id that means the same
16
+ thing to both. RAG Connector defines that once. A pipeline that implements the
17
+ contract — or is wrapped by a connector that does — can be measured by RAGauge
18
+ and curated by Pelorus Query, and a chunk one of them cites resolves to the
19
+ same text in the other. The bundled demo corpus and its frozen chunk set exist
20
+ for exactly that: both products are built against the same 605 chunk ids.
21
+
22
+ Interoperability between those two products is why this library exists, and it
23
+ is why the contract is narrow. Nothing here evaluates or curates: evaluation
24
+ policy belongs to RAGauge, curation behavior to Pelorus Query. Any other tool
25
+ that needs to treat a RAG pipeline as a black box can build on the same
26
+ contract. (RAGauge and Pelorus Query are not public yet.)
27
+
28
+ It defines:
29
+
30
+ - the retrieval contract and portable result models;
31
+ - stable chunk identity, score, and retrieval-mode semantics;
32
+ - optional capabilities such as corpus reading, query embedding, indexed chunk
33
+ vectors, generation, dataset listing, run snapshots, publication, and telemetry;
34
+ - reconstructable, secret-free connector registration;
35
+ - a conformance validator and reusable connector test kit;
36
+ - an optional FastEmbed + Chroma Reference RAG that can be provisioned from a
37
+ document folder, or from the 92-document demo corpus bundled with the package;
38
+ - an ingest kit (`rag_connector.ingest`) for hosts that build their own corpus —
39
+ deliberately outside the read-only connector contract.
40
+
41
+ The bundled Reference RAG is a known-good development and learning
42
+ implementation, not a production service.
43
+
44
+ ## Status
45
+
46
+ Pre-1.0, and used in production by its own authors. The retrieval contract,
47
+ chunk identity, the registry, the validator, document loading, fingerprinting
48
+ and the Reference RAG are all implemented and covered by the test suite; they
49
+ are what this package is for, and they are stable in practice.
50
+
51
+ What that means for depending on it: pin a version. The core contract —
52
+ `RagPipeline`, `ChunkRecord`, `RetrievedChunk`, retrieval modes, the registry
53
+ entry point — is the part least likely to move, and a change to it would break
54
+ the authors' own connectors first. The newer optional capabilities are younger
55
+ and may still gain fields. Anything listed under
56
+ [Reserved surface](https://github.com/staceyfarias/rag-connector/blob/main/docs/contract.md#reserved-surface) is exported but unwired:
57
+ nothing produces it, the validator does not check it, and it is not something
58
+ to build against yet.
59
+
60
+ ## Reference RAG
61
+
62
+ Install the optional implementation, provision a corpus, and validate the
63
+ result through the same black-box contract an external connector is held to:
64
+
65
+ ```bash
66
+ python -m pip install "rag-connector[reference]"
67
+ rag-connector reference provision \
68
+ --collection example \
69
+ --persist-dir ./rag_connector_data/reference
70
+ rag-connector validate \
71
+ --connector-type reference \
72
+ --params "{\"collection_name\":\"example\",\"persist_dir\":\"./rag_connector_data/reference\"}" \
73
+ --query "a question your documents can answer"
74
+ ```
75
+
76
+ With no folder named, `provision` uses the bundled **Pelorus Space** corpus — 92
77
+ short documents about a fictional interplanetary travel operator, MIT-licensed
78
+ like the rest of the package. Your own documents go in its place, as the first
79
+ positional argument:
80
+
81
+ ```bash
82
+ rag-connector reference provision ./my-documents \
83
+ --collection example \
84
+ --persist-dir ./rag_connector_data/reference
85
+ ```
86
+
87
+ Add `--normalized true|false` to declare whether the embedding space's vectors
88
+ are L2-normalized. Omit it and the space reports "not stated", which a host
89
+ reads as unverifiable — deliberately distinct from an asserted `false`. See
90
+ [the embedding space descriptor](https://github.com/staceyfarias/rag-connector/blob/main/docs/contract.md#the-embedding-space-descriptor).
91
+
92
+ The provision command prints the secret-free connection document needed to
93
+ reopen the same instance. Supported source formats are `.txt`, `.md`, `.pdf`,
94
+ `.docx`, and `.rtf`.
95
+
96
+ Two things to expect on a first run. The embedding model
97
+ (`BAAI/bge-small-en-v1.5`) is **downloaded on first use** into fastembed's
98
+ cache, so the first provision needs network and later ones do not. And **on
99
+ Windows**, install the `reference` extra into a virtualenv on a short path, or
100
+ enable long-path support first: onnxruntime, which fastembed depends on, ships
101
+ files whose full paths exceed the 260-character `MAX_PATH` limit, and pip
102
+ cannot unpack them under a deep directory.
103
+
104
+ ## Ingest kit
105
+
106
+ Loading a folder of documents, chunking it, and fingerprinting the result is
107
+ **not** part of the connector contract — a connector is a read interface onto a
108
+ system someone else already built, and nothing in that contract writes a corpus.
109
+ That work lives in `rag_connector.ingest`, which is declared API for the other
110
+ audience: a host building a corpus it owns.
111
+
112
+ ```python
113
+ from rag_connector.ingest import chunk_loaded_documents, load_documents
114
+
115
+ documents = load_documents("./my-documents")
116
+ chunks = chunk_loaded_documents(documents, chunk_size=1000, chunk_overlap=200)
117
+ ```
118
+
119
+ Each chunk's `char_start`/`char_end` are an invariant, not a hint:
120
+ `document.content[char_start:char_end] == chunk.text` exactly, so a stored
121
+ citation resolves back to its source. They index that **normalized text** — the
122
+ document as loaded, line endings read as `\n` — and not the bytes of the file,
123
+ so a CRLF source resolves by the same offsets a LF one does. A chunk id is
124
+ `<relative path>:chunk-<n>` — `booking_cancellation_policy.md:chunk-0` — and
125
+ subfolders are walked while dot-directories (`.git`, and any tool's state
126
+ folder) are not. The kit imports with no optional extras installed; a parser is
127
+ imported only when a file needs one.
128
+
129
+ ### The frozen chunk set
130
+
131
+ The bundled corpus also ships **pre-chunked**: 605 records at 1000/200, with
132
+ the ids anything downstream cites. Read it instead of re-chunking whenever a
133
+ stored artifact refers to a chunk — re-chunking is a re-derivation, and a
134
+ re-derivation that lands one character differently regrounds every citation.
135
+
136
+ ```python
137
+ from rag_connector.ingest import load_bundled_chunks
138
+ from rag_connector.reference_provision import provision_reference_rag_from_chunks
139
+
140
+ chunks = load_bundled_chunks() # 605 ChunkRecords
141
+ connector = provision_reference_rag_from_chunks(collection_name="frozen")
142
+ ```
143
+
144
+ The file is one JSON object per line in the shape RAGauge already writes, so
145
+ both products load the same corpus, the same split, and the same ids with no
146
+ adapter between them. Regenerating it is one library call —
147
+ `chunks_to_jsonl(build_bundled_chunks())`, written to
148
+ `rag_connector/corpus/<name>.chunks.jsonl` — and the suite fails if the
149
+ committed file and the kit ever disagree. (The repository keeps that call as
150
+ `tools/freeze_bundled_chunks.py`, which is development tooling and is not part
151
+ of the installed package.)
152
+
153
+ #### Reproducing it
154
+
155
+ Below is everything `build_bundled_chunks()` applies to the bundled documents.
156
+ That call **re-derives** the split; `load_bundled_chunks()` reads the frozen
157
+ file instead of re-deriving it, and is what anything citing a chunk id should
158
+ use. The settings are written down because a chunk set is only trustworthy if
159
+ someone else can arrive at the same one.
160
+
161
+ | Setting | Value |
162
+ |---|---|
163
+ | `chunk_size` | `1000` characters |
164
+ | `chunk_overlap` | `200` characters |
165
+ | Boundary search | last 20% of the window, preferring `
166
+
167
+ ` → `
168
+ ` → `. ` → `? ` → `! ` → space |
169
+ | Trailing chunk | dropped when its span lies wholly inside its predecessor |
170
+ | Whitespace | each chunk is stripped, and `char_start`/`char_end` record the **stripped** span |
171
+ | Source identity | the corpus-relative path, forward slashes, casefolded |
172
+ | Chunk id | `<source identity>:chunk-<n>`, `<n>` dense over emitted chunks |
173
+ | File discovery | recursive, sorted by relative POSIX path, dot-directories skipped, `.txt .md .pdf .docx .rtf` |
174
+ | Corpus layout | the 92 bundled documents, flat |
175
+
176
+ Verify a reproduction byte-for-byte. This re-derives the split from the
177
+ bundled documents and digests it, and needs no optional extras:
178
+
179
+ ```bash
180
+ python -c "import hashlib; from rag_connector.ingest import build_bundled_chunks, chunks_to_jsonl; b = chunks_to_jsonl(build_bundled_chunks()).encode('utf-8'); print(hashlib.sha256(b).hexdigest(), len(b), b.count(b'\n'))"
181
+ ```
182
+
183
+ ```
184
+ 290ca86076d9b878e54bdd2d5c1a1c870c4ab20c57edd1a013983e8782ba11fc 857219 605
185
+ ```
186
+
187
+ That is the sha256, the byte count and the line count of the shipped file
188
+ itself (LF; it is pinned to LF in `.gitattributes`), so digesting the file
189
+ rather than a re-derivation gives the same three numbers.
190
+
191
+ Two things that are **not** in the table, deliberately:
192
+
193
+ - **The embedding model does not affect the chunk set.** It decides the vectors,
194
+ not the split. The bundled datasets happen to use `BAAI/bge-small-en-v1.5`
195
+ (384-dim, cosine), and a different model over these same 605 chunks is a
196
+ different index of the same substrate.
197
+ - **Everything above is version-bound.** The boundary rule and the trailing-chunk
198
+ rule are behaviour, not configuration, so reproducing the hash needs the same
199
+ `rag-connector` release as well as the same parameters. A change to either
200
+ moves the frozen set, which is why the drift test exists.
201
+
202
+ One caveat if you reproduce this recipe on **your own** corpus rather than the
203
+ bundled one: text and Markdown are deterministic, but PDF, DOCX and RTF go
204
+ through third-party parsers whose output can change between their own releases.
205
+ A corpus of `.txt`/`.md` reproduces exactly; one full of PDFs reproduces only
206
+ against a pinned parser.
207
+
208
+ ## Installed connectors
209
+
210
+ List every connector registered in the current environment — the bundled
211
+ Reference RAG plus any independently installed connector package:
212
+
213
+ ```bash
214
+ rag-connector list
215
+ ```
216
+
217
+ A connector that does not appear here is not installed in this environment, or
218
+ its package publishes no `rag_connector.connectors` entry point.
219
+
220
+ ## Development
221
+
222
+ ```bash
223
+ python -m pip install -e ".[dev,reference]"
224
+ python -m pytest
225
+ python -m ruff check .
226
+ ```
227
+
228
+ `reference` is in there because the Reference RAG's own tests need chromadb and
229
+ fastembed. Without it the suite still runs clean — those tests skip, naming the
230
+ install that turns them back on.
231
+
232
+ See [Contract](https://github.com/staceyfarias/rag-connector/blob/main/docs/contract.md) and the
233
+ [Connector author guide](https://github.com/staceyfarias/rag-connector/blob/main/docs/connector-author-guide.md).
234
+
235
+ RAG Connector is licensed under the [MIT License](https://github.com/staceyfarias/rag-connector/blob/main/LICENSE).