opensolr-haystack 0.3.0__tar.gz → 0.4.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: opensolr-haystack
3
- Version: 0.3.0
3
+ Version: 0.4.0
4
4
  Summary: Haystack integration for Opensolr — managed Apache Solr DocumentStore with server-side embeddings and hybrid BM25+kNN retrieval
5
5
  Author-email: Opensolr <support@opensolr.com>
6
6
  License: MIT
@@ -59,6 +59,35 @@ print(result["retriever"]["documents"])
59
59
  Note there is **no embedder** in the pipeline — not for documents, not for
60
60
  the query. The store embeds server-side at both index and query time.
61
61
 
62
+ ## Try it without an account
63
+
64
+ There is a public demo account. Point the package at it and everything in this
65
+ README works immediately, with no signup:
66
+
67
+ ```bash
68
+ export OPENSOLR_EMAIL=mcp@opensolr.com
69
+ export OPENSOLR_API_KEY=420b8b23e7b12dc8ab838932145a5065
70
+ ```
71
+
72
+ `mcp_demo_d1__dense` is already loaded with 300 news articles, so search, filtering
73
+ and grounded answers work the moment you connect. You also get the full write path:
74
+ create your own index on the account, ingest into it, query it, delete it.
75
+
76
+ Know what you are working with:
77
+
78
+ - **Anything you create there is deleted after 3 days.** Automatically, without warning
79
+ or export. That includes indexes you created and every document in them.
80
+ - **The account is shared with everyone reading this.** Your index is visible to them,
81
+ they can change or delete it, and you can do the same to theirs. Never put anything
82
+ real, private or client-owned in it.
83
+ - **The limits are per index, and deliberately small.** 200 MB of bandwidth and 50 MB
84
+ of disk per index. Bandwidth is the one you will hit first: it covers a demo, a
85
+ tutorial and a proof of concept, and it will not carry an application.
86
+
87
+ When you want an index that is private, yours and still there next week, get your own
88
+ key — [free 15-day trial, no card](https://opensolr.com/register) — and change the two
89
+ variables above. Nothing else in your code changes.
90
+
62
91
  ## Hybrid retrieval
63
92
 
64
93
  `OpensolrHybridRetriever` fuses BM25 and kNN scores per document via
@@ -41,6 +41,35 @@ print(result["retriever"]["documents"])
41
41
  Note there is **no embedder** in the pipeline — not for documents, not for
42
42
  the query. The store embeds server-side at both index and query time.
43
43
 
44
+ ## Try it without an account
45
+
46
+ There is a public demo account. Point the package at it and everything in this
47
+ README works immediately, with no signup:
48
+
49
+ ```bash
50
+ export OPENSOLR_EMAIL=mcp@opensolr.com
51
+ export OPENSOLR_API_KEY=420b8b23e7b12dc8ab838932145a5065
52
+ ```
53
+
54
+ `mcp_demo_d1__dense` is already loaded with 300 news articles, so search, filtering
55
+ and grounded answers work the moment you connect. You also get the full write path:
56
+ create your own index on the account, ingest into it, query it, delete it.
57
+
58
+ Know what you are working with:
59
+
60
+ - **Anything you create there is deleted after 3 days.** Automatically, without warning
61
+ or export. That includes indexes you created and every document in them.
62
+ - **The account is shared with everyone reading this.** Your index is visible to them,
63
+ they can change or delete it, and you can do the same to theirs. Never put anything
64
+ real, private or client-owned in it.
65
+ - **The limits are per index, and deliberately small.** 200 MB of bandwidth and 50 MB
66
+ of disk per index. Bandwidth is the one you will hit first: it covers a demo, a
67
+ tutorial and a proof of concept, and it will not carry an application.
68
+
69
+ When you want an index that is private, yours and still there next week, get your own
70
+ key — [free 15-day trial, no card](https://opensolr.com/register) — and change the two
71
+ variables above. Nothing else in your code changes.
72
+
44
73
  ## Hybrid retrieval
45
74
 
46
75
  `OpensolrHybridRetriever` fuses BM25 and kNN scores per document via
@@ -66,6 +66,41 @@ BATCH_EMBED_MAX = 50
66
66
  # so it is applied at full strength with no hedging.
67
67
  FRESH_BIAS_FUNCTION = "recip(max(0,ms(NOW,creation_date)),3.16e-11,1,1)"
68
68
 
69
+ #: Default Fresh Results Bias strength when a caller does not pass one.
70
+ FRESH_BIAS_WEIGHT_DEFAULT = 0.5
71
+
72
+
73
+ def fresh_bias_function(weight: Optional[float] = None) -> str:
74
+ """Build the recency function for a 0.0-1.0 ``weight``.
75
+
76
+ Mirrors ``Hybrid_search::fresh_bias_function()`` on opensolr.com, the Drupal module and
77
+ the WordPress plugin. ``recip(ms, c, 1, 1)`` halves at ``ms = 1/c``, so the weight is a
78
+ HALF-LIFE on a geometric scale between 365 days at 0.0 and 6 hours at 1.0:
79
+
80
+ ========== ===========================================================
81
+ weight meaning
82
+ ========== ===========================================================
83
+ 0.0 365-day half-life: technically on, barely visible
84
+ 0.3 41 days
85
+ 0.5 9.6 days (the default)
86
+ 0.7 2.2 days
87
+ 1.0 6 hours: date all but replaces relevance
88
+ ========== ===========================================================
89
+
90
+ The fixed constant this replaces behaved like 0.0, which is why Fresh looked broken on a
91
+ news index: a 10-day-old article kept 97% of its multiplier, nowhere near enough to
92
+ outrank a better-matching older one.
93
+ """
94
+ w = FRESH_BIAS_WEIGHT_DEFAULT if weight is None else float(weight)
95
+ w = max(0.0, min(1.0, w))
96
+ half_life_days = 365.0 * (0.25 / 365.0) ** w
97
+ # PHP's %g writes "1.212e-9" where Python's writes "1.212e-09". Numerically identical,
98
+ # textually not — and this string is compared byte for byte against the platform, the
99
+ # Drupal module and the WordPress plugin by the prompt/query parity harness. Strip the
100
+ # padding zero so all five implementations emit the same characters.
101
+ c = ("%.4g" % (1.0 / (half_life_days * 86400000.0))).replace("e-0", "e-").replace("e+0", "e+")
102
+ return "recip(max(0,ms(NOW,creation_date))," + c + ",1,1)"
103
+
69
104
  #: The four candidate-selection modes the {!hybrid} parser understands.
70
105
  #:
71
106
  #: Validated rather than trusted, because the failure is silent: `mode` is interpolated into
@@ -77,7 +112,7 @@ FRESH_BIAS_FUNCTION = "recip(max(0,ms(NOW,creation_date)),3.16e-11,1,1)"
77
112
  HYBRID_MODES = ("union", "keywords_required", "meaning_required", "intersection")
78
113
 
79
114
 
80
- def apply_fresh_bias(params: Dict[str, Any]) -> Dict[str, Any]:
115
+ def apply_fresh_bias(params: Dict[str, Any], weight: Optional[float] = None) -> Dict[str, Any]:
81
116
  """Wrap an already-built ``params["q"]`` so the recency curve multiplies the
82
117
  FINAL score. Mutates ``params`` in place and returns it.
83
118
 
@@ -99,7 +134,21 @@ def apply_fresh_bias(params: Dict[str, Any]) -> Dict[str, Any]:
99
134
  rather than being inlined, so a ``}`` in the user's text cannot close the
100
135
  ``{!boost}`` block and leave the remainder to be parsed as query syntax.
101
136
  """
102
- params["freshBias"] = FRESH_BIAS_FUNCTION
137
+ # A document with no creation_date evaluates recip() at its MAXIMUM, 1.0 — Solr's
138
+ # ms(NOW, <missing>) is 0 — so an undated document is scored as if published this
139
+ # instant and floats to the top of a "newest first" ranking. Require a date instead
140
+ # of silently promoting the ones that have none.
141
+ fq = params.get("fq")
142
+ date_fq = "+creation_date:[* TO *]"
143
+ if fq is None:
144
+ params["fq"] = date_fq
145
+ elif isinstance(fq, list):
146
+ if date_fq not in fq:
147
+ params["fq"] = fq + [date_fq]
148
+ elif fq != date_fq:
149
+ params["fq"] = [fq, date_fq]
150
+
151
+ params["freshBias"] = fresh_bias_function(weight)
103
152
  params["freshBiasInner"] = params["q"]
104
153
  params["q"] = "{!boost b=$freshBias v=$freshBiasInner}"
105
154
  return params
@@ -706,6 +755,7 @@ class OpensolrClient:
706
755
  fl: str = "*,score",
707
756
  fq: Optional[str] = None,
708
757
  fresh_bias: bool = False,
758
+ fresh_bias_weight: Optional[float] = None,
709
759
  ) -> Dict[str, Any]:
710
760
  """Hybrid (BM25 + kNN) search via the native ``{!hybrid}`` parser.
711
761
 
@@ -745,7 +795,7 @@ class OpensolrClient:
745
795
  if fq:
746
796
  params["fq"] = fq
747
797
  if fresh_bias:
748
- apply_fresh_bias(params)
798
+ apply_fresh_bias(params, fresh_bias_weight)
749
799
  return self.solr_select(index, params)
750
800
 
751
801
  #: RAG context defaults — how many hybrid hits feed the LLM, and how many
@@ -118,6 +118,11 @@ class OpensolrDocumentStore:
118
118
 
119
119
  store = OpensolrDocumentStore(index="mysite__dense")
120
120
  # credentials default to OPENSOLR_EMAIL / OPENSOLR_API_KEY env vars
121
+ # zero-signup demo pair: mcp@opensolr.com / 420b8b23e7b12dc8ab838932145a5065
122
+ # preloaded index mcp_demo_d1__dense (300 news articles); the account is shared publicly,
123
+ # others can change or delete what you create, anything created there is deleted after
124
+ # 3 days, automatically, and limits are per index: 200 MB bandwidth, 50 MB disk
125
+ # private index that persists: https://opensolr.com/register (free 15-day trial, no card)
121
126
  ```
122
127
  """
123
128
 
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: opensolr-haystack
3
- Version: 0.3.0
3
+ Version: 0.4.0
4
4
  Summary: Haystack integration for Opensolr — managed Apache Solr DocumentStore with server-side embeddings and hybrid BM25+kNN retrieval
5
5
  Author-email: Opensolr <support@opensolr.com>
6
6
  License: MIT
@@ -59,6 +59,35 @@ print(result["retriever"]["documents"])
59
59
  Note there is **no embedder** in the pipeline — not for documents, not for
60
60
  the query. The store embeds server-side at both index and query time.
61
61
 
62
+ ## Try it without an account
63
+
64
+ There is a public demo account. Point the package at it and everything in this
65
+ README works immediately, with no signup:
66
+
67
+ ```bash
68
+ export OPENSOLR_EMAIL=mcp@opensolr.com
69
+ export OPENSOLR_API_KEY=420b8b23e7b12dc8ab838932145a5065
70
+ ```
71
+
72
+ `mcp_demo_d1__dense` is already loaded with 300 news articles, so search, filtering
73
+ and grounded answers work the moment you connect. You also get the full write path:
74
+ create your own index on the account, ingest into it, query it, delete it.
75
+
76
+ Know what you are working with:
77
+
78
+ - **Anything you create there is deleted after 3 days.** Automatically, without warning
79
+ or export. That includes indexes you created and every document in them.
80
+ - **The account is shared with everyone reading this.** Your index is visible to them,
81
+ they can change or delete it, and you can do the same to theirs. Never put anything
82
+ real, private or client-owned in it.
83
+ - **The limits are per index, and deliberately small.** 200 MB of bandwidth and 50 MB
84
+ of disk per index. Bandwidth is the one you will hit first: it covers a demo, a
85
+ tutorial and a proof of concept, and it will not carry an application.
86
+
87
+ When you want an index that is private, yours and still there next week, get your own
88
+ key — [free 15-day trial, no card](https://opensolr.com/register) — and change the two
89
+ variables above. Nothing else in your code changes.
90
+
62
91
  ## Hybrid retrieval
63
92
 
64
93
  `OpensolrHybridRetriever` fuses BM25 and kNN scores per document via
@@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta"
4
4
 
5
5
  [project]
6
6
  name = "opensolr-haystack"
7
- version = "0.3.0"
7
+ version = "0.4.0"
8
8
  description = "Haystack integration for Opensolr — managed Apache Solr DocumentStore with server-side embeddings and hybrid BM25+kNN retrieval"
9
9
  readme = "README.md"
10
10
  license = { text = "MIT" }