parse-stack-next 5.7.3 → 5.7.5
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +190 -0
- data/README.md +3 -0
- data/docs/atlas_vector_search_guide.md +125 -10
- data/docs/mcp_guide.md +81 -0
- data/lib/parse/agent/log_levels.rb +11 -0
- data/lib/parse/agent/mcp_dispatcher.rb +250 -13
- data/lib/parse/agent/mcp_rack_app.rb +120 -17
- data/lib/parse/agent.rb +40 -0
- data/lib/parse/atlas_search.rb +5 -1
- data/lib/parse/client.rb +85 -12
- data/lib/parse/embeddings/cache.rb +17 -5
- data/lib/parse/embeddings/provider.rb +26 -0
- data/lib/parse/embeddings/voyage.rb +214 -15
- data/lib/parse/embeddings.rb +3 -1
- data/lib/parse/model/core/embed_managed.rb +8 -2
- data/lib/parse/mongodb.rb +49 -2
- data/lib/parse/query.rb +5 -5
- data/lib/parse/retrieval/reranker/voyage.rb +282 -0
- data/lib/parse/retrieval/reranker.rb +2 -0
- data/lib/parse/stack/version.rb +1 -1
- data/parse-stack-next.gemspec +1 -0
- metadata +3 -1
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: 2d06f0d8932e328bae92fde22a568b9669614a02e75a5deed1a220884213d43e
|
|
4
|
+
data.tar.gz: edaaf3bb456a5396fdc08706188ca07d66aba66b17c7fce95220f130041644d8
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: 202606fe6edad9aafda2d48a216addfca376d24b01f1a8a39febce4fa511fcf5253f736632bf7aeff64645fe398d015fccb694ddf622688f3a78a6431b570ed2
|
|
7
|
+
data.tar.gz: a09f557be7a37d9f12937e0c70da314c617bd8ec18993cef2952336b7b55b6d1d9bbe3e0381933b70a4f3bdc3b2dcf3a71125edf83c460f5d15d4b0c58d3b3c8
|
data/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,195 @@
|
|
|
1
1
|
## parse-stack-next Changelog
|
|
2
2
|
|
|
3
|
+
### 5.7.5
|
|
4
|
+
|
|
5
|
+
#### MongoDB 9.0 support
|
|
6
|
+
|
|
7
|
+
Parse Server 9.10 and the SDK's mongo-direct paths run against MongoDB 9.0.
|
|
8
|
+
This release moves the test stack to 9.0, re-runs reads that 9.0 kills
|
|
9
|
+
mid-flight, and documents the server-side behavior changes that reach SDK
|
|
10
|
+
callers.
|
|
11
|
+
|
|
12
|
+
- **CHANGED**: The integration test stack runs `mongo:9` by default. Set
|
|
13
|
+
`MONGO_VERSION=8` to run it against the previous major. A data volume
|
|
14
|
+
written by one major is not guaranteed to start under another, so switch
|
|
15
|
+
with `docker-compose down -v` or a separate `PSNEXT_PREFIX`. Atlas Local
|
|
16
|
+
stays on 8.0, since no 9.x image of it is published yet.
|
|
17
|
+
- **CHANGED**: Applications using `Parse::MongoDB`, `Parse::AtlasSearch`,
|
|
18
|
+
or the `*_direct` query methods against a 9.0 server should require the
|
|
19
|
+
`mongo` driver 2.26 or newer, which adds handling for MongoDB 9.0's
|
|
20
|
+
overload (Intelligent Workload Management) errors. The gem's development
|
|
21
|
+
lock is already on 2.26.0.
|
|
22
|
+
- **NEW**: Mongo-direct reads (`Parse::MongoDB.aggregate`,
|
|
23
|
+
`Parse::MongoDB.find`, everything routed through them such as
|
|
24
|
+
`results_direct`, and `Parse::AtlasSearch` searches) re-run a query the
|
|
25
|
+
server killed with `QueryPlanKilled` (code 175; the 9.0 release notes call
|
|
26
|
+
it a `QueryKilledError`). MongoDB 9.0 kills a running query this way when an
|
|
27
|
+
indexed field it references becomes multikey because a concurrent write
|
|
28
|
+
stored an array there. The read is retried up to
|
|
29
|
+
`Parse::MongoDB::QUERY_KILLED_RETRIES` (2) times before the error
|
|
30
|
+
propagates. Retries are immediate, since the kill means a concurrent write
|
|
31
|
+
changed the index rather than that the server is overloaded. Each retry
|
|
32
|
+
emits a `parse.mongodb.query_killed_retry` notification. Other driver
|
|
33
|
+
errors are not retried.
|
|
34
|
+
|
|
35
|
+
#### Voyage `voyage-code-4` and contextualized chunk embeddings
|
|
36
|
+
|
|
37
|
+
- **NEW**: `Parse::Embeddings::Voyage` accepts `voyage-code-4` (1024
|
|
38
|
+
native, Matryoshka 256/512/1024/2048, 32k tokens). It routes through
|
|
39
|
+
`/v1/embeddings` like the other text models.
|
|
40
|
+
- **NEW**: `voyage-context-4` and `voyage-context-3` are supported. These
|
|
41
|
+
models embed each chunk together with the document it came from, through
|
|
42
|
+
Voyage's `/v1/contextualizedembeddings` endpoint. `embed_text` sends each
|
|
43
|
+
string as a one-chunk document, which is the right shape for queries. The
|
|
44
|
+
`embed` class macro uses the same path, so stored fields are embedded
|
|
45
|
+
without surrounding-document context; call `embed_chunks` for that. These
|
|
46
|
+
models default to `embed_batch_size: 32` rather than 128, since every input
|
|
47
|
+
is a whole document and Voyage caps a request at 120k tokens.
|
|
48
|
+
- **NEW**: `Voyage#embed_chunks(documents, input_type:)` takes one Array of
|
|
49
|
+
chunk Strings per document and returns one Array of chunk vectors per
|
|
50
|
+
document, aligned with the input. It enforces Voyage's per-request limits
|
|
51
|
+
(1,000 documents and 16,000 chunks) before any network call, and raises
|
|
52
|
+
`BadRequestError` on a model that is not contextualized. The endpoint has
|
|
53
|
+
no `truncation` field, so none is sent for these models. Large inputs are
|
|
54
|
+
sent as several requests, grouping whole documents so each response stays
|
|
55
|
+
within the provider's response-size cap; a document is never split across
|
|
56
|
+
requests. That split is sized by response vectors, not input tokens, so
|
|
57
|
+
it does not by itself keep a request under Voyage's 120k-token input cap.
|
|
58
|
+
- **NEW**: `Parse::Retrieval::Reranker::Voyage` wraps Voyage's `/v1/rerank`
|
|
59
|
+
and plugs into `Parse::Retrieval.retrieve(rerank:)` like the Cohere
|
|
60
|
+
reranker. It defaults to `rerank-3` and accepts `rerank-3-lite` and the
|
|
61
|
+
2.5 and 2 series. An Atlas model API key (`al-` prefix) routes to the
|
|
62
|
+
Atlas Embedding and Reranking API automatically. `truncation: false`
|
|
63
|
+
makes over-length inputs an error instead of truncating them.
|
|
64
|
+
|
|
65
|
+
#### MCP server speaks protocol version 2025-11-25
|
|
66
|
+
|
|
67
|
+
- **NEW**: The MCP server negotiates `2025-11-25` and advertises it as its
|
|
68
|
+
preferred version; `2025-06-18`, `2025-03-26`, and `2024-11-05` are still
|
|
69
|
+
accepted. `serverInfo` now carries `title` and `description`.
|
|
70
|
+
- **NEW**: `completion/complete` is implemented and the `completions`
|
|
71
|
+
capability advertised. Prompt arguments named `class_name`,
|
|
72
|
+
`parent_class`, `child_class`, or `classes`, and the `{className}`
|
|
73
|
+
variable of the `parse://` resource templates, complete to the class
|
|
74
|
+
names the connecting agent can see. `group_by` and `pointer_field`
|
|
75
|
+
complete to field names of the class given in the request's
|
|
76
|
+
`context.arguments`. Candidates come from the same tools that back
|
|
77
|
+
`resources/list` and `get_schema`, so hidden classes and fields are never
|
|
78
|
+
offered. Each completion runs those tools through `agent.execute`, so it
|
|
79
|
+
counts against the agent's rate limiter; clients should debounce.
|
|
80
|
+
- **NEW**: `logging/setLevel` is implemented. On a streaming request,
|
|
81
|
+
`notifications/message` events at or above the session's level are sent
|
|
82
|
+
on that request's response stream. Nothing is sent until the client sets a
|
|
83
|
+
level. The `logging` capability is advertised only by `MCPRackApp` with
|
|
84
|
+
streaming on, since no other transport can deliver the messages. A level
|
|
85
|
+
can be set only for a session that was initialized by the same principal,
|
|
86
|
+
so one caller cannot change another session's level. Tools emit messages with
|
|
87
|
+
`agent.log(level, data, logger:)`, and failed tool calls are logged at
|
|
88
|
+
`warning` with the tool name and error code.
|
|
89
|
+
- **FIXED**: Approval prompts are sent only to clients that accept form-mode
|
|
90
|
+
elicitation. Under `2025-11-25` a client declares its modes, and one that
|
|
91
|
+
declares only `url` would reject the form; the approval was then refused
|
|
92
|
+
and reported as a user cancellation. Such a client is now treated like one
|
|
93
|
+
without elicitation, so the destructive call is refused up front with
|
|
94
|
+
that reason. An empty `elicitation: {}` (the earlier shape) still means
|
|
95
|
+
form support.
|
|
96
|
+
- **CHANGED**: `initialize` with an `Mcp-Session-Id` already bound to a
|
|
97
|
+
different principal is refused with 403 instead of rebinding the session
|
|
98
|
+
to the new caller. Previously, knowing another session's id was enough to
|
|
99
|
+
take over its owner binding, and with it the listening stream, the
|
|
100
|
+
recorded elicitation capability, and the log level. The owning principal
|
|
101
|
+
can still re-initialize its own session.
|
|
102
|
+
- **CHANGED**: A `tools/call` whose `arguments` is not a JSON object now
|
|
103
|
+
returns a tool result with `isError: true` instead of an internal error,
|
|
104
|
+
as `2025-11-25` requires for input validation failures.
|
|
105
|
+
|
|
106
|
+
#### `server/discover` no longer fails the MCP version check
|
|
107
|
+
|
|
108
|
+
- **FIXED**: Newer MCP clients send `server/discover` before `initialize`,
|
|
109
|
+
carrying a protocol version this server does not support. The transport
|
|
110
|
+
rejected it with a 400 before it reached the dispatcher, so the client
|
|
111
|
+
treated the server as broken instead of falling back to `initialize`. The
|
|
112
|
+
request now reaches the dispatcher, which answers `-32601`, and the client
|
|
113
|
+
negotiates a supported version through `initialize`. Other methods still
|
|
114
|
+
get a 400 for an unsupported version.
|
|
115
|
+
|
|
116
|
+
#### Embedding cache entries are separated by deployment
|
|
117
|
+
|
|
118
|
+
- **FIXED**: `Parse::Embeddings::Cache` keyed entries by provider class,
|
|
119
|
+
model, dimensions, and input type, but not by endpoint. Two providers of
|
|
120
|
+
the same class and model pointed at different deployments (two
|
|
121
|
+
self-hosted `LocalHTTP` servers serving different weights under one model
|
|
122
|
+
name, or a provider behind a proxy) shared cache entries, so one could be
|
|
123
|
+
served the other's vector for the same query. The key now includes the
|
|
124
|
+
provider's deployment identity: the endpoint's scheme, host, non-default
|
|
125
|
+
port, and path, never credentials, userinfo, or a query string. Built-in
|
|
126
|
+
HTTP providers derive it from their `base_url`; `Provider#cache_identity`
|
|
127
|
+
can be overridden. Providers with no endpoint keep their existing keys,
|
|
128
|
+
so custom providers are unaffected. Existing cached entries for built-in
|
|
129
|
+
providers miss once and are re-filled.
|
|
130
|
+
|
|
131
|
+
#### Test infrastructure
|
|
132
|
+
|
|
133
|
+
- **CHANGED**: Development and CI use Bundler 4.0.22. The lockfile now
|
|
134
|
+
includes gem checksums; dependency versions are unchanged.
|
|
135
|
+
- **CHANGED**: The CI matrix's `3.5` lane, which resolved to a 2025
|
|
136
|
+
`3.5.0preview1` build, is replaced by Ruby 4.0. The unit suite passes on
|
|
137
|
+
Ruby 4.0.6.
|
|
138
|
+
- **NEW**: Snapshot fixtures pin the `$vectorSearch` pipeline (master,
|
|
139
|
+
user-session, strict-role, and caller-filter scopes), the native
|
|
140
|
+
`$rankFusion` pipeline (including that the ACL `$match` and final `$limit`
|
|
141
|
+
run after fusion), and `clp_scope` protected-field resolution and
|
|
142
|
+
redaction.
|
|
143
|
+
- **FIXED**: A streaming heartbeat test asserted how many heartbeats fired
|
|
144
|
+
before a tool's first progress report, which depends on scheduler timing
|
|
145
|
+
and failed intermittently on slower CI runners. It now asserts the
|
|
146
|
+
property under test: no heartbeat follows the first report.
|
|
147
|
+
|
|
148
|
+
### Behavior Notes
|
|
149
|
+
|
|
150
|
+
These are MongoDB 9.0 server changes, not SDK changes. They apply to REST
|
|
151
|
+
queries (Parse Server passes them to MongoDB) and to mongo-direct queries
|
|
152
|
+
alike.
|
|
153
|
+
|
|
154
|
+
- Equality and range comparisons against `null` (`$eq`, `$ne`, `$in`,
|
|
155
|
+
`$nin`, `$gte`, `$lte`, and `$lookup` equality) treat a dotted path that
|
|
156
|
+
traverses an array and resolves to no non-null value as `null`. For
|
|
157
|
+
`{ a: [] }` or `{ a: [1] }`, a query for `"a.b"` equal to `nil` now
|
|
158
|
+
matches, and `"a.b"` not equal to `nil` no longer does.
|
|
159
|
+
- `$group` rejects an accumulator with an empty field name.
|
|
160
|
+
- A query fails with `QueryPlanKilled` if an indexed field becomes
|
|
161
|
+
multikey while it runs. The SDK's mongo-direct reads re-run it (see
|
|
162
|
+
above); REST queries surface Parse Server's error.
|
|
163
|
+
- `$where`, `$function`, and `$accumulator` are no longer deprecated in
|
|
164
|
+
9.0. The SDK continues to block them in every pipeline and constraint it
|
|
165
|
+
validates.
|
|
166
|
+
|
|
167
|
+
### 5.7.4
|
|
168
|
+
|
|
169
|
+
#### Reset connections are retried instead of surfacing a raw Faraday error
|
|
170
|
+
|
|
171
|
+
A focused fix release for the request retry mechanism. A stale keep-alive
|
|
172
|
+
connection (a pooled persistent connection closed by the server or a load
|
|
173
|
+
balancer after idling) failed the next request with a raw
|
|
174
|
+
`Faraday::ConnectionFailed` / `Errno::ECONNRESET` instead of retrying, even
|
|
175
|
+
though an immediate re-send on a fresh connection succeeds. Reset connections
|
|
176
|
+
now retry under the same idempotency rules as read timeouts.
|
|
177
|
+
|
|
178
|
+
- **FIXED**: `Parse::Client#request` now rescues `Faraday::ConnectionFailed`
|
|
179
|
+
and inspects the wrapped cause. Reset-class causes (`Errno::ECONNRESET`,
|
|
180
|
+
`Errno::EPIPE`, `Errno::ECONNABORTED`, and the `EOFError` raised when the
|
|
181
|
+
remote end closes a keep-alive socket cleanly) are transient, so idempotent
|
|
182
|
+
requests (GET, DELETE, op-free PUT, and any write covered by asserted
|
|
183
|
+
server-side request-id dedup) retry with the standard backoff. Connection
|
|
184
|
+
refused and DNS failures keep the previous fail-fast behavior and propagate
|
|
185
|
+
the raw `Faraday::ConnectionFailed` with no retry latency.
|
|
186
|
+
- **CHANGED**: A reset connection that persists through the whole retry budget
|
|
187
|
+
now raises `Parse::Error::ConnectionError` (consistent with the read-timeout
|
|
188
|
+
path) instead of the raw `Faraday::ConnectionFailed`. Code that rescued
|
|
189
|
+
`Faraday::ConnectionFailed` to catch resets should rescue
|
|
190
|
+
`Parse::Error::ConnectionError` instead; refused and DNS failures still
|
|
191
|
+
raise `Faraday::ConnectionFailed`.
|
|
192
|
+
|
|
3
193
|
### 5.7.3
|
|
4
194
|
|
|
5
195
|
#### Stored values can no longer drive the operator's terminal
|
data/README.md
CHANGED
|
@@ -6,6 +6,9 @@ A full-featured Ruby client SDK for [Parse Server](http://parseplatform.org/). [
|
|
|
6
6
|
|
|
7
7
|
## What's new in 5.7
|
|
8
8
|
|
|
9
|
+
- **5.7.5: MongoDB 9.0 support.** The test stack runs MongoDB 9 by default (`MONGO_VERSION=8` selects the previous major), and use the `mongo` driver 2.26 or newer for 9.0's overload handling. MongoDB 9.0 changes how `null` comparisons treat dotted paths through arrays; see the behavior notes in [CHANGELOG.md](./CHANGELOG.md)
|
|
10
|
+
- **5.7.5: `voyage-code-4` and contextualized chunk embeddings.** The Voyage provider accepts `voyage-code-4`, `voyage-context-4`, and `voyage-context-3`. `Voyage#embed_chunks` embeds whole chunked documents so each chunk's vector carries its document's context. See [CHANGELOG.md](./CHANGELOG.md)
|
|
11
|
+
- **5.7.5: Voyage reranker and MCP 2025-11-25.** `Parse::Retrieval::Reranker::Voyage` adds `rerank-3` and `rerank-3-lite` reranking, including through the Atlas endpoint. The MCP server negotiates protocol `2025-11-25` and adds `completion/complete` (class and field names scoped to the agent) and `logging/setLevel`. Mongo-direct reads re-run queries MongoDB 9.0 kills mid-flight. See [CHANGELOG.md](./CHANGELOG.md)
|
|
9
12
|
- **5.7.0: Reserved, app-scoped cache keyspace.** `Parse::Cache::Keyspace` lays out response, identity, and role-cache keys on a shared backend (`parse-stack:v1:<app_scope>[:<namespace>]:<family>[:T:<tenant>]:<rest>`) and owns the glob patterns that clear them again, so key generation and eviction can no longer drift apart. `app_scope` is a digest of the application id and server URL, so two apps sharing one Redis no longer collide. Enable with `cache_keyspace: true` on `Parse.setup`; left unset, behavior is unchanged. See [CHANGELOG.md](./CHANGELOG.md)
|
|
10
13
|
- **5.7.0: `clear_cache!` stops falling back to `FLUSHDB`.** With `cache_keyspace: true`, `Parse::Client#clear_cache!` performs a scoped SCAN inside the client's own keys instead of flushing the whole database, which previously could destroy co-tenant data and drop `first_or_create!` create-locks on a shared Redis. `flush_db!` remains the explicit opt-in for a full flush. See [CHANGELOG.md](./CHANGELOG.md)
|
|
11
14
|
- **5.7.0: Response-cache auth separation enforced by construction.** `Parse::Cache::Keyspace#cache_key` now requires an `auth:` discriminator for the response-cache family and refuses to build a key without one, so a master-key body and a session-token body can no longer land under the same key by accident. A non-GET write now invalidates every auth variant of a resource in one scoped pattern instead of only the variants the process has already seen. See [CHANGELOG.md](./CHANGELOG.md)
|
|
@@ -120,8 +120,11 @@ providers:
|
|
|
120
120
|
`{ type: "image_url", image_url: { url: ... } }` content rows.
|
|
121
121
|
* `Parse::Embeddings::Voyage` — voyage-4 family (`voyage-4-large` 2048,
|
|
122
122
|
Matryoshka; `voyage-4` 1024; `voyage-4-lite` 512; `voyage-4-nano` 256),
|
|
123
|
-
voyage-3 family, domain models (`voyage-code-
|
|
124
|
-
`voyage-
|
|
123
|
+
voyage-3 family, domain models (`voyage-code-4`, `voyage-code-3`,
|
|
124
|
+
`voyage-finance-2`, `voyage-law-2`), contextualized chunk models
|
|
125
|
+
(`voyage-context-4`, `voyage-context-3`; route to
|
|
126
|
+
`/v1/contextualizedembeddings`, with whole chunked documents embedded
|
|
127
|
+
through `embed_chunks`), and `voyage-multimodal-3` (1024-dim, 32k token
|
|
125
128
|
context, routes to `/v1/multimodalembeddings` with the wrapped
|
|
126
129
|
`{inputs: [{content: [{type: "text", text: ...}]}]}` envelope for
|
|
127
130
|
text and `{type: "image_url", image_url: <url>}` content rows for
|
|
@@ -330,7 +333,8 @@ Every `text:`-overload query funnels through one embed path
|
|
|
330
333
|
```ruby
|
|
331
334
|
# Opt-in query-embed cache: repeated identical queries skip the
|
|
332
335
|
# provider round-trip. Keyed by (provider, model, dimensions,
|
|
333
|
-
# input_type, SHA-256(input))
|
|
336
|
+
# input_type, deployment endpoint, SHA-256(input)); plaintext and
|
|
337
|
+
# credentials never land in the store.
|
|
334
338
|
Parse::Embeddings::Cache.enable!(max_entries: 2048, ttl: 600)
|
|
335
339
|
Parse::Embeddings::Cache.stats # => { enabled:, hits:, misses:, size: }
|
|
336
340
|
|
|
@@ -451,6 +455,69 @@ embed-time chunking), use one of these patterns:
|
|
|
451
455
|
similarity search against the chunk collection, then hydrate
|
|
452
456
|
parents as needed.
|
|
453
457
|
|
|
458
|
+
### Contextualized chunk embeddings (`voyage-context-4`)
|
|
459
|
+
|
|
460
|
+
An ordinary embedding of a chunk sees only that chunk. "It ships Friday."
|
|
461
|
+
embeds the same no matter which product the surrounding document is
|
|
462
|
+
about. Voyage's contextualized models (`voyage-context-4`,
|
|
463
|
+
`voyage-context-3`) embed every chunk together with the rest of its
|
|
464
|
+
document, so each chunk vector also carries what the document is about.
|
|
465
|
+
|
|
466
|
+
To get that, send a document's chunks together with `embed_chunks`:
|
|
467
|
+
|
|
468
|
+
```ruby
|
|
469
|
+
voyage = Parse::Embeddings::Voyage.new(
|
|
470
|
+
api_key: ENV.fetch("VOYAGE_API_KEY"),
|
|
471
|
+
model: "voyage-context-4", # 1024 dims; 256/512/2048 via dimensions:
|
|
472
|
+
)
|
|
473
|
+
|
|
474
|
+
documents = [
|
|
475
|
+
["The Atlas release slipped a week.", "It ships Friday."], # one document, two chunks
|
|
476
|
+
["Billing moves to the new provider.", "It ships Friday."],
|
|
477
|
+
]
|
|
478
|
+
vectors = voyage.embed_chunks(documents, input_type: :search_document)
|
|
479
|
+
# => [[vec_a1, vec_a2], [vec_b1, vec_b2]] one Array per document, aligned with its chunks
|
|
480
|
+
# vec_a2 != vec_b2: the same sentence, embedded with different documents.
|
|
481
|
+
|
|
482
|
+
# Store each chunk vector on its own record. The chunk class declares the
|
|
483
|
+
# vector property but NOT the `embed` macro: `embed :content` would
|
|
484
|
+
# recompute the vector from `content` alone on save, silently replacing the
|
|
485
|
+
# contextual vector with an ordinary one.
|
|
486
|
+
class PostChunk < Parse::Object
|
|
487
|
+
belongs_to :post
|
|
488
|
+
property :content, :string
|
|
489
|
+
property :embedding, :vector, dimensions: 1024
|
|
490
|
+
end
|
|
491
|
+
|
|
492
|
+
documents.zip(vectors).each do |chunks, chunk_vectors|
|
|
493
|
+
chunks.zip(chunk_vectors).each { |text, vec| PostChunk.create!(content: text, embedding: vec) }
|
|
494
|
+
end
|
|
495
|
+
|
|
496
|
+
# Queries are single strings: embed_text sends each as a one-chunk document.
|
|
497
|
+
query_vec = voyage.embed_text(["when does it ship?"], input_type: :search_query).first
|
|
498
|
+
PostChunk.find_similar(vector: query_vec, k: 5)
|
|
499
|
+
```
|
|
500
|
+
|
|
501
|
+
Two things to know:
|
|
502
|
+
|
|
503
|
+
* **The `embed` macro does not contextualize.** It sends each record's
|
|
504
|
+
source text through `embed_text`, which treats every input as its own
|
|
505
|
+
one-chunk document. Registering a context model as an `embed` provider
|
|
506
|
+
works, but the stored vectors carry no surrounding-document context.
|
|
507
|
+
Use `embed_chunks` when chunk-level context is the point.
|
|
508
|
+
* **Request sizing is partly handled for you.** Voyage caps one request
|
|
509
|
+
at 1,000 documents, 16,000 chunks, and 120k tokens. `embed_chunks`
|
|
510
|
+
checks the document and chunk caps before sending, and splits large
|
|
511
|
+
inputs across several requests by whole document (a document's chunks
|
|
512
|
+
always travel together) so that each **response** stays under the
|
|
513
|
+
SDK's response-size cap. That split is sized by the vectors coming
|
|
514
|
+
back, not by tokens going out, so it does **not** guarantee a request
|
|
515
|
+
stays under the 120k input-token cap. The SDK has no tokenizer to
|
|
516
|
+
check that; keep long documents to a few per call, and expect a
|
|
517
|
+
`BadRequestError` from Voyage if a request exceeds it. Context models
|
|
518
|
+
default to `embed_batch_size: 32` to keep `embed_text` batches clear of
|
|
519
|
+
the token cap for typical inputs, which is a heuristic, not a check.
|
|
520
|
+
|
|
454
521
|
---
|
|
455
522
|
|
|
456
523
|
## Retrieval (RAG)
|
|
@@ -541,6 +608,19 @@ when the cluster does not support it; the default `:rrf` always fuses
|
|
|
541
608
|
client-side, which is the fully-enforced, deterministic path. `$rankFusion`
|
|
542
609
|
is admitted to `PipelineSecurity::ALLOWED_STAGES` for the native path.
|
|
543
610
|
|
|
611
|
+
*Status:* the native path is shipped and its pipeline shape is pinned by
|
|
612
|
+
unit and snapshot tests. Unlike the client-side path, where each branch
|
|
613
|
+
enforces ACL and CLP before fusion, the native path applies the ACL
|
|
614
|
+
`$match` **after** `$rankFusion`: the stage fuses the unfiltered candidate
|
|
615
|
+
sets, then rows the caller cannot read are dropped. To keep a scoped caller
|
|
616
|
+
from being underfilled by those drops, the branches request a wider
|
|
617
|
+
candidate window and the final `$limit` runs after the ACL match; a caller
|
|
618
|
+
who can read only a small fraction of the collection can still receive
|
|
619
|
+
fewer than `k` results. Fused scores are recomputed from surviving rows so
|
|
620
|
+
an unreadable row's rank does not leak through. It has not yet been
|
|
621
|
+
validated end to end against a live Atlas cluster, which is why it stays
|
|
622
|
+
opt-in rather than the default.
|
|
623
|
+
|
|
544
624
|
`Parse::Retrieval.retrieve(hybrid: true, ...)` routes through
|
|
545
625
|
`hybrid_search` and chunks the fused results; pass `hybrid: { lexical:,
|
|
546
626
|
vector:, fusion: }` to configure the branches. Tenant scope is folded into
|
|
@@ -565,6 +645,27 @@ chunks = Parse::Retrieval.retrieve(
|
|
|
565
645
|
# Reranked chunks' score is the cross-encoder relevance_score.
|
|
566
646
|
```
|
|
567
647
|
|
|
648
|
+
Voyage's rerankers plug in the same way. `rerank-3` is the default and
|
|
649
|
+
`rerank-3-lite` is the cheaper option. An Atlas model API key (`al-`
|
|
650
|
+
prefix) routes to MongoDB's Atlas Embedding and Reranking API
|
|
651
|
+
automatically, exactly like the Voyage embeddings provider:
|
|
652
|
+
|
|
653
|
+
```ruby
|
|
654
|
+
reranker = Parse::Retrieval::Reranker::Voyage.new(
|
|
655
|
+
api_key: ENV.fetch("VOYAGE_API_KEY"), # or an Atlas "al-..." key
|
|
656
|
+
model: "rerank-3-lite",
|
|
657
|
+
truncation: false, # over-length input raises instead of truncating
|
|
658
|
+
)
|
|
659
|
+
chunks = Parse::Retrieval.retrieve(
|
|
660
|
+
query: "reset my password", klass: Article, k: 30,
|
|
661
|
+
rerank: reranker, rerank_top_n: 5,
|
|
662
|
+
)
|
|
663
|
+
|
|
664
|
+
# Or directly, outside retrieve:
|
|
665
|
+
reranker.rerank(query: "capital of France", documents: docs, top_n: 3)
|
|
666
|
+
# => [#<struct Result index=1, relevance_score=0.91>, ...]
|
|
667
|
+
```
|
|
668
|
+
|
|
568
669
|
`Reranker::Fixture` is a deterministic, zero-network reranker (lexical
|
|
569
670
|
token overlap) for tests. The `Reranker::Base` protocol validates inputs,
|
|
570
671
|
bounds `top_n`, rejects out-of-range indices, and sorts descending —
|
|
@@ -655,8 +756,9 @@ envelope. See the [MCP guide's Token Economy section](./mcp_guide.md#token-econo
|
|
|
655
756
|
|
|
656
757
|
`embed_image` is the image-source counterpart to `embed`. The source
|
|
657
758
|
property must be `:file`-typed; the target must be a `:vector` property
|
|
658
|
-
whose declared `provider:` supports multimodal input (
|
|
659
|
-
|
|
759
|
+
whose declared `provider:` supports multimodal input (`:voyage` with
|
|
760
|
+
`voyage-multimodal-3.5` or `voyage-multimodal-3`, or `:cohere` with
|
|
761
|
+
`embed-v4.0`).
|
|
660
762
|
|
|
661
763
|
Two fetch modes, selected per declaration with `source:`:
|
|
662
764
|
|
|
@@ -782,7 +884,17 @@ Direct provider calls accept the same shape:
|
|
|
782
884
|
|
|
783
885
|
### Save-side semantics
|
|
784
886
|
|
|
785
|
-
*
|
|
887
|
+
* **Private buckets.** When the file adapter returns signed URLs (S3 or GCS
|
|
888
|
+
with `presignedUrl: true`), `file.url` holds the canonical URL with the
|
|
889
|
+
signature stripped and the signed form is kept in `file.presigned_url`.
|
|
890
|
+
Recompute fetches through `file.presigned_url` while it is still valid
|
|
891
|
+
(for both `source: :url` and `source: :bytes`) and falls back to the bare
|
|
892
|
+
URL otherwise, so a private object is reachable without making the
|
|
893
|
+
bucket public. Make sure the URL's lifetime outlasts the save: an expired
|
|
894
|
+
signature falls back to the bare URL, which a private bucket refuses.
|
|
895
|
+
* Digest is the **SHA-256 of the canonical URL String** (signature
|
|
896
|
+
stripped), not the file bytes, so a save that only rotates the
|
|
897
|
+
signature does not re-embed.
|
|
786
898
|
Replacing the `Parse::File` with one pointing at a different URL
|
|
787
899
|
re-embeds; resaving the same URL is a no-op (zero provider calls).
|
|
788
900
|
Parse-managed file URLs are stable unless overwritten in place — if
|
|
@@ -841,9 +953,12 @@ each row's digest sibling (so the save-path recompute cannot elide the
|
|
|
841
953
|
provider call), and saves. Unlike `embed_pending!` — which only fills
|
|
842
954
|
NULL vectors — `reembed!` recomputes populated rows too. Run it with a
|
|
843
955
|
master-key client (or pass `save_opts:` with a session token that can
|
|
844
|
-
write every row).
|
|
845
|
-
|
|
846
|
-
|
|
956
|
+
write every row). `batch_size:` is the query page size: it controls how
|
|
957
|
+
many records are fetched per round, not how many inputs go into a provider
|
|
958
|
+
request. Each row's save makes its own provider call; pace bulk runs
|
|
959
|
+
against provider rate limits (see `BatchEmbedder` below for the pattern,
|
|
960
|
+
or just throttle the loop). `embed_pending!(batch_size:)` pages the same
|
|
961
|
+
way.
|
|
847
962
|
|
|
848
963
|
### Changed-width migrations: dual-field workflow
|
|
849
964
|
|
|
@@ -932,7 +1047,7 @@ Payload contract (keys always present; values may be nil):
|
|
|
932
1047
|
| `:input_count` | `Integer` | batch size |
|
|
933
1048
|
| `:input_type` | `Symbol` | `:search_query` / `:search_document` |
|
|
934
1049
|
| `:total_tokens`| `Integer`/nil | provider-reported usage; nil for Fixture and providers without usage |
|
|
935
|
-
| `:cached` | `Boolean` |
|
|
1050
|
+
| `:cached` | `Boolean` | true when `Parse::Embeddings::Cache` served the vector (no provider call) |
|
|
936
1051
|
| `:error` | `String`/nil | `exception.class.name` when the block raised — class name only |
|
|
937
1052
|
|
|
938
1053
|
Notes:
|
data/docs/mcp_guide.md
CHANGED
|
@@ -766,6 +766,87 @@ out-of-band / clustered publisher that needs the lower-level `publish` seam.
|
|
|
766
766
|
|
|
767
767
|
---
|
|
768
768
|
|
|
769
|
+
## Argument Completion (`completion/complete`)
|
|
770
|
+
|
|
771
|
+
The server advertises the `completions` capability, so clients that support
|
|
772
|
+
it (MCP Inspector, IDE integrations) can autocomplete prompt arguments and
|
|
773
|
+
resource-template variables:
|
|
774
|
+
|
|
775
|
+
| Argument | Completes to |
|
|
776
|
+
|---|---|
|
|
777
|
+
| `class_name`, `parent_class`, `child_class` on any prompt; `{className}` in `parse://{className}/...` | Class names the connecting agent can see |
|
|
778
|
+
| `classes` (comma-separated) | The last segment of the list |
|
|
779
|
+
| `group_by`, `pointer_field` | Field names of the class named by `class_name` (or `child_class`) in the request's `context.arguments` |
|
|
780
|
+
|
|
781
|
+
```json
|
|
782
|
+
{"jsonrpc":"2.0","id":7,"method":"completion/complete","params":{
|
|
783
|
+
"ref":{"type":"ref/prompt","name":"count_by"},
|
|
784
|
+
"argument":{"name":"group_by","value":"st"},
|
|
785
|
+
"context":{"arguments":{"class_name":"Post"}}}}
|
|
786
|
+
```
|
|
787
|
+
|
|
788
|
+
```json
|
|
789
|
+
{"jsonrpc":"2.0","id":7,"result":{"completion":{"values":["status","state"],"total":2,"hasMore":false}}}
|
|
790
|
+
```
|
|
791
|
+
|
|
792
|
+
Candidates come from the same `get_all_schemas` / `get_schema` tools that back
|
|
793
|
+
`resources/list`, so `agent_hidden` classes, the `classes:` allowlist, and
|
|
794
|
+
`agent_fields` all apply: a completion never offers a name the agent could not
|
|
795
|
+
already see. Each completion is a tool call for rate-limiting and audit
|
|
796
|
+
purposes, so clients that complete on every keystroke should debounce. Custom
|
|
797
|
+
prompts get class-name completion automatically by naming an argument
|
|
798
|
+
`class_name`, `parent_class`, `child_class`, or `classes`.
|
|
799
|
+
|
|
800
|
+
---
|
|
801
|
+
|
|
802
|
+
## Logging (`logging/setLevel`)
|
|
803
|
+
|
|
804
|
+
On `MCPRackApp` with streaming on, the server advertises the `logging`
|
|
805
|
+
capability. A client opts in by setting a minimum level for its session; from
|
|
806
|
+
then on, log messages at or above that level ride the response stream of each
|
|
807
|
+
streamed request as `notifications/message`, ahead of the final response.
|
|
808
|
+
Nothing is sent before the client sets a level, and the WEBrick `MCPServer`
|
|
809
|
+
and non-streaming Rack mounts do not advertise the capability at all.
|
|
810
|
+
|
|
811
|
+
```json
|
|
812
|
+
{"jsonrpc":"2.0","id":3,"method":"logging/setLevel","params":{"level":"warning"}}
|
|
813
|
+
```
|
|
814
|
+
|
|
815
|
+
Built-in behavior: a failed `tools/call` logs at `warning` with the tool name
|
|
816
|
+
and error code. Custom tools log through the agent:
|
|
817
|
+
|
|
818
|
+
```ruby
|
|
819
|
+
Parse::Agent::Tools.register(
|
|
820
|
+
name: :reindex_posts,
|
|
821
|
+
description: "Rebuild the Post search index",
|
|
822
|
+
parameters: { type: "object", properties: {} },
|
|
823
|
+
permission: :readonly,
|
|
824
|
+
handler: lambda do |agent, **_args|
|
|
825
|
+
agent.log(:info, { "step" => "scanning", "rows" => 1200 }, logger: "reindex")
|
|
826
|
+
# ...
|
|
827
|
+
agent.log(:warning, "3 rows skipped: missing title")
|
|
828
|
+
{ reindexed: 1197 }
|
|
829
|
+
end,
|
|
830
|
+
)
|
|
831
|
+
```
|
|
832
|
+
|
|
833
|
+
`agent.log` takes an RFC 5424 level (`:debug`, `:info`, `:notice`,
|
|
834
|
+
`:warning`, `:error`, `:critical`, `:alert`, `:emergency`; anything else
|
|
835
|
+
raises `ArgumentError` on every transport) and any JSON-serializable data.
|
|
836
|
+
It is a no-op when no client is listening. Do not log secrets or rows the
|
|
837
|
+
agent's scope cannot read; the message goes to whoever holds the session.
|
|
838
|
+
|
|
839
|
+
A level is stored per `Mcp-Session-Id` and can be set only by the principal
|
|
840
|
+
that initialized that session, so one caller cannot silence or flood another
|
|
841
|
+
session's logs. `DELETE` on the session forgets it.
|
|
842
|
+
|
|
843
|
+
> **Spec note.** MCP `2026-07-28` removes `logging/setLevel` in favor of a
|
|
844
|
+
> per-request log level in `_meta` and deprecates the logging feature. This
|
|
845
|
+
> server targets `2025-11-25` and earlier, where `logging/setLevel` is the
|
|
846
|
+
> mechanism. `agent.log` is the stable API either way.
|
|
847
|
+
|
|
848
|
+
---
|
|
849
|
+
|
|
769
850
|
## Built-in Agent Hardening & Telemetry
|
|
770
851
|
|
|
771
852
|
5.2 adds several agent-side controls, all configured on `Parse::Agent`:
|
|
@@ -0,0 +1,11 @@
|
|
|
1
|
+
# encoding: UTF-8
|
|
2
|
+
# frozen_string_literal: true
|
|
3
|
+
|
|
4
|
+
module Parse
|
|
5
|
+
class Agent
|
|
6
|
+
# RFC 5424 severities used by MCP logging, least to most severe. Shared
|
|
7
|
+
# by {Parse::Agent#log} and {Parse::Agent::MCPDispatcher} so a level is
|
|
8
|
+
# validated the same way whether or not a transport is attached.
|
|
9
|
+
LOG_LEVELS = %w[debug info notice warning error critical alert emergency].freeze
|
|
10
|
+
end
|
|
11
|
+
end
|