goodmem-deepeval 0.3.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- goodmem_deepeval-0.3.0/.gitignore +21 -0
- goodmem_deepeval-0.3.0/CHANGELOG.md +171 -0
- goodmem_deepeval-0.3.0/LICENSE +201 -0
- goodmem_deepeval-0.3.0/PKG-INFO +232 -0
- goodmem_deepeval-0.3.0/README.md +201 -0
- goodmem_deepeval-0.3.0/pyproject.toml +81 -0
- goodmem_deepeval-0.3.0/src/goodmem_deepeval/__init__.py +35 -0
- goodmem_deepeval-0.3.0/src/goodmem_deepeval/_connection.py +100 -0
- goodmem_deepeval-0.3.0/src/goodmem_deepeval/_results.py +159 -0
- goodmem_deepeval-0.3.0/src/goodmem_deepeval/_spaces.py +124 -0
- goodmem_deepeval-0.3.0/src/goodmem_deepeval/_typing.py +49 -0
- goodmem_deepeval-0.3.0/src/goodmem_deepeval/filters.py +133 -0
- goodmem_deepeval-0.3.0/src/goodmem_deepeval/metrics.py +89 -0
- goodmem_deepeval-0.3.0/src/goodmem_deepeval/py.typed +0 -0
- goodmem_deepeval-0.3.0/src/goodmem_deepeval/retriever.py +320 -0
- goodmem_deepeval-0.3.0/tests/__init__.py +0 -0
- goodmem_deepeval-0.3.0/tests/conftest.py +96 -0
- goodmem_deepeval-0.3.0/tests/fixtures/retrieve_broken_reranker.ndjson +11 -0
- goodmem_deepeval-0.3.0/tests/fixtures/retrieve_empty.ndjson +2 -0
- goodmem_deepeval-0.3.0/tests/fixtures/retrieve_reranked.ndjson +9 -0
- goodmem_deepeval-0.3.0/tests/fixtures/retrieve_vector.ndjson +8 -0
- goodmem_deepeval-0.3.0/tests/fixtures/space.json +39 -0
- goodmem_deepeval-0.3.0/tests/fixtures/spaces_list.json +83 -0
- goodmem_deepeval-0.3.0/tests/test_e2e.py +281 -0
- goodmem_deepeval-0.3.0/tests/test_readme.py +81 -0
- goodmem_deepeval-0.3.0/tests/test_regressions.py +390 -0
- goodmem_deepeval-0.3.0/tests/test_reranker_fallback.py +210 -0
- goodmem_deepeval-0.3.0/tests/test_typed_filters.py +175 -0
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
__pycache__/
|
|
2
|
+
*.py[cod]
|
|
3
|
+
*.egg-info/
|
|
4
|
+
.eggs/
|
|
5
|
+
build/
|
|
6
|
+
dist/
|
|
7
|
+
.pytest_cache/
|
|
8
|
+
.tox/
|
|
9
|
+
.venv/
|
|
10
|
+
.env
|
|
11
|
+
.coverage
|
|
12
|
+
htmlcov/
|
|
13
|
+
.mypy_cache/
|
|
14
|
+
.ruff_cache/
|
|
15
|
+
.DS_Store
|
|
16
|
+
.venv/
|
|
17
|
+
.venv-smoke/
|
|
18
|
+
.ruff_cache/
|
|
19
|
+
.mypy_cache/
|
|
20
|
+
.pytest_cache/
|
|
21
|
+
dist/
|
|
@@ -0,0 +1,171 @@
|
|
|
1
|
+
# Changelog
|
|
2
|
+
|
|
3
|
+
## 0.3.0
|
|
4
|
+
|
|
5
|
+
### Changed
|
|
6
|
+
|
|
7
|
+
- **Breaking:** renamed to `goodmem-deepeval` (import `goodmem_deepeval`),
|
|
8
|
+
the goodmem-<framework> naming used by goodmem-adk and
|
|
9
|
+
goodmem-semantic-kernel. Update imports from `deepeval_goodmem` to
|
|
10
|
+
`goodmem_deepeval`.
|
|
11
|
+
|
|
12
|
+
## 0.2.2
|
|
13
|
+
|
|
14
|
+
Both fixes were reproduced against a live GoodMem server (v1.0.320) on
|
|
15
|
+
0.2.1 and re-measured with the identical probe after the fix.
|
|
16
|
+
|
|
17
|
+
### Fixed
|
|
18
|
+
|
|
19
|
+
- **A failed reranker's fallback was reported as reranked, and `min_score`
|
|
20
|
+
discarded it.** With `reranker_id` set and the reranker failing, the server
|
|
21
|
+
sends `NOT_FOUND` (naming the reranker) and `RERANKING_FAILED` and still
|
|
22
|
+
returns the vector-stage hits (`stageName: "retrieve"`, raw scores
|
|
23
|
+
`-0.785, -0.577, -0.112` live). 0.2.1 decided `score_kind` from
|
|
24
|
+
configuration, so every hit was `"reranker"`, and `min_score` was applied
|
|
25
|
+
to vector distances: live, `min_score=0.0` returned **0 of the 3 hits** the
|
|
26
|
+
server sent, with a warning blaming the reranker's score range. Now the
|
|
27
|
+
response decides, once the whole stream is in: the hits are reranked only
|
|
28
|
+
if a reranker was requested and the server reported neither
|
|
29
|
+
`RERANKING_FAILED` nor a `NOT_FOUND` naming the reranker (`reranker_id` /
|
|
30
|
+
`rerankerId` in its details, or "reranker" in its message). Fallback hits
|
|
31
|
+
are `score_kind: "vector"` with their raw scores, `min_score` is not
|
|
32
|
+
applied to them, and the result stays `partial: true` with its statuses
|
|
33
|
+
(retrieval status contract Q4a). An unrelated status leaves reranker
|
|
34
|
+
scores and the threshold as they were; a threshold that empties genuinely
|
|
35
|
+
reranked hits still warns and names the observed range.
|
|
36
|
+
- **Boolean and float `metadata_filter` values silently matched nothing.**
|
|
37
|
+
Every value became `text_equals(field, str(value))`, so `{"flag": True}` was
|
|
38
|
+
sent as `CAST(val('$.flag') AS TEXT) = 'True'`: live, HTTP 200 and **0
|
|
39
|
+
results** against a memory stored with `flag: true` (`{"flag": False}` the
|
|
40
|
+
same), and `{"n": 5.0}` sent as `'5.0'` missed a stored `5`, all of which
|
|
41
|
+
read as "nothing stored". `None` was sent as the text `'None'`. Values are
|
|
42
|
+
now compared as their own type, ported from langchain-goodmem's filter
|
|
43
|
+
builder: `bool` as `BOOLEAN` (`true`/`false`, checked before `int`),
|
|
44
|
+
`int`/`float` as `NUMERIC` (finite, written as plain decimals), `str` as
|
|
45
|
+
`TEXT` with the same escaping as before. `None` and any other type raise
|
|
46
|
+
`ValueError` before a request instead of being turned into text. Live,
|
|
47
|
+
`{"flag": True}`, `{"flag": False}`, `{"n": 5}`, `{"n": 5.0}` and
|
|
48
|
+
`{"flag": True, "n": 5}` now each return the one matching memory; string
|
|
49
|
+
filters (including `O'Brien`, `a\\b` and the `x' OR '1'='1` injection) send
|
|
50
|
+
the same expression as 0.2.1. Field names are checked with a full match, so
|
|
51
|
+
a trailing newline is refused.
|
|
52
|
+
|
|
53
|
+
### Tests
|
|
54
|
+
|
|
55
|
+
- `tests/test_reranker_fallback.py`: 15 offline tests over the captured
|
|
56
|
+
broken-reranker and reranked streams (reordered, or with one status added
|
|
57
|
+
or removed), decoded by the real SDK. 10 fail on 0.2.1; the 5 controls
|
|
58
|
+
(working reranker, unrelated statuses, threshold warning) pass on both.
|
|
59
|
+
- `tests/test_typed_filters.py`: 36 offline tests driving `search()` through
|
|
60
|
+
the real SDK (the expression on the wire, or a refusal with nothing sent),
|
|
61
|
+
one of them checking the README's filter table against the builder.
|
|
62
|
+
23 fail on 0.2.1; the 13 controls
|
|
63
|
+
(strings and their escaping, control characters, unsafe field names, the
|
|
64
|
+
empty mapping) pass on both.
|
|
65
|
+
- Two live tests: a missing reranker with `min_score=0.0` keeps its 3
|
|
66
|
+
vector hits, and boolean/number filters match. Both fail on 0.2.1.
|
|
67
|
+
- 87 offline tests (was 36); 14 live tests (was 12).
|
|
68
|
+
|
|
69
|
+
## 0.2.1
|
|
70
|
+
|
|
71
|
+
Documentation only; no code change.
|
|
72
|
+
|
|
73
|
+
### Fixed
|
|
74
|
+
|
|
75
|
+
- **README: the headline example did not run.** It built the test case with
|
|
76
|
+
`to_test_case(result, actual_output=answer)` and then evaluated it with
|
|
77
|
+
`ContextualRecallMetric`, which requires `expected_output`. Executed as
|
|
78
|
+
written, `evaluate()` stopped with `MissingTestCaseParamsError: 'expected_output'
|
|
79
|
+
cannot be None for the 'Contextual Recall' metric` (both the first example
|
|
80
|
+
and the health-metric example). The example now passes `expected_output`,
|
|
81
|
+
the health-metric example imports what it uses, and the README says which
|
|
82
|
+
RAG metrics need a reference answer. Measured: 4/6 README Python blocks ran
|
|
83
|
+
before, 6/6 after.
|
|
84
|
+
- README: the `search()` return table now lists the `query` key it returns.
|
|
85
|
+
|
|
86
|
+
### Added
|
|
87
|
+
|
|
88
|
+
- `tests/test_readme.py` executes every README Python block through the real
|
|
89
|
+
SDK over the captured-bytes transport and fails if a metric the README
|
|
90
|
+
passes to `evaluate()` lacks a test-case field it requires. 36 offline
|
|
91
|
+
tests (was 34).
|
|
92
|
+
|
|
93
|
+
## 0.2.0
|
|
94
|
+
|
|
95
|
+
A rewrite. 0.1.0 shipped no DeepEval integration: it was fourteen classes
|
|
96
|
+
wrapping REST endpoints with `httpx`, and `deepeval` was not even a
|
|
97
|
+
dependency. 0.2.0 is what the name promises, built on the official `goodmem`
|
|
98
|
+
SDK.
|
|
99
|
+
|
|
100
|
+
Every claim below was reproduced against a live GoodMem server (v1.0.320)
|
|
101
|
+
before the fix, and the offline tests replay bytes captured from it.
|
|
102
|
+
|
|
103
|
+
### Added
|
|
104
|
+
|
|
105
|
+
- **`GoodMemRetriever`** — `search()` is decorated with
|
|
106
|
+
`@observe(type="retriever")` and reports `top_k`/embedder through
|
|
107
|
+
`update_retriever_span`, so a retrieval is a DeepEval retriever span.
|
|
108
|
+
- **`retrieve()`** — returns `list[str]`, the shape `retrieval_context` wants.
|
|
109
|
+
- **`to_test_case()`** — builds an `LLMTestCase` with `retrieval_context`
|
|
110
|
+
populated and the retrieval's diagnostics in `metadata["goodmem"]`.
|
|
111
|
+
- **`GoodMemRetrievalHealthMetric`** — a real `BaseMetric` that fails a case
|
|
112
|
+
whose retrieval was degraded, so a broken reranker is its own failure
|
|
113
|
+
instead of a mysterious drop in someone else's recall number.
|
|
114
|
+
- 34 offline tests over captured server bytes and 12 live tests with verified
|
|
115
|
+
teardown. 0.1.0 had 0 tests in CI, and no CI.
|
|
116
|
+
|
|
117
|
+
### Fixed
|
|
118
|
+
|
|
119
|
+
| Behaviour | 0.1.0 (reproduced live) | 0.2.0 |
|
|
120
|
+
| --- | --- | --- |
|
|
121
|
+
| Retrieval with a broken reranker | `success: true`, `totalResults: 1` — the server sent `NOT_FOUND` **and** `RERANKING_FAILED` and the NDJSON parser had no `status` branch, so both were dropped. The result was an *unreranked* fallback presented as a successful reranked search | `partial: true` with both statuses; the fallback chunks are kept |
|
|
122
|
+
| Retrieval that failed outright | `success: true, totalResults: 0` with a message blaming indexing | Empty `hits`, `partial: true`, statuses, and a warning — never raised |
|
|
123
|
+
| Searching an empty space | **63.8 seconds** (`wait_for_indexing=True`, `max_wait_seconds=60`, `poll_interval=5` → 12 re-POSTs, each re-invoking the LLM when `llm_id` was set) | **0.3 seconds**, one request |
|
|
124
|
+
| Relevance scores | Documented as "Minimum relevance score (0-1)"; live values were `-0.57` (vector) and `0.17` (reranker) | Reported as sent, in the server's order, with `score_kind`; `min_score` is client-side and reranker-only |
|
|
125
|
+
| `metadata_filter` | A raw string pasted into the request. Live, `x' OR '1'='1` subverted the filter and returned a row it should not have | A dict, quoted for the live filter grammar; control characters refused |
|
|
126
|
+
| Any error | `run()` caught **everything** and returned `{"success": false, "error": "Client error '400 Bad Request' … check https://developer.mozilla.org/…"}` as a *string* — the server's actual reason (`Invalid space ID format: not-a-uuid`) was discarded for an MDN link | Errors propagate with the server's message |
|
|
127
|
+
| `update_space(public_read=…)` | HTTP 400 — the field was removed from the API | Gone |
|
|
128
|
+
| Space name collisions | Not checked | Reuse when the embedder matches; a mismatch or an ambiguous name is an error |
|
|
129
|
+
| API key | A plain attribute on every operation object | A private connection; never in `model_dump()`, `repr()` or a trace |
|
|
130
|
+
|
|
131
|
+
### Migration
|
|
132
|
+
|
|
133
|
+
The fourteen `GoodMem*` operation classes are gone; they were a REST client,
|
|
134
|
+
and the `goodmem` SDK does that job better.
|
|
135
|
+
|
|
136
|
+
```python
|
|
137
|
+
# 0.1.0
|
|
138
|
+
from deepeval_goodmem import GoodMemRetrieveMemories
|
|
139
|
+
out = json.loads(GoodMemRetrieveMemories(...).run(query="…", space_ids=sid))
|
|
140
|
+
for r in out["results"]:
|
|
141
|
+
print(r["chunkText"], r["relevanceScore"])
|
|
142
|
+
|
|
143
|
+
# 0.2.0
|
|
144
|
+
from deepeval_goodmem import GoodMemRetriever
|
|
145
|
+
result = GoodMemRetriever(space_id=sid).search("…")
|
|
146
|
+
for h in result["hits"]:
|
|
147
|
+
print(h["chunk_text"], h["score"], h["score_kind"])
|
|
148
|
+
|
|
149
|
+
# 0.2.0 — for plain API access, use the SDK directly
|
|
150
|
+
from goodmem import Goodmem
|
|
151
|
+
Goodmem(base_url=…, api_key=…).spaces.create(...)
|
|
152
|
+
```
|
|
153
|
+
|
|
154
|
+
| 0.1.0 | 0.2.0 |
|
|
155
|
+
| --- | --- |
|
|
156
|
+
| `GoodMemRetrieveMemories(...).run(query=, space_ids=)` | `GoodMemRetriever(space_ids=[...]).search(query)` |
|
|
157
|
+
| `out["results"][i]["chunkText"]` | `result["hits"][i]["chunk_text"]` |
|
|
158
|
+
| `out["results"][i]["relevanceScore"]` | `result["hits"][i]["score"]` + `["score_kind"]` |
|
|
159
|
+
| `relevance_threshold=` | `min_score=`, with `reranker_id` set |
|
|
160
|
+
| `metadata_filter="CAST(...)"` | `metadata_filter={"field": "value"}` (or `filter=` for an expression) |
|
|
161
|
+
| `GoodMemCreateSpace` / `GoodMemGetSpace` / … | `goodmem.Goodmem().spaces.*` |
|
|
162
|
+
| `GoodMemCreateMemory` / `GoodMemListMemories` / … | `goodmem.Goodmem().memories.*` |
|
|
163
|
+
|
|
164
|
+
Prior art: the in-tree integration in the `deepeval-goodmem` fork of upstream
|
|
165
|
+
DeepEval established the `@observe(type="retriever")` + `update_retriever_span`
|
|
166
|
+
shape, the `list[str]` return for `retrieval_context`, and a
|
|
167
|
+
`wait_for_indexing` default of `False`. All three are carried over here.
|
|
168
|
+
|
|
169
|
+
## 0.1.0
|
|
170
|
+
|
|
171
|
+
Initial release.
|
|
@@ -0,0 +1,201 @@
|
|
|
1
|
+
Apache License
|
|
2
|
+
Version 2.0, January 2004
|
|
3
|
+
http://www.apache.org/licenses/
|
|
4
|
+
|
|
5
|
+
TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
|
|
6
|
+
|
|
7
|
+
1. Definitions.
|
|
8
|
+
|
|
9
|
+
"License" shall mean the terms and conditions for use, reproduction,
|
|
10
|
+
and distribution as defined by Sections 1 through 9 of this document.
|
|
11
|
+
|
|
12
|
+
"Licensor" shall mean the copyright owner or entity authorized by
|
|
13
|
+
the copyright owner that is granting the License.
|
|
14
|
+
|
|
15
|
+
"Legal Entity" shall mean the union of the acting entity and all
|
|
16
|
+
other entities that control, are controlled by, or are under common
|
|
17
|
+
control with that entity. For the purposes of this definition,
|
|
18
|
+
"control" means (i) the power, direct or indirect, to cause the
|
|
19
|
+
direction or management of such entity, whether by contract or
|
|
20
|
+
otherwise, or (ii) ownership of fifty percent (50%) or more of the
|
|
21
|
+
outstanding shares, or (iii) beneficial ownership of such entity.
|
|
22
|
+
|
|
23
|
+
"You" (or "Your") shall mean an individual or Legal Entity
|
|
24
|
+
exercising permissions granted by this License.
|
|
25
|
+
|
|
26
|
+
"Source" form shall mean the preferred form for making modifications,
|
|
27
|
+
including but not limited to software source code, documentation
|
|
28
|
+
source, and configuration files.
|
|
29
|
+
|
|
30
|
+
"Object" form shall mean any form resulting from mechanical
|
|
31
|
+
transformation or translation of a Source form, including but
|
|
32
|
+
not limited to compiled object code, generated documentation,
|
|
33
|
+
and conversions to other media types.
|
|
34
|
+
|
|
35
|
+
"Work" shall mean the work of authorship, whether in Source or
|
|
36
|
+
Object form, made available under the License, as indicated by a
|
|
37
|
+
copyright notice that is included in or attached to the work
|
|
38
|
+
(an example is provided in the Appendix below).
|
|
39
|
+
|
|
40
|
+
"Derivative Works" shall mean any work, whether in Source or Object
|
|
41
|
+
form, that is based on (or derived from) the Work and for which the
|
|
42
|
+
editorial revisions, annotations, elaborations, or other modifications
|
|
43
|
+
represent, as a whole, an original work of authorship. For the purposes
|
|
44
|
+
of this License, Derivative Works shall not include works that remain
|
|
45
|
+
separable from, or merely link (or bind by name) to the interfaces of,
|
|
46
|
+
the Work and Derivative Works thereof.
|
|
47
|
+
|
|
48
|
+
"Contribution" shall mean any work of authorship, including
|
|
49
|
+
the original version of the Work and any modifications or additions
|
|
50
|
+
to that Work or Derivative Works thereof, that is intentionally
|
|
51
|
+
submitted to Licensor for inclusion in the Work by the copyright owner
|
|
52
|
+
or by an individual or Legal Entity authorized to submit on behalf of
|
|
53
|
+
the copyright owner. For the purposes of this definition, "submitted"
|
|
54
|
+
means any form of electronic, verbal, or written communication sent
|
|
55
|
+
to the Licensor or its representatives, including but not limited to
|
|
56
|
+
communication on electronic mailing lists, source code control systems,
|
|
57
|
+
and issue tracking systems that are managed by, or on behalf of, the
|
|
58
|
+
Licensor for the purpose of discussing and improving the Work, but
|
|
59
|
+
excluding communication that is conspicuously marked or otherwise
|
|
60
|
+
designated in writing by the copyright owner as "Not a Contribution."
|
|
61
|
+
|
|
62
|
+
"Contributor" shall mean Licensor and any individual or Legal Entity
|
|
63
|
+
on behalf of whom a Contribution has been received by Licensor and
|
|
64
|
+
subsequently incorporated within the Work.
|
|
65
|
+
|
|
66
|
+
2. Grant of Copyright License. Subject to the terms and conditions of
|
|
67
|
+
this License, each Contributor hereby grants to You a perpetual,
|
|
68
|
+
worldwide, non-exclusive, no-charge, royalty-free, irrevocable
|
|
69
|
+
copyright license to reproduce, prepare Derivative Works of,
|
|
70
|
+
publicly display, publicly perform, sublicense, and distribute the
|
|
71
|
+
Work and such Derivative Works in Source or Object form.
|
|
72
|
+
|
|
73
|
+
3. Grant of Patent License. Subject to the terms and conditions of
|
|
74
|
+
this License, each Contributor hereby grants to You a perpetual,
|
|
75
|
+
worldwide, non-exclusive, no-charge, royalty-free, irrevocable
|
|
76
|
+
(except as stated in this section) patent license to make, have made,
|
|
77
|
+
use, offer to sell, sell, import, and otherwise transfer the Work,
|
|
78
|
+
where such license applies only to those patent claims licensable
|
|
79
|
+
by such Contributor that are necessarily infringed by their
|
|
80
|
+
Contribution(s) alone or by combination of their Contribution(s)
|
|
81
|
+
with the Work to which such Contribution(s) was submitted. If You
|
|
82
|
+
institute patent litigation against any entity (including a
|
|
83
|
+
cross-claim or counterclaim in a lawsuit) alleging that the Work
|
|
84
|
+
or a Contribution incorporated within the Work constitutes direct
|
|
85
|
+
or contributory patent infringement, then any patent licenses
|
|
86
|
+
granted to You under this License for that Work shall terminate
|
|
87
|
+
as of the date such litigation is filed.
|
|
88
|
+
|
|
89
|
+
4. Redistribution. You may reproduce and distribute copies of the
|
|
90
|
+
Work or Derivative Works thereof in any medium, with or without
|
|
91
|
+
modifications, and in Source or Object form, provided that You
|
|
92
|
+
meet the following conditions:
|
|
93
|
+
|
|
94
|
+
(a) You must give any other recipients of the Work or
|
|
95
|
+
Derivative Works a copy of this License; and
|
|
96
|
+
|
|
97
|
+
(b) You must cause any modified files to carry prominent notices
|
|
98
|
+
stating that You changed the files; and
|
|
99
|
+
|
|
100
|
+
(c) You must retain, in the Source form of any Derivative Works
|
|
101
|
+
that You distribute, all copyright, patent, trademark, and
|
|
102
|
+
attribution notices from the Source form of the Work,
|
|
103
|
+
excluding those notices that do not pertain to any part of
|
|
104
|
+
the Derivative Works; and
|
|
105
|
+
|
|
106
|
+
(d) If the Work includes a "NOTICE" text file as part of its
|
|
107
|
+
distribution, then any Derivative Works that You distribute must
|
|
108
|
+
include a readable copy of the attribution notices contained
|
|
109
|
+
within such NOTICE file, excluding those notices that do not
|
|
110
|
+
pertain to any part of the Derivative Works, in at least one
|
|
111
|
+
of the following places: within a NOTICE text file distributed
|
|
112
|
+
as part of the Derivative Works; within the Source form or
|
|
113
|
+
documentation, if provided along with the Derivative Works; or,
|
|
114
|
+
within a display generated by the Derivative Works, if and
|
|
115
|
+
wherever such third-party notices normally appear. The contents
|
|
116
|
+
of the NOTICE file are for informational purposes only and
|
|
117
|
+
do not modify the License. You may add Your own attribution
|
|
118
|
+
notices within Derivative Works that You distribute, alongside
|
|
119
|
+
or as an addendum to the NOTICE text from the Work, provided
|
|
120
|
+
that such additional attribution notices cannot be construed
|
|
121
|
+
as modifying the License.
|
|
122
|
+
|
|
123
|
+
You may add Your own copyright statement to Your modifications and
|
|
124
|
+
may provide additional or different license terms and conditions
|
|
125
|
+
for use, reproduction, or distribution of Your modifications, or
|
|
126
|
+
for any such Derivative Works as a whole, provided Your use,
|
|
127
|
+
reproduction, and distribution of the Work otherwise complies with
|
|
128
|
+
the conditions stated in this License.
|
|
129
|
+
|
|
130
|
+
5. Submission of Contributions. Unless You explicitly state otherwise,
|
|
131
|
+
any Contribution intentionally submitted for inclusion in the Work
|
|
132
|
+
by You to the Licensor shall be under the terms and conditions of
|
|
133
|
+
this License, without any additional terms or conditions.
|
|
134
|
+
Notwithstanding the above, nothing herein shall supersede or modify
|
|
135
|
+
the terms of any separate license agreement you may have executed
|
|
136
|
+
with Licensor regarding such Contributions.
|
|
137
|
+
|
|
138
|
+
6. Trademarks. This License does not grant permission to use the trade
|
|
139
|
+
names, trademarks, service marks, or product names of the Licensor,
|
|
140
|
+
except as required for describing the origin of the Work and
|
|
141
|
+
reproducing the content of the NOTICE file.
|
|
142
|
+
|
|
143
|
+
7. Disclaimer of Warranty. Unless required by applicable law or
|
|
144
|
+
agreed to in writing, Licensor provides the Work (and each
|
|
145
|
+
Contributor provides its Contributions) on an "AS IS" BASIS,
|
|
146
|
+
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
|
|
147
|
+
implied, including, without limitation, any warranties or conditions
|
|
148
|
+
of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
|
|
149
|
+
PARTICULAR PURPOSE. You are solely responsible for determining the
|
|
150
|
+
appropriateness of using or redistributing the Work and assume any
|
|
151
|
+
risks associated with Your exercise of permissions under this License.
|
|
152
|
+
|
|
153
|
+
8. Limitation of Liability. In no event and under no legal theory,
|
|
154
|
+
whether in tort (including negligence), contract, or otherwise,
|
|
155
|
+
unless required by applicable law (such as deliberate and grossly
|
|
156
|
+
negligent acts) or agreed to in writing, shall any Contributor be
|
|
157
|
+
liable to You for damages, including any direct, indirect, special,
|
|
158
|
+
incidental, or consequential damages of any character arising as a
|
|
159
|
+
result of this License or out of the use or inability to use the
|
|
160
|
+
Work (including but not limited to damages for loss of goodwill,
|
|
161
|
+
work stoppage, computer failure or malfunction, or any and all
|
|
162
|
+
other commercial damages or losses), even if such Contributor
|
|
163
|
+
has been advised of the possibility of such damages.
|
|
164
|
+
|
|
165
|
+
9. Accepting Warranty or Support. While redistributing the Work or
|
|
166
|
+
Derivative Works thereof, You may choose to offer, and charge a
|
|
167
|
+
fee for, acceptance of support, warranty, indemnity, or other
|
|
168
|
+
liability obligations and/or rights consistent with this License.
|
|
169
|
+
However, in accepting such obligations, You may act only on Your
|
|
170
|
+
own behalf and on Your sole responsibility, not on behalf of any
|
|
171
|
+
other Contributor, and only if You agree to indemnify, defend,
|
|
172
|
+
and hold each Contributor harmless for any liability incurred by,
|
|
173
|
+
or claims asserted against, such Contributor by reason of your
|
|
174
|
+
accepting any such warranty or support.
|
|
175
|
+
|
|
176
|
+
END OF TERMS AND CONDITIONS
|
|
177
|
+
|
|
178
|
+
APPENDIX: How to apply the Apache License to your work.
|
|
179
|
+
|
|
180
|
+
To apply the Apache License to your work, attach the following
|
|
181
|
+
boilerplate notice, with the fields enclosed by brackets "[]"
|
|
182
|
+
replaced with your own identifying information. (Don't include
|
|
183
|
+
the brackets!) The text should be enclosed in the appropriate
|
|
184
|
+
comment syntax for the file format. We also recommend that a
|
|
185
|
+
file or class name and description of purpose be included on the
|
|
186
|
+
same "printed page" as the copyright notice for easier
|
|
187
|
+
identification within third-party archives.
|
|
188
|
+
|
|
189
|
+
Copyright 2026 GoodMem
|
|
190
|
+
|
|
191
|
+
Licensed under the Apache License, Version 2.0 (the "License");
|
|
192
|
+
you may not use this file except in compliance with the License.
|
|
193
|
+
You may obtain a copy of the License at
|
|
194
|
+
|
|
195
|
+
http://www.apache.org/licenses/LICENSE-2.0
|
|
196
|
+
|
|
197
|
+
Unless required by applicable law or agreed to in writing, software
|
|
198
|
+
distributed under the License is distributed on an "AS IS" BASIS,
|
|
199
|
+
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
|
200
|
+
See the License for the specific language governing permissions and
|
|
201
|
+
limitations under the License.
|
|
@@ -0,0 +1,232 @@
|
|
|
1
|
+
Metadata-Version: 2.5
|
|
2
|
+
Name: goodmem-deepeval
|
|
3
|
+
Version: 0.3.0
|
|
4
|
+
Summary: GoodMem retrieval for DeepEval — a traced retriever, LLMTestCase construction, and a retrieval-health metric.
|
|
5
|
+
Project-URL: homepage, https://github.com/PAIR-Systems-Inc/goodmem_deepeval
|
|
6
|
+
Project-URL: source, https://github.com/PAIR-Systems-Inc/goodmem_deepeval
|
|
7
|
+
Project-URL: issues, https://github.com/PAIR-Systems-Inc/goodmem_deepeval/issues
|
|
8
|
+
Author: PAIR Systems
|
|
9
|
+
License-File: LICENSE
|
|
10
|
+
Keywords: deepeval,evaluation,goodmem,llm,rag,retrieval,tracing
|
|
11
|
+
Classifier: Development Status :: 4 - Beta
|
|
12
|
+
Classifier: Intended Audience :: Developers
|
|
13
|
+
Classifier: Operating System :: OS Independent
|
|
14
|
+
Classifier: Programming Language :: Python :: 3
|
|
15
|
+
Classifier: Programming Language :: Python :: 3.10
|
|
16
|
+
Classifier: Programming Language :: Python :: 3.11
|
|
17
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
18
|
+
Classifier: Programming Language :: Python :: 3.13
|
|
19
|
+
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
|
|
20
|
+
Classifier: Topic :: Software Development :: Libraries :: Python Modules
|
|
21
|
+
Requires-Python: >=3.10
|
|
22
|
+
Requires-Dist: deepeval>=4.0
|
|
23
|
+
Requires-Dist: goodmem>=0.1.34
|
|
24
|
+
Requires-Dist: pydantic>=2.7
|
|
25
|
+
Provides-Extra: dev
|
|
26
|
+
Requires-Dist: httpx>=0.27; extra == 'dev'
|
|
27
|
+
Requires-Dist: mypy==1.18.2; extra == 'dev'
|
|
28
|
+
Requires-Dist: pytest>=7.0; extra == 'dev'
|
|
29
|
+
Requires-Dist: ruff==0.16.8; extra == 'dev'
|
|
30
|
+
Description-Content-Type: text/markdown
|
|
31
|
+
|
|
32
|
+
# goodmem-deepeval
|
|
33
|
+
|
|
34
|
+
GoodMem retrieval for [DeepEval](https://deepeval.com).
|
|
35
|
+
|
|
36
|
+
Every retrieval is a DeepEval **retriever span**, and a retrieval turns
|
|
37
|
+
directly into an `LLMTestCase` with `retrieval_context` populated — which is
|
|
38
|
+
what `ContextualPrecisionMetric`, `ContextualRecallMetric`,
|
|
39
|
+
`ContextualRelevancyMetric` and `FaithfulnessMetric` actually read.
|
|
40
|
+
|
|
41
|
+
```bash
|
|
42
|
+
pip install goodmem-deepeval
|
|
43
|
+
```
|
|
44
|
+
|
|
45
|
+
## Retrieve and evaluate
|
|
46
|
+
|
|
47
|
+
```python
|
|
48
|
+
from deepeval import evaluate
|
|
49
|
+
from deepeval.metrics import ContextualRelevancyMetric, ContextualRecallMetric
|
|
50
|
+
from goodmem_deepeval import GoodMemRetriever, GoodMemRetrievalHealthMetric
|
|
51
|
+
|
|
52
|
+
retriever = GoodMemRetriever(space_name="docs", limit=5) # credentials from the environment
|
|
53
|
+
|
|
54
|
+
result = retriever.search("how do I rotate an API key?")
|
|
55
|
+
answer = my_llm(result["hits"]) # your generation step
|
|
56
|
+
expected = "…" # your reference answer
|
|
57
|
+
|
|
58
|
+
case = retriever.to_test_case(result, actual_output=answer, expected_output=expected)
|
|
59
|
+
|
|
60
|
+
evaluate([case], [
|
|
61
|
+
ContextualRelevancyMetric(),
|
|
62
|
+
ContextualRecallMetric(),
|
|
63
|
+
GoodMemRetrievalHealthMetric(),
|
|
64
|
+
])
|
|
65
|
+
```
|
|
66
|
+
|
|
67
|
+
`ContextualRecallMetric` and `ContextualPrecisionMetric` compare against a
|
|
68
|
+
reference answer, so they need `expected_output`; without it DeepEval's
|
|
69
|
+
`evaluate()` raises `MissingTestCaseParamsError`.
|
|
70
|
+
`ContextualRelevancyMetric` and `FaithfulnessMetric` do not.
|
|
71
|
+
|
|
72
|
+
`search()` is decorated with `@observe(type="retriever")` and reports `top_k`
|
|
73
|
+
and the embedder through `update_retriever_span`, so the call appears as a
|
|
74
|
+
retriever span with its query, latency and results.
|
|
75
|
+
|
|
76
|
+
If you only need the context, `retrieve()` returns the chunk text directly:
|
|
77
|
+
|
|
78
|
+
```python
|
|
79
|
+
context = retriever.retrieve("how do I rotate an API key?") # list[str]
|
|
80
|
+
```
|
|
81
|
+
|
|
82
|
+
### What `search()` returns
|
|
83
|
+
|
|
84
|
+
| Key | What it is |
|
|
85
|
+
| --- | --- |
|
|
86
|
+
| `query` | The query, as passed |
|
|
87
|
+
| `hits` | `chunk_id`, `chunk_text`, `memory_id`, `space_id`, `source`, `score`, `score_kind`, `metadata` — in the server's order |
|
|
88
|
+
| `score_kind` | `"vector"` or `"reranker"`: what the server returned, not what was configured. Different scales; see below |
|
|
89
|
+
| `statuses` | Statuses indicating a real problem, `[]` when clean |
|
|
90
|
+
| `partial` | `True` when the server reported a real problem during this retrieval, with or without hits |
|
|
91
|
+
| `abstract_reply` | The server-generated summary, only when `llm_id` is set |
|
|
92
|
+
| `space_ids` | Which spaces were actually searched |
|
|
93
|
+
|
|
94
|
+
A retrieval that reported a problem and returned nothing is an empty `hits`
|
|
95
|
+
with `partial: True` **and a warning** — never an exception, and never
|
|
96
|
+
indistinguishable from "no matches".
|
|
97
|
+
|
|
98
|
+
## The health metric
|
|
99
|
+
|
|
100
|
+
```python
|
|
101
|
+
from deepeval import evaluate
|
|
102
|
+
from deepeval.metrics import ContextualRecallMetric
|
|
103
|
+
from goodmem_deepeval import GoodMemRetrievalHealthMetric
|
|
104
|
+
|
|
105
|
+
# cases: LLMTestCases built with retriever.to_test_case(...), as above
|
|
106
|
+
evaluate(cases, [ContextualRecallMetric(), GoodMemRetrievalHealthMetric()])
|
|
107
|
+
```
|
|
108
|
+
|
|
109
|
+
Every other RAG metric scores the *content* that came back, so a broken
|
|
110
|
+
reranker and a thin corpus look identical: both are poor recall.
|
|
111
|
+
`GoodMemRetrievalHealthMetric` reads the diagnostics `to_test_case` puts in
|
|
112
|
+
`metadata["goodmem"]` and scores the retrieval itself — `1.0` clean, `0.0`
|
|
113
|
+
degraded, with the server's status codes in `reason` and `score_breakdown`.
|
|
114
|
+
No LLM, so it costs nothing and never flakes. A test case with no GoodMem
|
|
115
|
+
metadata is skipped rather than failed.
|
|
116
|
+
|
|
117
|
+
## Scores
|
|
118
|
+
|
|
119
|
+
GoodMem returns two different things in the same field:
|
|
120
|
+
|
|
121
|
+
| | Range observed live | Best match is |
|
|
122
|
+
| --- | --- | --- |
|
|
123
|
+
| Vector score | negative, e.g. `-0.57` | the **lowest** number |
|
|
124
|
+
| Reranker score | can also go negative | the **highest** number |
|
|
125
|
+
|
|
126
|
+
So results keep **the server's order** and are never re-sorted; `score_kind`
|
|
127
|
+
says which scale you have; and `min_score` is applied client-side, only to
|
|
128
|
+
reranker scores. The server's `relevance_threshold` is never sent.
|
|
129
|
+
|
|
130
|
+
`score_kind` says what the server actually did, not what was configured. When
|
|
131
|
+
`reranker_id` is set but the reranker fails, the server reports
|
|
132
|
+
`RERANKING_FAILED` (and `NOT_FOUND` for a missing reranker) and still returns
|
|
133
|
+
the vector-stage hits. Those hits are `score_kind: "vector"` with their raw
|
|
134
|
+
vector scores, and `min_score` is not applied to them, so a reranker threshold
|
|
135
|
+
cannot discard what the server returned; `partial` is `True` and `statuses`
|
|
136
|
+
carries both codes. (0.2.1 labelled them `"reranker"`, and live, with a
|
|
137
|
+
missing reranker, `min_score=0.0` returned none of the 3 hits the server sent.)
|
|
138
|
+
|
|
139
|
+
Even with a reranker the scale is **model-dependent**: on the same documents
|
|
140
|
+
Voyage `rerank-2.5` scored `0.27..0.93` and Jina `jina-reranker-v3` scored
|
|
141
|
+
`-0.14..0.43`. A `min_score` tuned for one empties the other, so when a
|
|
142
|
+
threshold removes every hit the retriever warns and names the observed range.
|
|
143
|
+
Calibrate `min_score` for the reranker you use; there is no default.
|
|
144
|
+
|
|
145
|
+
## Filtering
|
|
146
|
+
|
|
147
|
+
```python
|
|
148
|
+
retriever = GoodMemRetriever(space_name="docs", metadata_filter={"category": "billing"})
|
|
149
|
+
retriever.search("refunds", metadata_filter={"lang": "en", "archived": False}) # AND-ed per call
|
|
150
|
+
```
|
|
151
|
+
|
|
152
|
+
Each value is compared as its own type; the Python type picks the cast:
|
|
153
|
+
|
|
154
|
+
| `metadata_filter` | Sent to the server |
|
|
155
|
+
| --- | --- |
|
|
156
|
+
| `{"category": "billing"}` | `CAST(val('$.category') AS TEXT) = 'billing'` |
|
|
157
|
+
| `{"archived": False}` | `CAST(val('$.archived') AS BOOLEAN) = false` |
|
|
158
|
+
| `{"year": 2026}` | `CAST(val('$.year') AS NUMERIC) = 2026` |
|
|
159
|
+
| `{"score": 2.5}` | `CAST(val('$.score') AS NUMERIC) = 2.5` |
|
|
160
|
+
|
|
161
|
+
Pass the type your metadata stores. Live (v1.0.320), 0.2.1 sent
|
|
162
|
+
`{"flag": True}` as `AS TEXT = 'True'`, which the server accepts with HTTP 200
|
|
163
|
+
and matches nothing, so a boolean filter looked like "nothing stored"; and
|
|
164
|
+
`{"n": 5.0}` as `'5.0'`, which misses a stored `5`. `AS BOOLEAN` and
|
|
165
|
+
`AS NUMERIC` match both. `None` and any other type (lists, dicts, bytes, NaN,
|
|
166
|
+
infinity) raise `ValueError` before a request rather than being turned into
|
|
167
|
+
text that silently matches nothing.
|
|
168
|
+
|
|
169
|
+
Text values are quoted for the GoodMem filter grammar (backslash escaping,
|
|
170
|
+
verified against a live server; control characters refused). For anything
|
|
171
|
+
more complex, pass an expression directly with `filter=...`.
|
|
172
|
+
|
|
173
|
+
## Attaching to a space by name
|
|
174
|
+
|
|
175
|
+
```python
|
|
176
|
+
GoodMemRetriever(space_name="docs", embedder_id="…") # reuse or fail
|
|
177
|
+
GoodMemRetriever(space_name="docs", embedder_id="…", create_space=True) # or create it
|
|
178
|
+
```
|
|
179
|
+
|
|
180
|
+
Attach-by-name is idempotent reuse: an existing space whose embedder matches
|
|
181
|
+
is reused; a **different** embedder is an error, because retrieving across
|
|
182
|
+
mismatched embedders returns plausible-looking nonsense. An ambiguous name is
|
|
183
|
+
an error too.
|
|
184
|
+
|
|
185
|
+
## Credentials
|
|
186
|
+
|
|
187
|
+
From `GOODMEM_BASE_URL` and `GOODMEM_API_KEY`, or as constructor keywords:
|
|
188
|
+
|
|
189
|
+
```python
|
|
190
|
+
GoodMemRetriever(space_name="docs", base_url="https://goodmem.example.com",
|
|
191
|
+
api_key="gm_…")
|
|
192
|
+
```
|
|
193
|
+
|
|
194
|
+
`verify_ssl` defaults to on and exists for a local server with a self-signed
|
|
195
|
+
certificate only; no example here turns it off.
|
|
196
|
+
|
|
197
|
+
They are deliberately **not** model fields, so they cannot reach a
|
|
198
|
+
`model_dump()`, a `repr()` or a trace payload.
|
|
199
|
+
|
|
200
|
+
## Development
|
|
201
|
+
|
|
202
|
+
```bash
|
|
203
|
+
pip install -e ".[dev]"
|
|
204
|
+
ruff check src tests examples && mypy && pytest -m "not integration"
|
|
205
|
+
```
|
|
206
|
+
|
|
207
|
+
CI runs the same three on Python 3.10–3.13, then builds the sdist and wheel,
|
|
208
|
+
imports the wheel in a clean virtualenv, and fails if a GoodMem API key
|
|
209
|
+
(`gm_` followed by 20 or more lowercase letters or digits) is anywhere in the
|
|
210
|
+
tree.
|
|
211
|
+
|
|
212
|
+
The offline suite replays NDJSON captured from a live GoodMem server
|
|
213
|
+
(v1.0.320) through the real SDK decoders, so the wire format is never
|
|
214
|
+
invented. The live suite needs a server and is skipped without one:
|
|
215
|
+
|
|
216
|
+
```bash
|
|
217
|
+
GOODMEM_BASE_URL=… GOODMEM_API_KEY=… GOODMEM_EMBEDDER_ID=… \
|
|
218
|
+
GOODMEM_RERANKER_ID=… SSL_CERT_FILE=/path/to/local-ca.pem \
|
|
219
|
+
pytest -m integration
|
|
220
|
+
```
|
|
221
|
+
|
|
222
|
+
`GOODMEM_RERANKER_ID` is optional — the working-reranker test skips without
|
|
223
|
+
it. For a local server with a self-signed certificate, point `SSL_CERT_FILE`
|
|
224
|
+
at its CA so TLS verification stays on (`GOODMEM_VERIFY_SSL=0` turns it off).
|
|
225
|
+
`GOODMEM_E2E_SPACE_PREFIX` names the temporary test space (default
|
|
226
|
+
`goodmem-deepeval-e2e-`), so its owner is recognisable on a shared server.
|
|
227
|
+
|
|
228
|
+
There is no default credential anywhere in this repository.
|
|
229
|
+
|
|
230
|
+
## License
|
|
231
|
+
|
|
232
|
+
MIT
|