sessionmemory 0.6.0__tar.gz → 0.6.1__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (36) hide show
  1. {sessionmemory-0.6.0 → sessionmemory-0.6.1}/PKG-INFO +12 -12
  2. {sessionmemory-0.6.0 → sessionmemory-0.6.1}/README.md +10 -10
  3. {sessionmemory-0.6.0 → sessionmemory-0.6.1}/pyproject.toml +11 -11
  4. {sessionmemory-0.6.0 → sessionmemory-0.6.1}/pyproject.toml.orig +11 -11
  5. {sessionmemory-0.6.0 → sessionmemory-0.6.1}/src/sessionmemory/commands/search.py +16 -5
  6. {sessionmemory-0.6.0 → sessionmemory-0.6.1}/src/sessionmemory/lib/fieldindex.py +35 -11
  7. {sessionmemory-0.6.0 → sessionmemory-0.6.1}/src/sessionmemory/lib/inject.py +6 -3
  8. {sessionmemory-0.6.0 → sessionmemory-0.6.1}/src/sessionmemory/__init__.py +0 -0
  9. {sessionmemory-0.6.0 → sessionmemory-0.6.1}/src/sessionmemory/cli.py +0 -0
  10. {sessionmemory-0.6.0 → sessionmemory-0.6.1}/src/sessionmemory/commands/__init__.py +0 -0
  11. {sessionmemory-0.6.0 → sessionmemory-0.6.1}/src/sessionmemory/commands/_common.py +0 -0
  12. {sessionmemory-0.6.0 → sessionmemory-0.6.1}/src/sessionmemory/commands/delete.py +0 -0
  13. {sessionmemory-0.6.0 → sessionmemory-0.6.1}/src/sessionmemory/commands/doctor.py +0 -0
  14. {sessionmemory-0.6.0 → sessionmemory-0.6.1}/src/sessionmemory/commands/export.py +0 -0
  15. {sessionmemory-0.6.0 → sessionmemory-0.6.1}/src/sessionmemory/commands/init.py +0 -0
  16. {sessionmemory-0.6.0 → sessionmemory-0.6.1}/src/sessionmemory/commands/inject.py +0 -0
  17. {sessionmemory-0.6.0 → sessionmemory-0.6.1}/src/sessionmemory/commands/log.py +0 -0
  18. {sessionmemory-0.6.0 → sessionmemory-0.6.1}/src/sessionmemory/commands/new.py +0 -0
  19. {sessionmemory-0.6.0 → sessionmemory-0.6.1}/src/sessionmemory/commands/project.py +0 -0
  20. {sessionmemory-0.6.0 → sessionmemory-0.6.1}/src/sessionmemory/commands/reindex.py +0 -0
  21. {sessionmemory-0.6.0 → sessionmemory-0.6.1}/src/sessionmemory/lib/__init__.py +0 -0
  22. {sessionmemory-0.6.0 → sessionmemory-0.6.1}/src/sessionmemory/lib/atomic.py +0 -0
  23. {sessionmemory-0.6.0 → sessionmemory-0.6.1}/src/sessionmemory/lib/backlog.py +0 -0
  24. {sessionmemory-0.6.0 → sessionmemory-0.6.1}/src/sessionmemory/lib/bootstrap.py +0 -0
  25. {sessionmemory-0.6.0 → sessionmemory-0.6.1}/src/sessionmemory/lib/config.py +0 -0
  26. {sessionmemory-0.6.0 → sessionmemory-0.6.1}/src/sessionmemory/lib/doctor.py +0 -0
  27. {sessionmemory-0.6.0 → sessionmemory-0.6.1}/src/sessionmemory/lib/embed.py +0 -0
  28. {sessionmemory-0.6.0 → sessionmemory-0.6.1}/src/sessionmemory/lib/export.py +0 -0
  29. {sessionmemory-0.6.0 → sessionmemory-0.6.1}/src/sessionmemory/lib/field.py +0 -0
  30. {sessionmemory-0.6.0 → sessionmemory-0.6.1}/src/sessionmemory/lib/frontmatter.py +0 -0
  31. {sessionmemory-0.6.0 → sessionmemory-0.6.1}/src/sessionmemory/lib/gitinfo.py +0 -0
  32. {sessionmemory-0.6.0 → sessionmemory-0.6.1}/src/sessionmemory/lib/ids.py +0 -0
  33. {sessionmemory-0.6.0 → sessionmemory-0.6.1}/src/sessionmemory/lib/log.py +0 -0
  34. {sessionmemory-0.6.0 → sessionmemory-0.6.1}/src/sessionmemory/lib/paths.py +0 -0
  35. {sessionmemory-0.6.0 → sessionmemory-0.6.1}/src/sessionmemory/lib/registry.py +0 -0
  36. {sessionmemory-0.6.0 → sessionmemory-0.6.1}/src/sessionmemory/lib/resolve.py +0 -0
@@ -1,10 +1,10 @@
1
1
  Metadata-Version: 2.3
2
2
  Name: sessionmemory
3
- Version: 0.6.0
3
+ Version: 0.6.1
4
4
  Summary: Durable memory for coding agents, one folder of searchable pages per project.
5
5
  Author: Nathaniel Landau
6
6
  Author-email: Nathaniel Landau <github@natelandau.com>
7
- Requires-Dist: fastembed>=0.8.0
7
+ Requires-Dist: fastembed>=0.8.1
8
8
  Requires-Dist: nclutils>=3.4.4
9
9
  Requires-Dist: pyyaml>=6.0.3
10
10
  Requires-Dist: sqlite-vec>=0.1.9
@@ -159,22 +159,19 @@ The CLI does two things. It finds pages by meaning, and it creates pages. Readin
159
159
  editing a page is a job for your editor or your agent's own tools.
160
160
 
161
161
  ```bash
162
- sessionmemory search "why does the same stripe event arrive twice" --limit 2 --cwd .
162
+ sessionmemory search "stripe event delivered twice" --cwd .
163
163
  ```
164
164
 
165
165
  ```
166
166
  ~/repos/my-vault/projects/invoice-api/learnings/stripe-retries-a-webhook-for-72-hours-so-the-handler-must-be-idempotent.md
167
167
  Stripe retries a webhook for 72 hours, so the handler must be idempotent
168
168
  Stripe redelivers an unacknowledged webhook for up to 72 hours, so the handler records the event id and ignores a repeat.
169
-
170
- ~/repos/my-vault/projects/invoice-api/learnings/the-nightly-reconciliation-job-must-start-after-the-02-00-bank-feed.md
171
- The nightly reconciliation job must start after the 02:00 bank feed
172
- The bank feed lands at 02:00 UTC; a reconciliation run before it reports every open invoice as unpaid.
173
169
  ```
174
170
 
175
171
  A result is a path, a title, and a summary. A paraphrase finds the page, because search
176
- ranks by meaning and not by words in common. A query that nothing answers returns no
177
- results rather than the nearest pages. Pass `--read` to print every hit in full.
172
+ ranks by meaning and not by words in common. A hit has to stand out from the rest of the
173
+ project's pages, so a query that nothing answers returns no results rather than the
174
+ nearest pages. Pass `--read` to print every hit in full.
178
175
 
179
176
  ```bash
180
177
  sessionmemory new learning \
@@ -210,9 +207,12 @@ away. The project's folder has `learnings/` and `logs/`, searched by meaning, be
210
207
 
211
208
  - Before assuming nothing was written down, search: `sessionmemory search "<words>"`
212
209
  prints each hit's path, title, and summary, and `--read` prints every hit's whole
213
- page in one call. A paraphrase still matches. No hits means nothing is recorded,
214
- not that the query needs loosening.
215
- - Past sessions, one page each: `sessionmemory search "<words>" --logs`.
210
+ page in one call. Search with a few distinctive words, such as names, identifiers,
211
+ or error text, rather than a sentence. No hits means nothing is recorded, not that
212
+ the query needs loosening.
213
+ - Past sessions, one page each: `sessionmemory search "<words>" --logs`. Search them
214
+ for why something was decided, what happened the last time an area changed, or
215
+ whether a fix was already tried.
216
216
  - Open work: read `backlog.md`. An item is one line under a `## <kind>` heading
217
217
  (feat, fix, refactor, perf, docs, test, build, ci), sized S, M, or L:
218
218
  `- [S] <imperative description> - <YYYY-MM-DD> [#topic]`. Add one with
@@ -144,22 +144,19 @@ The CLI does two things. It finds pages by meaning, and it creates pages. Readin
144
144
  editing a page is a job for your editor or your agent's own tools.
145
145
 
146
146
  ```bash
147
- sessionmemory search "why does the same stripe event arrive twice" --limit 2 --cwd .
147
+ sessionmemory search "stripe event delivered twice" --cwd .
148
148
  ```
149
149
 
150
150
  ```
151
151
  ~/repos/my-vault/projects/invoice-api/learnings/stripe-retries-a-webhook-for-72-hours-so-the-handler-must-be-idempotent.md
152
152
  Stripe retries a webhook for 72 hours, so the handler must be idempotent
153
153
  Stripe redelivers an unacknowledged webhook for up to 72 hours, so the handler records the event id and ignores a repeat.
154
-
155
- ~/repos/my-vault/projects/invoice-api/learnings/the-nightly-reconciliation-job-must-start-after-the-02-00-bank-feed.md
156
- The nightly reconciliation job must start after the 02:00 bank feed
157
- The bank feed lands at 02:00 UTC; a reconciliation run before it reports every open invoice as unpaid.
158
154
  ```
159
155
 
160
156
  A result is a path, a title, and a summary. A paraphrase finds the page, because search
161
- ranks by meaning and not by words in common. A query that nothing answers returns no
162
- results rather than the nearest pages. Pass `--read` to print every hit in full.
157
+ ranks by meaning and not by words in common. A hit has to stand out from the rest of the
158
+ project's pages, so a query that nothing answers returns no results rather than the
159
+ nearest pages. Pass `--read` to print every hit in full.
163
160
 
164
161
  ```bash
165
162
  sessionmemory new learning \
@@ -195,9 +192,12 @@ away. The project's folder has `learnings/` and `logs/`, searched by meaning, be
195
192
 
196
193
  - Before assuming nothing was written down, search: `sessionmemory search "<words>"`
197
194
  prints each hit's path, title, and summary, and `--read` prints every hit's whole
198
- page in one call. A paraphrase still matches. No hits means nothing is recorded,
199
- not that the query needs loosening.
200
- - Past sessions, one page each: `sessionmemory search "<words>" --logs`.
195
+ page in one call. Search with a few distinctive words, such as names, identifiers,
196
+ or error text, rather than a sentence. No hits means nothing is recorded, not that
197
+ the query needs loosening.
198
+ - Past sessions, one page each: `sessionmemory search "<words>" --logs`. Search them
199
+ for why something was decided, what happened the last time an area changed, or
200
+ whether a fix was already tried.
201
201
  - Open work: read `backlog.md`. An item is one line under a `## <kind>` heading
202
202
  (feat, fix, refactor, perf, docs, test, build, ci), sized S, M, or L:
203
203
  `- [S] <imperative description> - <YYYY-MM-DD> [#topic]`. Add one with
@@ -1,6 +1,6 @@
1
1
  [project]
2
2
  dependencies = [
3
- "fastembed>=0.8.0",
3
+ "fastembed>=0.8.1",
4
4
  "nclutils>=3.4.4",
5
5
  "pyyaml>=6.0.3",
6
6
  "sqlite-vec>=0.1.9",
@@ -11,7 +11,7 @@ description = "Durable memory for coding agents, one folder of searchable pages
11
11
  name = "sessionmemory"
12
12
  readme = "README.md"
13
13
  requires-python = ">=3.13,<3.15"
14
- version = "0.6.0"
14
+ version = "0.6.1"
15
15
 
16
16
  [[project.authors]]
17
17
  name = "Nathaniel Landau"
@@ -22,22 +22,22 @@ sessionmemory = "sessionmemory.cli:main"
22
22
 
23
23
  [dependency-groups]
24
24
  dev = [
25
- "commitizen>=4.18.0",
26
- "coverage>=7.16.0",
27
- "duty>=1.9.0",
28
- "prek>=0.5.2",
25
+ "commitizen>=4.19.0",
26
+ "coverage>=7.16.2",
27
+ "duty>=1.10.0",
28
+ "prek>=0.5.4",
29
29
  "pytest-clarity>=1.0.1",
30
30
  "pytest-cov>=7.1.0",
31
31
  "pytest-devtools>=1.3.0",
32
- "pytest-mock>=3.15.1",
32
+ "pytest-mock>=3.16.0",
33
33
  "pytest-xdist>=3.8.0",
34
34
  "pytest>=9.1.1",
35
35
  "rich>=15.0.0",
36
- "ruff>=0.16.5",
36
+ "ruff>=0.16.10",
37
37
  "shellcheck-py>=0.11.0.1",
38
- "ty>=0.0.78",
39
- "types-pyyaml>=6.0.12.20260815",
40
- "typos>=1.50.1",
38
+ "ty>=0.0.84",
39
+ "types-pyyaml>=6.0.12.20260906",
40
+ "typos>=1.50.3",
41
41
  "yamllint>=1.38.0",
42
42
  ]
43
43
 
@@ -1,7 +1,7 @@
1
1
  [project]
2
2
  authors = [{ name = "Nathaniel Landau", email = "github@natelandau.com" }]
3
3
  dependencies = [
4
- "fastembed>=0.8.0",
4
+ "fastembed>=0.8.1",
5
5
  "nclutils>=3.4.4",
6
6
  "pyyaml>=6.0.3",
7
7
  "sqlite-vec>=0.1.9",
@@ -12,29 +12,29 @@
12
12
  name = "sessionmemory"
13
13
  readme = "README.md"
14
14
  requires-python = ">=3.13,<3.15"
15
- version = "0.6.0"
15
+ version = "0.6.1"
16
16
 
17
17
  [project.scripts]
18
18
  sessionmemory = "sessionmemory.cli:main"
19
19
 
20
20
  [dependency-groups]
21
21
  dev = [
22
- "commitizen>=4.18.0",
23
- "coverage>=7.16.0",
24
- "duty>=1.9.0",
25
- "prek>=0.5.2",
22
+ "commitizen>=4.19.0",
23
+ "coverage>=7.16.2",
24
+ "duty>=1.10.0",
25
+ "prek>=0.5.4",
26
26
  "pytest-clarity>=1.0.1",
27
27
  "pytest-cov>=7.1.0",
28
28
  "pytest-devtools>=1.3.0",
29
- "pytest-mock>=3.15.1",
29
+ "pytest-mock>=3.16.0",
30
30
  "pytest-xdist>=3.8.0",
31
31
  "pytest>=9.1.1",
32
32
  "rich>=15.0.0",
33
- "ruff>=0.16.5",
33
+ "ruff>=0.16.10",
34
34
  "shellcheck-py>=0.11.0.1",
35
- "ty>=0.0.78",
36
- "types-pyyaml>=6.0.12.20260815",
37
- "typos>=1.50.1",
35
+ "ty>=0.0.84",
36
+ "types-pyyaml>=6.0.12.20260906",
37
+ "typos>=1.50.3",
38
38
  "yamllint>=1.38.0",
39
39
  ]
40
40
 
@@ -27,17 +27,25 @@ MAX_DISTANCE = typer.Option(
27
27
  max=2.0,
28
28
  help="Farthest cosine distance that still counts as a hit.",
29
29
  )
30
+ MIN_MARGIN = typer.Option(
31
+ fieldindex.DEFAULT_MIN_MARGIN,
32
+ "--min-margin",
33
+ min=0.0,
34
+ max=2.0,
35
+ help="How much nearer than the field's median page a hit must sit. 0 turns this off.",
36
+ )
30
37
  READ = typer.Option(False, "--read", help="Print each hit's whole file under its path.") # noqa: FBT003
31
38
  CWD = typer.Option(None, "--cwd", help="Directory to resolve the project from.")
32
39
  JSON = typer.Option(False, "--json", help="Emit JSON instead of prose.") # noqa: FBT003
33
40
 
34
41
 
35
- def search_command(
42
+ def search_command( # noqa: PLR0913
36
43
  query: str = QUERY,
37
44
  *,
38
45
  logs: bool = LOGS,
39
46
  limit: int = LIMIT,
40
47
  max_distance: float = MAX_DISTANCE,
48
+ min_margin: float = MIN_MARGIN,
41
49
  read: bool = READ,
42
50
  cwd: Path | None = CWD,
43
51
  as_json: bool = JSON,
@@ -49,7 +57,12 @@ def search_command(
49
57
  slug = require_project(vault, cwd)
50
58
  directory = paths.logs_dir(vault, slug) if logs else paths.learnings_dir(vault, slug)
51
59
  hits = fieldindex.search(
52
- directory, build_embedder(), query, limit=limit, max_distance=max_distance
60
+ directory,
61
+ build_embedder(),
62
+ query,
63
+ limit=limit,
64
+ max_distance=max_distance,
65
+ min_margin=min_margin,
53
66
  )
54
67
 
55
68
  if as_json:
@@ -67,9 +80,7 @@ def search_command(
67
80
  emit_json(payload)
68
81
  return
69
82
  if not hits:
70
- pp.info(
71
- f"no results within distance {max_distance}; raise --max-distance to see farther pages"
72
- )
83
+ pp.info("no results: nothing recorded matches this query")
73
84
  return
74
85
  # A path, a title, a summary, and a page are all things a caller copies or parses,
75
86
  # so nothing here may be styled.
@@ -14,6 +14,7 @@ import datetime
14
14
  import hashlib
15
15
  import json
16
16
  import sqlite3
17
+ import statistics
17
18
  from dataclasses import dataclass
18
19
  from typing import TYPE_CHECKING
19
20
 
@@ -26,11 +27,22 @@ if TYPE_CHECKING:
26
27
 
27
28
  from sessionmemory.lib.embed import Embedder
28
29
 
29
- # Measured on a real vault with nomic-embed-text-v1.5: a page that answers the query sits
30
- # under 0.25, a related neighbor under 0.40, and the nearest page to an unrelated query
31
- # sits at 0.45 or beyond. It matches the reference implementation's default for the model.
30
+ # The reference implementation's default cutoff for nomic-embed-text-v1.5. On its own it
31
+ # admits most unrelated queries, so it is only a ceiling and the margin below decides.
32
32
  DEFAULT_MAX_DISTANCE = 0.45
33
33
 
34
+ # Measured with nomic-embed-text-v1.5 on 120 labeled queries across four projects' learnings
35
+ # and logs: a page that answers the query sits at least this much nearer than the field's
36
+ # median page, and the nearest page to an unrelated query does not. The rule is relative
37
+ # because phrasing a query as a question lowers every distance at once, and because a log,
38
+ # which summarizes a whole session, sits near every query about its project.
39
+ DEFAULT_MIN_MARGIN = 0.11
40
+
41
+ # A median over fewer pages than this is noise, so a small field measures against the
42
+ # median background seen across the labeled queries instead.
43
+ MIN_BACKGROUND_PAGES = 8
44
+ FALLBACK_BACKGROUND = 0.49
45
+
34
46
  _SCHEMA = """
35
47
  CREATE TABLE IF NOT EXISTS pages (
36
48
  filename TEXT PRIMARY KEY,
@@ -165,11 +177,14 @@ def search(
165
177
  *,
166
178
  limit: int,
167
179
  max_distance: float = DEFAULT_MAX_DISTANCE,
180
+ min_margin: float = DEFAULT_MIN_MARGIN,
168
181
  ) -> list[Hit]:
169
- """Return the pages within `max_distance` of `query`, nearest first, refreshing the index first.
182
+ """Return the pages that stand out as nearest to `query`, nearest first, refreshing the index first.
170
183
 
171
- A cutoff rather than a bare top-k, so a query nothing answers returns nothing instead
172
- of the nearest pages dressed up as hits.
184
+ A hit sits within `max_distance` and at least `min_margin` nearer than the field's
185
+ median page. A cutoff rather than a bare top-k, so a query nothing answers returns
186
+ nothing instead of the nearest pages dressed up as hits. A `min_margin` of 0 leaves
187
+ only the absolute cutoff.
173
188
  """
174
189
  if not field_dir.is_dir():
175
190
  return []
@@ -177,16 +192,25 @@ def search(
177
192
  try:
178
193
  _refresh(conn, field_dir, embedder)
179
194
  rows = conn.execute(
180
- "SELECT * FROM ("
181
- " SELECT filename, frontmatter, vec_distance_cosine(embedding, ?) AS distance"
182
- " FROM pages)"
183
- " WHERE distance <= ? ORDER BY distance LIMIT ?",
184
- (sqlite_vec.serialize_float32(embedder.encode_query(query)), max_distance, limit),
195
+ "SELECT filename, frontmatter, vec_distance_cosine(embedding, ?) AS distance"
196
+ " FROM pages ORDER BY distance",
197
+ (sqlite_vec.serialize_float32(embedder.encode_query(query)),),
185
198
  ).fetchall()
186
199
  finally:
187
200
  conn.close()
201
+ ceiling = max_distance
202
+ if min_margin > 0:
203
+ distances = [float(row["distance"]) for row in rows]
204
+ background = (
205
+ statistics.median(distances)
206
+ if len(distances) >= MIN_BACKGROUND_PAGES
207
+ else FALLBACK_BACKGROUND
208
+ )
209
+ ceiling = min(ceiling, background - min_margin)
188
210
  hits = []
189
211
  for row in rows:
212
+ if float(row["distance"]) > ceiling or len(hits) == limit:
213
+ break
190
214
  meta = json.loads(row["frontmatter"])
191
215
  title = meta.get("title")
192
216
  summary = meta.get("summary")
@@ -73,9 +73,12 @@ away. The project's folder has `learnings/` and `logs/`, searched by meaning, be
73
73
 
74
74
  - Before assuming nothing was written down, search: `{command} search "<words>"`
75
75
  prints each hit's path, title, and summary, and `--read` prints every hit's whole
76
- page in one call. A paraphrase still matches. No hits means nothing is recorded,
77
- not that the query needs loosening.
78
- - Past sessions, one page each: `{command} search "<words>" --logs`.
76
+ page in one call. Search with a few distinctive words, such as names, identifiers,
77
+ or error text, rather than a sentence. No hits means nothing is recorded, not that
78
+ the query needs loosening.
79
+ - Past sessions, one page each: `{command} search "<words>" --logs`. Search them
80
+ for why something was decided, what happened the last time an area changed, or
81
+ whether a fix was already tried.
79
82
  - Open work: read `backlog.md`. An item is one line under a `## <kind>` heading
80
83
  (feat, fix, refactor, perf, docs, test, build, ci), sized S, M, or L:
81
84
  `- [S] <imperative description> - <YYYY-MM-DD> [#topic]`. Add one with