@lotics/cli 0.152.2 → 0.154.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -72,36 +72,57 @@ verbs respect access + activation, so only docs the caller may use ever surface.
72
72
 
73
73
  1. **`list_knowledge`** — `ls`. The corpus as a folder tree: name, id, folder, and a truncated
74
74
  description. No bodies.
75
- 2. **`grep_knowledge`** — `grep -rn`. Substring match across every readable doc (or one doc, or
76
- one folder), returning **doc, line number, and the matching line**. It runs inside Postgres,
75
+ 2. **`grep_knowledge`** — `grep -rn`. Match across every readable doc (or one doc, or one
76
+ folder), returning **doc, line number, and the matching line**. It runs inside Postgres,
77
77
  so only matching lines cross the wire.
78
78
  3. **`read_knowledge`** — `cat` / `sed -n 'X,Yp'`. Read a doc whole or by line range, to see
79
79
  the context around a hit.
80
80
 
81
- Matching is **substring, not ranked** — there is no index and no tokenizer, which is why a
82
- corpus in any script works and why nothing goes stale. The consequence is that the *agent*
83
- does the narrowing: a distinctive phrase is sharply selective, a whole question matches
84
- everything. `grep_knowledge` always reports `total_matches`, so "too broad" is visible and
85
- cheap to fix.
81
+ Matching is **not ranked** — there is no index and no tokenizer, which is why a corpus in any
82
+ script works and why nothing goes stale. The consequence is that the *agent* does the
83
+ narrowing: a distinctive phrase is sharply selective, a whole question matches everything.
84
+ `grep_knowledge` always reports `total_matches`, so "too broad" is visible and cheap to fix.
85
+
86
+ ### Literal by default, expression on request
87
+
88
+ `pattern` is **literal text** — every character matches itself. Set `regex: true` to read it as
89
+ a POSIX regular expression instead.
90
+
91
+ ```
92
+ grep_knowledge({ pattern: "C/O (Form E)" })
93
+ grep_knowledge({ pattern: "\\yNK\\y", regex: true, case_sensitive: true })
94
+ ```
95
+
96
+ The default is the safe one rather than the conventional one, and the asymmetry is worth
97
+ knowing. As an expression, `C/O (Form E)` searches for `C/O Form E` — not in the document — and
98
+ returns a confident zero; `[CŨ]` becomes a one-character class and matches every `C` in the
99
+ corpus. Both look like ordinary answers. A literal search has no such failure: the worst case
100
+ is an expression sent without the flag, which finds nothing and says so, naming the search that
101
+ ran and what to pass instead.
86
102
 
87
103
  ### Matching options
88
104
 
89
- - **Diacritics fold by default** — `ca phe` matches `cà phê`. `diacritic_insensitive: false` matches
90
- tone marks exactly.
105
+ - **Tone marks match exactly by default.** `diacritic_insensitive: true` folds them, so `ca phe`
106
+ matches `cà phê` — reach for it when the pattern was typed without them, not by habit: folding
107
+ strips the pattern to ASCII, where a short Vietnamese word matches inside unrelated words
108
+ (`mã` finds `manifest`). An empty result tells you when folding would have matched.
91
109
  - **Case folds by default**, independently of diacritics. `case_sensitive: true` matches case
92
110
  exactly — useful for an acronym (`NK` vs `nk`) that a folded search would blur.
93
111
  - **Whitespace is normalized on both sides.** A body converted from PDF, Word or Excel carries
94
112
  non-breaking spaces, soft hyphens, zero-width marks and padded runs that nobody types into a
95
113
  query; those fold to ordinary single spaces before matching, so a correct search does not return
96
- a silent zero on text that is present. Your pattern is normalized the same way.
97
- - **`regex: true`** treats the pattern as a POSIX regular expression: quantifiers, character
98
- classes, alternation, anchors, and **`\y` for a word boundary**. Note `\y`, not `\b` — Postgres
99
- spells it differently, and `\b` is rewritten for you rather than silently matching nothing.
100
- Literal is the default on purpose: `0901.11.20` as a regex would also match `0901X11Y20`.
101
-
102
- ```
103
- grep_knowledge({ pattern: "\\yNK\\y", regex: true, case_sensitive: true })
104
- ```
114
+ a silent zero on text that is present. A literal pattern is normalized the same way; an
115
+ expression is left byte-exact, since collapsing its whitespace would rewrite it.
116
+ - **What `regex: true` supports**: quantifiers, character classes, alternation, anchors,
117
+ backreferences, non-greedy forms, `(?i)`, and both lookahead and lookbehind. A word boundary
118
+ is **`\y`**, not `\b` — Postgres spells it differently, and `\b` is rewritten for you rather
119
+ than silently matching nothing. Named groups and `\p{…}` are likewise rewritten to their POSIX
120
+ equivalents. A `\p{…}` with no POSIX equivalent, and a malformed expression, are both refused
121
+ with the reason.
122
+
123
+ Where an expression cannot express the question at all — a value that must be computed, a
124
+ layout matching cannot address — stage the doc into a code run instead (`code_exec` with
125
+ `knowledge_doc_ids` puts it at `inputs/knowledge/<id>.md` as a real file).
105
126
 
106
127
  ### What a result is bounded by
107
128
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@lotics/cli",
3
- "version": "0.152.2",
3
+ "version": "0.154.0",
4
4
  "description": "Lotics SDK and CLI for AI agents",
5
5
  "type": "module",
6
6
  "bin": {