@lotics/cli 0.152.2 → 0.153.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/docs/knowledge_docs.md +35 -16
- package/package.json +1 -1
package/docs/knowledge_docs.md
CHANGED
|
@@ -72,17 +72,33 @@ verbs respect access + activation, so only docs the caller may use ever surface.
|
|
|
72
72
|
|
|
73
73
|
1. **`list_knowledge`** — `ls`. The corpus as a folder tree: name, id, folder, and a truncated
|
|
74
74
|
description. No bodies.
|
|
75
|
-
2. **`grep_knowledge`** — `grep -rn`.
|
|
76
|
-
|
|
75
|
+
2. **`grep_knowledge`** — `grep -rn`. Match across every readable doc (or one doc, or one
|
|
76
|
+
folder), returning **doc, line number, and the matching line**. It runs inside Postgres,
|
|
77
77
|
so only matching lines cross the wire.
|
|
78
78
|
3. **`read_knowledge`** — `cat` / `sed -n 'X,Yp'`. Read a doc whole or by line range, to see
|
|
79
79
|
the context around a hit.
|
|
80
80
|
|
|
81
|
-
Matching is **
|
|
82
|
-
|
|
83
|
-
|
|
84
|
-
|
|
85
|
-
|
|
81
|
+
Matching is **not ranked** — there is no index and no tokenizer, which is why a corpus in any
|
|
82
|
+
script works and why nothing goes stale. The consequence is that the *agent* does the
|
|
83
|
+
narrowing: a distinctive phrase is sharply selective, a whole question matches everything.
|
|
84
|
+
`grep_knowledge` always reports `total_matches`, so "too broad" is visible and cheap to fix.
|
|
85
|
+
|
|
86
|
+
### Literal by default, expression on request
|
|
87
|
+
|
|
88
|
+
`pattern` is **literal text** — every character matches itself. Set `regex: true` to read it as
|
|
89
|
+
a POSIX regular expression instead.
|
|
90
|
+
|
|
91
|
+
```
|
|
92
|
+
grep_knowledge({ pattern: "C/O (Form E)" })
|
|
93
|
+
grep_knowledge({ pattern: "\\yNK\\y", regex: true, case_sensitive: true })
|
|
94
|
+
```
|
|
95
|
+
|
|
96
|
+
The default is the safe one rather than the conventional one, and the asymmetry is worth
|
|
97
|
+
knowing. As an expression, `C/O (Form E)` searches for `C/O Form E` — not in the document — and
|
|
98
|
+
returns a confident zero; `[CŨ]` becomes a one-character class and matches every `C` in the
|
|
99
|
+
corpus. Both look like ordinary answers. A literal search has no such failure: the worst case
|
|
100
|
+
is an expression sent without the flag, which finds nothing and says so, naming the search that
|
|
101
|
+
ran and what to pass instead.
|
|
86
102
|
|
|
87
103
|
### Matching options
|
|
88
104
|
|
|
@@ -93,15 +109,18 @@ cheap to fix.
|
|
|
93
109
|
- **Whitespace is normalized on both sides.** A body converted from PDF, Word or Excel carries
|
|
94
110
|
non-breaking spaces, soft hyphens, zero-width marks and padded runs that nobody types into a
|
|
95
111
|
query; those fold to ordinary single spaces before matching, so a correct search does not return
|
|
96
|
-
a silent zero on text that is present.
|
|
97
|
-
|
|
98
|
-
|
|
99
|
-
|
|
100
|
-
|
|
101
|
-
|
|
102
|
-
|
|
103
|
-
|
|
104
|
-
|
|
112
|
+
a silent zero on text that is present. A literal pattern is normalized the same way; an
|
|
113
|
+
expression is left byte-exact, since collapsing its whitespace would rewrite it.
|
|
114
|
+
- **What `regex: true` supports**: quantifiers, character classes, alternation, anchors,
|
|
115
|
+
backreferences, non-greedy forms, `(?i)`, and both lookahead and lookbehind. A word boundary
|
|
116
|
+
is **`\y`**, not `\b` — Postgres spells it differently, and `\b` is rewritten for you rather
|
|
117
|
+
than silently matching nothing. Named groups and `\p{…}` are likewise rewritten to their POSIX
|
|
118
|
+
equivalents. A `\p{…}` with no POSIX equivalent, and a malformed expression, are both refused
|
|
119
|
+
with the reason.
|
|
120
|
+
|
|
121
|
+
Where an expression cannot express the question at all — a value that must be computed, a
|
|
122
|
+
layout matching cannot address — stage the doc into a code run instead (`code_exec` with
|
|
123
|
+
`knowledge_doc_ids` puts it at `inputs/knowledge/<id>.md` as a real file).
|
|
105
124
|
|
|
106
125
|
### What a result is bounded by
|
|
107
126
|
|