@lotics/cli 0.98.0 → 0.99.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/src/cli.js +3 -0
- package/docs/knowledge_docs.md +83 -27
- package/package.json +1 -1
package/dist/src/cli.js
CHANGED
|
@@ -64486,6 +64486,9 @@ var queryOutputColumnSchema = zod_default.object({
|
|
|
64486
64486
|
),
|
|
64487
64487
|
source_field_key: zod_default.string().optional().describe(
|
|
64488
64488
|
"Source field this column originates from. Populated alongside source_table_id for passthrough columns."
|
|
64489
|
+
),
|
|
64490
|
+
computed: zod_default.boolean().optional().describe(
|
|
64491
|
+
"True when the value is DERIVED rather than a passthrough of the addressed field \u2014 a link extraction reads a field on the link TARGET, so the row's own source record is a different table. Such a column may still carry source addressing (it is what resolves labels and option sets), but that addressing is NOT a write target: writable_target is refused on a computed column and on anything that passes one through."
|
|
64489
64492
|
)
|
|
64490
64493
|
});
|
|
64491
64494
|
var queryOutputSchemaSchema = zod_default.object({
|
package/docs/knowledge_docs.md
CHANGED
|
@@ -7,8 +7,8 @@ them on demand** while it works, pulling in only the lines it needs.
|
|
|
7
7
|
|
|
8
8
|
They are deliberately **not** injected into the agent wholesale. Bulk-loading every doc into
|
|
9
9
|
every request would burn the context budget and drown the signal. Instead the agent retrieves
|
|
10
|
-
from them
|
|
11
|
-
actually touches it.
|
|
10
|
+
from them line by line — grep, then read — so a 10,000-line tariff book costs nothing until a
|
|
11
|
+
question actually touches it, and a 10MB one is no different.
|
|
12
12
|
|
|
13
13
|
This is a **capability + usage** guide. For the exact input schema of any tool named here, run
|
|
14
14
|
`lotics tools <tool_name>`.
|
|
@@ -34,9 +34,9 @@ This reads the body from your filesystem and creates the doc through `create_kno
|
|
|
34
34
|
`content` from a file or stdin. Either way, `create_knowledge` takes three things:
|
|
35
35
|
|
|
36
36
|
- `name` — what it is.
|
|
37
|
-
- `description` — what the doc is **and
|
|
38
|
-
|
|
39
|
-
|
|
37
|
+
- `description` — what the doc is **and what to grep it for**: its vocabulary and the synonyms
|
|
38
|
+
an ambiguous query would use. Surfaced in the tree before any body is read, and truncated at
|
|
39
|
+
200 characters. See *Write for grep* below.
|
|
40
40
|
- `content` — Markdown.
|
|
41
41
|
|
|
42
42
|
A new doc is created **owned by you, active in your own agent context, and private** — no one
|
|
@@ -61,28 +61,84 @@ So: shared + active → the agent can find and read it. Shared but deactivated
|
|
|
61
61
|
that member's agent. This is the token economy in action — activation is how a member curates
|
|
62
62
|
which rulebooks their agent carries.
|
|
63
63
|
|
|
64
|
-
## How an agent uses a doc —
|
|
64
|
+
## How an agent uses a doc — ls, grep, cat
|
|
65
|
+
|
|
66
|
+
Docs are not injected wholesale. The agent works the corpus like a filesystem, and all three
|
|
67
|
+
verbs respect access + activation, so only docs the caller may use ever surface.
|
|
68
|
+
|
|
69
|
+
1. **`list_knowledge`** — `ls`. The corpus as a folder tree: name, id, folder, and a truncated
|
|
70
|
+
description. No bodies.
|
|
71
|
+
2. **`grep_knowledge`** — `grep -rn`. Substring match across every readable doc (or one doc, or
|
|
72
|
+
one folder), returning **doc, line number, and the matching line**. It runs inside Postgres,
|
|
73
|
+
so only matching lines cross the wire.
|
|
74
|
+
3. **`read_knowledge`** — `cat` / `sed -n 'X,Yp'`. Read a doc whole or by line range, to see
|
|
75
|
+
the context around a hit.
|
|
76
|
+
|
|
77
|
+
Matching is **substring, not ranked** — there is no index and no tokenizer, which is why a
|
|
78
|
+
corpus in any script works and why nothing goes stale. The consequence is that the *agent*
|
|
79
|
+
does the narrowing: a distinctive phrase is sharply selective, a whole question matches
|
|
80
|
+
everything. `grep_knowledge` always reports `total_matches`, so "too broad" is visible and
|
|
81
|
+
cheap to fix.
|
|
82
|
+
|
|
83
|
+
### Matching options
|
|
84
|
+
|
|
85
|
+
- **Diacritics fold by default** — `ca phe` matches `cà phê`. `diacritic_insensitive: false` matches
|
|
86
|
+
tone marks exactly.
|
|
87
|
+
- **Case folds by default**, independently of diacritics. `case_sensitive: true` matches case
|
|
88
|
+
exactly — useful for an acronym (`NK` vs `nk`) that a folded search would blur.
|
|
89
|
+
- **Whitespace is normalized on both sides.** A body converted from PDF, Word or Excel carries
|
|
90
|
+
non-breaking spaces, soft hyphens, zero-width marks and padded runs that nobody types into a
|
|
91
|
+
query; those fold to ordinary single spaces before matching, so a correct search does not return
|
|
92
|
+
a silent zero on text that is present. Your pattern is normalized the same way.
|
|
93
|
+
- **`regex: true`** treats the pattern as a POSIX regular expression: quantifiers, character
|
|
94
|
+
classes, alternation, anchors, and **`\y` for a word boundary**. Note `\y`, not `\b` — Postgres
|
|
95
|
+
spells it differently, and `\b` is rewritten for you rather than silently matching nothing.
|
|
96
|
+
Literal is the default on purpose: `0901.11.20` as a regex would also match `0901X11Y20`.
|
|
65
97
|
|
|
66
|
-
|
|
67
|
-
|
|
68
|
-
|
|
69
|
-
|
|
70
|
-
1. **Catalog** — `list_knowledge` returns every usable doc as `{ id, name, description }` — no
|
|
71
|
-
bodies. This is why the *description* carries the weight: it is what the agent reads before
|
|
72
|
-
deciding to open a doc. Write it to sell the doc — its vocabulary, the colloquial synonyms an
|
|
73
|
-
ambiguous query would use, and what it covers.
|
|
74
|
-
2. **Read** — the chat agent stages the chosen doc's content file into a code run and greps it
|
|
75
|
-
there (the body arrives as a file to `cat`/`grep`). Over this CLI you read a body directly
|
|
76
|
-
with `lotics knowledge get <id>`.
|
|
77
|
-
|
|
78
|
-
Structure your content so a reader lands on the answer without loading the rest:
|
|
98
|
+
```
|
|
99
|
+
grep_knowledge({ pattern: "\\yNK\\y", regex: true, case_sensitive: true })
|
|
100
|
+
```
|
|
79
101
|
|
|
80
|
-
|
|
81
|
-
|
|
82
|
-
|
|
83
|
-
|
|
84
|
-
|
|
85
|
-
|
|
102
|
+
### What a result is bounded by
|
|
103
|
+
|
|
104
|
+
A result carries at most 40.000 characters. Matched lines are admitted first and context fills
|
|
105
|
+
what remains, so an answer is never dropped to make room for a neighbouring line. Anything the
|
|
106
|
+
budget cut is REPORTED — `chars_elided_matches` (an answer was dropped) and
|
|
107
|
+
`context_omitted_matches` (an answer was kept without the context you asked for) mean different
|
|
108
|
+
things and call for different fixes: narrow the pattern, or ask for fewer `context_lines`.
|
|
109
|
+
Individual lines longer than 600 characters are clipped and marked, with `read_knowledge` giving
|
|
110
|
+
the rest.
|
|
111
|
+
|
|
112
|
+
`total_matches` is always the true total, even when fewer are shown.
|
|
113
|
+
|
|
114
|
+
## Write for grep — the rules that decide whether an answer is findable
|
|
115
|
+
|
|
116
|
+
Retrieval addresses **lines**. Every rule below follows from that one fact, and ignoring them
|
|
117
|
+
is the difference between a doc that answers and a doc that merely exists.
|
|
118
|
+
|
|
119
|
+
- **One self-contained fact per line.** A grep hit returns *that line*. For dense or tabular
|
|
120
|
+
data — a price row, a tariff code, a charge entry — put the whole record on one line.
|
|
121
|
+
- **Repeat the searchable terms on every line; headings do not carry down.** A line reading
|
|
122
|
+
`Rate: 15%` under a heading `Roasted coffee` will never match a search for `coffee`. Restate
|
|
123
|
+
the identifying terms inline, even when it reads redundantly to a human. This is the single
|
|
124
|
+
rule most often got wrong.
|
|
125
|
+
- **Keep a line under ~600 characters.** Past that the line is truncated in the agent's view
|
|
126
|
+
and marked as cut; the agent can still `read_knowledge` for the rest, but it costs a round
|
|
127
|
+
trip. Put the identifying terms early and the long tail late.
|
|
128
|
+
- **Put synonyms and translations on the line itself.** A bilingual row (`Cà phê, đã rang /
|
|
129
|
+
Coffee, roasted`) matches queries in either language for free. The same trick carries
|
|
130
|
+
colloquial terms next to official ones.
|
|
131
|
+
- **For prose, keep a rule and its exception close.** A hit returns one line, so a rule on line
|
|
132
|
+
40 and its exception on line 90 can be retrieved apart. Use `context_lines` when reading, and
|
|
133
|
+
keep related clauses adjacent when writing.
|
|
134
|
+
|
|
135
|
+
Markdown headings are ordinary text — useful for a human and for orienting a `read_knowledge`,
|
|
136
|
+
but they carry no retrieval weight of their own.
|
|
137
|
+
|
|
138
|
+
**The description is the discovery hint.** It is what the agent sees in the tree before it has
|
|
139
|
+
read a byte of the body, so it is the only clue for *what term to grep for*. Write it to name
|
|
140
|
+
the doc's vocabulary — the words that actually appear inside it. Keep it to a sentence: it is
|
|
141
|
+
truncated at 200 characters, and a keyword-stuffed description is cut, not rewarded.
|
|
86
142
|
|
|
87
143
|
## Updating a doc — `lotics knowledge update`
|
|
88
144
|
|
|
@@ -95,8 +151,8 @@ Send only the fields you're changing. `--from` / `--content` replaces the body;
|
|
|
95
151
|
re-reads the current content pointer and version-chains the new body — so there is no version
|
|
96
152
|
token to pass from the CLI. (The chat agent may instead send an `edits` array — anchored
|
|
97
153
|
replace / insert / append — for a surgical change; see `lotics tools update_knowledge`.) Refine
|
|
98
|
-
structure as you learn what users actually ask: add the synonym that failed to match, split
|
|
99
|
-
|
|
154
|
+
structure as you learn what users actually ask: add the synonym that failed to match, split a
|
|
155
|
+
line that was too coarse, restate a term the heading was carrying.
|
|
100
156
|
|
|
101
157
|
## Package-managed knowledge
|
|
102
158
|
|