okf 1.8.0 → 1.10.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (42) hide show
  1. checksums.yaml +4 -4
  2. data/CHANGELOG.md +615 -38
  3. data/README.md +109 -15
  4. data/lib/okf/bundle/folder.rb +20 -0
  5. data/lib/okf/bundle/search/index.rb +65 -0
  6. data/lib/okf/bundle/search/scan.rb +89 -0
  7. data/lib/okf/bundle/search.rb +262 -66
  8. data/lib/okf/bundle.rb +27 -3
  9. data/lib/okf/cli/catalog.rb +66 -0
  10. data/lib/okf/cli/command.rb +495 -0
  11. data/lib/okf/cli/files.rb +68 -0
  12. data/lib/okf/cli/graph.rb +82 -0
  13. data/lib/okf/cli/index.rb +127 -0
  14. data/lib/okf/cli/lint.rb +139 -0
  15. data/lib/okf/cli/loose.rb +78 -0
  16. data/lib/okf/cli/registry.rb +229 -0
  17. data/lib/okf/cli/render.rb +66 -0
  18. data/lib/okf/cli/search.rb +285 -0
  19. data/lib/okf/cli/server.rb +179 -0
  20. data/lib/okf/cli/skill.rb +57 -0
  21. data/lib/okf/cli/stats.rb +88 -0
  22. data/lib/okf/cli/tags.rb +122 -0
  23. data/lib/okf/cli/types.rb +37 -0
  24. data/lib/okf/cli/validate.rb +66 -0
  25. data/lib/okf/cli.rb +418 -1633
  26. data/lib/okf/{server → render}/graph/template.html.erb +1553 -175
  27. data/lib/okf/{server → render}/graph.rb +85 -9
  28. data/lib/okf/server/app.rb +17 -48
  29. data/lib/okf/server/hub/not_found.rb +663 -0
  30. data/lib/okf/server/hub.rb +504 -38
  31. data/lib/okf/skill/SKILL.md +41 -26
  32. data/lib/okf/skill/playbooks/consume.md +5 -3
  33. data/lib/okf/skill/playbooks/curate.md +3 -1
  34. data/lib/okf/skill/playbooks/maintain.md +4 -3
  35. data/lib/okf/skill/playbooks/menu.md +5 -0
  36. data/lib/okf/skill/playbooks/refine.md +92 -0
  37. data/lib/okf/skill/playbooks/search.md +47 -7
  38. data/lib/okf/skill/reference/authoring.md +3 -2
  39. data/lib/okf/skill/reference/cli.md +98 -21
  40. data/lib/okf/version.rb +1 -1
  41. data/lib/okf.rb +8 -0
  42. metadata +37 -3
@@ -5,14 +5,15 @@ description: >-
5
5
  directory of markdown files with YAML frontmatter that humans and agents read
6
6
  from one source. Use when capturing knowledge into a bundle (a service, schema,
7
7
  metric, decision, runbook: "document this in OKF", "capture this as a concept"),
8
+ converting existing docs into one ("migrate/OKFy our docs into a bundle"),
8
9
  retrieving from one without reading it whole ("what do we know about X?", "where
9
- is X documented?", "search the bundle"), updating one after code or docs change
10
- ("update the knowledge bundle"), checking its conformance or curation quality
11
- ("validate/lint the bundle"), serving or rendering it as a graph, or working in a
12
- repo that already carries an OKF bundle — a `.okf/` directory or a root `index.md`
13
- carrying `okf_version`.
10
+ is X documented?", "search the bundle"), updating one after code or docs
11
+ change ("update the knowledge bundle"), checking its conformance or curation
12
+ quality ("validate/lint the bundle"), serving or rendering it as a graph, or
13
+ working in a repo that already carries an OKF bundle — a `.okf/` directory or a
14
+ root `index.md` carrying `okf_version`.
14
15
  user-invocable: true
15
- argument-hint: "[search|produce|maintain|consume|<okf-cli-verb>] [dir] [--flags]"
16
+ argument-hint: "[search|produce|migrate|maintain|refine|consume|curate|doctor|<okf-cli-verb>] [dir|@slug] [--flags]"
16
17
  allowed-tools: Read Write Edit Grep Glob Bash
17
18
  ---
18
19
 
@@ -66,12 +67,16 @@ earn your keep as the expert, not the executable.
66
67
 
67
68
  ## The CLI is your eyes — you are the judgment
68
69
 
69
- Guard once, then trust it — the `okf` executable answers every mechanical question
70
- deterministically, and its read views show everything the browser UI does:
71
-
72
- ```bash
73
- command -v okf >/dev/null || echo "okf CLI missing — install: gem install okf (or from a checkout: cd gem && bundle exec rake install)"
74
- ```
70
+ The `okf` executable answers every mechanical question deterministically, and its
71
+ read views show everything the browser UI does. **Don't probe for it — just run
72
+ the verb.** A proactive `command -v okf` before every task spends a whole tool
73
+ round proving what the next command reveals for free; the CLI's own failure is a
74
+ cheaper, truer signal. (The two deliberate exceptions are [menu](playbooks/menu.md)
75
+ and [doctor](playbooks/doctor.md) — both decide *whether to install*, so they check
76
+ first.) The one distinction to hold: a shell `okf: command not
77
+ found` is the *only* thing that means "install it" (→ [doctor](playbooks/doctor.md));
78
+ every line that starts `error:` is okf *answering* — a bundle or usage result to
79
+ read and act on, never a missing toolchain to send to doctor.
75
80
 
76
81
  Don't memorize the surface — `okf --help` maps every verb, `okf <verb> --help` its
77
82
  flags. The division of labour is the whole game:
@@ -98,7 +103,7 @@ shapes, the tag-curation views, the server's trust boundary.
98
103
  ## Orient before you touch anything
99
104
 
100
105
  Picking up a bundle you don't already know — to consume or maintain — run `okf
101
- index <dir>` (the §6 map: every directory's index body, rollups, and listings) and
106
+ index <dir|@slug>` (the §6 map: every directory's index body, rollups, and listings) and
102
107
  read `log.md` (the §7 baseline of what changed last) **before** greping or opening
103
108
  leaves. It is the cheapest high-signal context, and the only reliable way to catch
104
109
  enumeration drift: **grep cannot find an index entry that is missing** — you can't
@@ -118,14 +123,22 @@ when you need chapter and verse.
118
123
 
119
124
  **No subcommand?** Infer intent: "document this / capture X" → `produce`;
120
125
  "convert / migrate / OKFy these existing docs into a bundle" → `migrate`; "the
121
- code changed, update the docs" → `maintain`; "what do we know about X / where
122
- is X documented" → `search`; a repo already carrying a bundle plus a task
123
- needing its knowledge → `consume`; "check / graph / preview it" → run the
124
- matching CLI verb and interpret the result. When genuinely ambiguous, ask.
125
-
126
- **Which directory?** Use the path given. Otherwise default to `.okf/` at the repo
127
- root, but first detect whether the project already keeps its bundle elsewhere
128
- (e.g. `docs/`) and prefer that. Commit the bundle alongside the code it describes.
126
+ code changed, update the docs" → `maintain`; "restructure / rebalance the
127
+ bundle / is the structure right / get more out of it" → `refine`; "what do we
128
+ know about X / where is X documented" → `search`; a repo already carrying a
129
+ bundle plus a task needing its knowledge → `consume`; "check / graph / preview
130
+ it" → run the matching CLI verb and interpret the result. When genuinely
131
+ ambiguous, ask.
132
+
133
+ **Which target?** A leading `@` is a *registry ref*, not a path: `@slug` names a
134
+ bundle registered with `okf registry set`, bare `@` the default — route it
135
+ straight to `okf <verb> @slug` and skip the directory hunt (`okf search` spans
136
+ several: `@a @b`, or `@all`). A plain path is used as given. Given no target and a
137
+ cwd that carries no bundle, `okf registry list` is the next move, not a hunt
138
+ across sibling directories. Producing a *new* bundle with no path? Default to
139
+ `.okf/` at the repo root, but first detect whether the project already keeps its
140
+ bundle elsewhere (e.g. `docs/`) and prefer that; commit it alongside the code it
141
+ describes.
129
142
 
130
143
  **Target isn't a bundle?** When a verb points at a directory that holds markdown
131
144
  but no root `index.md` carrying `okf_version` — `validate` failing wholesale on
@@ -147,15 +160,17 @@ Read the referenced playbook before executing — it *is* the procedure.
147
160
  | `produce` | Author | create or extend a bundle | [playbooks/produce.md](playbooks/produce.md) |
148
161
  | `migrate` | Author | convert existing docs in place: frontmatter + reserved files, bodies verbatim | [playbooks/migrate.md](playbooks/migrate.md) |
149
162
  | `maintain` | Author | sync the bundle's content with reality after a change | [playbooks/maintain.md](playbooks/maintain.md) |
163
+ | `refine` | Author | optimize the bundle's structure: evidence-driven, cohesion-first; proposes, never auto-applies | [playbooks/refine.md](playbooks/refine.md) |
150
164
  | `consume` | Use | use the bundle as context for a task | [playbooks/consume.md](playbooks/consume.md) |
151
165
  | `curate` | Curate | structural upkeep as it stands: validate + lint + loose | [playbooks/curate.md](playbooks/curate.md) |
152
166
  | `doctor` | Setup | install and verify the CLI, then doctor the bundle | [playbooks/doctor.md](playbooks/doctor.md) |
153
- | `<okf-cli-verb>` | Read | validate, lint, loose, index, catalog, files, tags, types, stats, graph, server, render, registry, skill | `okf <verb> --help` + [reference/cli.md](reference/cli.md) |
167
+ | `<okf-cli-verb>` | Read | validate, lint, loose, index, catalog, files, tags, types, stats, graph, server, render, registry, skill — **plus any verb an installed extension adds** (`okf help` is authoritative, this list is not) | `okf <verb> --help` + [reference/cli.md](reference/cli.md) |
154
168
 
155
- Two boundaries worth keeping sharp: `curate` is structural upkeep only — when
156
- the *content* no longer matches reality, that is `maintain` — and `doctor` is
157
- the one playbook that does not assume the CLI is installed. In Claude Code with
158
- the okf plugin, `/okf:gem` routes these same verbs.
169
+ Three boundaries worth keeping sharp: `curate` is structural upkeep only — when
170
+ the *content* no longer matches reality, that is `maintain`, and when the
171
+ content is right but the *shape* underserves retrieval, that is `refine` — and
172
+ `doctor` is the one playbook that does not assume the CLI is installed. In
173
+ Claude Code with the okf plugin, `/okf:gem` routes these same verbs.
159
174
 
160
175
  ## The lifecycle is a flywheel, not phases
161
176
 
@@ -1,8 +1,10 @@
1
1
  # Playbook: consume — use a bundle as context
2
2
 
3
- 1. **Orient first** (the [SKILL.md](../SKILL.md) reflex): `okf index <dir>` maps the
4
- whole bundle in one pass — every directory's index body, rollups, and listings —
5
- and `log.md` gives recent history. Then follow links only into the concepts the
3
+ 1. **Orient first** (the [SKILL.md](../SKILL.md) reflex): `okf index <dir|@slug>` maps
4
+ the whole bundle in one pass — every directory's index body, rollups, and listings —
5
+ and `log.md` gives recent history. Address a registered bundle by `@slug` (bare
6
+ `@` = the default); if the cwd carries no bundle, `okf registry list` finds one
7
+ instead of a directory hunt. Then follow links only into the concepts the
6
8
  task needs. For a *pointed question* rather than broad context, switch to the
7
9
  [search playbook](search.md): map → finder (`okf search`) → only the winning
8
10
  bodies. For a large bundle, `okf graph --json --minimal` gives the whole link
@@ -9,7 +9,9 @@ reachability, backlog, completeness, hygiene. It is not `maintain`, the
9
9
  skill's workflow for when the project changed and the bundle's *content*
10
10
  must catch up with reality; reach for that one when what is written stopped
11
11
  being true. Curating can surface semantic staleness, and when it does,
12
- switch to `maintain` for those concepts.
12
+ switch to `maintain` for those concepts. And it never moves knowledge: when
13
+ the content is right but the shape underserves retrieval — fat areas,
14
+ mis-homed hubs, an uncurated tag layer — that is [refine](refine.md).
13
15
 
14
16
  1. Locate the bundle: the directory you were given, if any; otherwise a
15
17
  `.okf/` directory or a root `index.md` whose frontmatter carries
@@ -1,8 +1,9 @@
1
1
  # Playbook: maintain — keep a bundle in sync with reality
2
2
 
3
3
  Reach for this when the project changed and the bundle's *content* must catch
4
- up. The modelling craft behind steps 3 and 7 lives in
5
- [authoring.md](../reference/authoring.md).
4
+ up. Restructuring the bundle itself — moving concepts, adding areas — is
5
+ [refine](refine.md), not maintain. The modelling craft behind steps 3 and 7
6
+ lives in [authoring.md](../reference/authoring.md).
6
7
 
7
8
  1. **Orient before hunting.** Run `okf index <dir>` (the §6 map — every directory's
8
9
  index body, rollups, and listings), read `log.md` (the §7 baseline: what changed
@@ -43,7 +44,7 @@ up. The modelling craft behind steps 3 and 7 lives in
43
44
  by design only through its index — leave it. **Terminal-by-design is not a
44
45
  defect.** Loose ≠ orphan: an index listing makes a file *reachable* (not an
45
46
  orphan) but is not a graph edge, so an indexed file can still float here.
46
- 7. **Curate the tag vocabulary** <!-- rule:okf-tag-vocabulary --> when the pass
47
+ 7. **Curate the tag vocabulary** when the pass
47
48
  touched tags, or when `okf tags <dir>` shows a long tail of singletons. Run `okf tags <dir> --by area` and
48
49
  `--by type` — the grouped view is the analysis; read each group top-down:
49
50
  - **twins** — two tags riding the exact same concepts (equal counts sort them
@@ -33,6 +33,11 @@ is the lede.
33
33
  the working tree has uncommitted changes to the code the bundle describes
34
34
  (`git status`), prefer **`maintain`**: that is exactly the drift it exists
35
35
  to close.
36
+ - **clean, but the shape strains** — one area dwarfing the rest in
37
+ `okf stats`, tags spread thin across areas in `okf tags --by area`, hubs
38
+ whose inbound links are mostly foreign in `okf graph --hubs` → offer
39
+ **`refine`** (evidence-driven restructuring; it proposes before it
40
+ touches anything).
36
41
  4. **Freshness is off by default.** If the bundle carries timestamps, note that a
37
42
  plain `lint` said nothing about staleness and `okf lint <root> --stale-after
38
43
  90d` is the check that would.
@@ -0,0 +1,92 @@
1
+ # Playbook: refine — restructure a bundle to get the most from OKF
2
+
3
+ Reach for this when the bundle's *content* is right but its *shape* may not be:
4
+ areas grown fat by additive passes, hubs homed by history, a tag layer that
5
+ never became the second index. Refine optimizes the projection — the same
6
+ knowledge, arranged to serve progressive disclosure, the emergent graph,
7
+ cross-cutting tags, and capture-once-link-many. It is not [curate](curate.md)
8
+ (upkeep of the structure as it stands) and not [maintain](maintain.md) (content
9
+ catching up with reality): refine changes where knowledge lives, never what it
10
+ says. Its permitted edits are structural — move a concept, extract a duplicated
11
+ fact to one home, section an index, retag, relink, and write the connective
12
+ sentence a link lives in; summarizing, updating, or correcting a body is
13
+ maintain's job, reached by switching verbs, not by stretching this one.
14
+
15
+ The frame that governs every move: the directory tree is a **lossy projection
16
+ of the link graph**. A tree gives each concept one parent, so the tree encodes
17
+ only the single dominant decomposition; every genuinely many-to-many
18
+ relationship rides links and tags, never new directories. And cohesion outranks
19
+ balance — a move has semantic cost, so balance is a tiebreaker and a fatness
20
+ alarm, never the objective. <!-- rule:okf-cohesion-over-balance -->
21
+
22
+ 1. **Orient.** `okf index <dir|@slug> --no-body` (areas, fan-out, depth),
23
+ `log.md` (how the bundle grew), `okf stats` (totals). Additive growth
24
+ optimizes each pass locally, never the whole — that is the drift this
25
+ playbook corrects.
26
+ 2. **Measure — the CLI is the evidence.** Baseline `validate` / `lint
27
+ --stale-after` / `loose` first: refine assumes a sound bundle, and hard
28
+ errors are [curate](curate.md)'s job. Then the two structural reads:
29
+ - `okf tags <dir> --by area` — each row carries `count/total`, so a tag's
30
+ **locality** reads directly: a tag wholly inside one area names a *domain*
31
+ (the directories are right); one spread across areas names a *concern*.
32
+ - `okf graph <dir> --hubs` — concepts ranked by inbound links, each with
33
+ the areas those links come from: the **origin test** for every hub.
34
+ 3. **Diagnose — you are the judgment.** The measurements are evidence, never
35
+ verdicts:
36
+ - **Concerns never become containers.** A directory built around a spread
37
+ tag ("everything async") prunes nothing — most needs would enter it.
38
+ The cross-cut stays a tag. <!-- rule:okf-concern-not-container -->
39
+ - **A directory must prune.** The positive test for any area, existing or
40
+ proposed: does knowing "it's in there" eliminate a large, even slice? A
41
+ good node splits its parent into chunks that are nameable, mutually
42
+ exclusive, and roughly comparable in size. And small is not merge-worthy
43
+ on its own — a two-concept area that is a genuinely distinct domain
44
+ stays. <!-- rule:okf-directory-prunes -->
45
+ - **The hub origin test.** Inbound majority from the hub's own area:
46
+ well-homed, leave it. A dominant *foreign* area: that area is the better
47
+ home. Foreign majority with *no* dominant area: a shared primitive — the
48
+ only admission ticket into a shared-core area (without that test, a
49
+ `foundation/` rots into a `misc/`). Two comparable strong ties, one of
50
+ them home: stay and carry the other as a tag — moving trades one
51
+ imbalance for another. And in a design bundle expect the central
52
+ decisions to fail this test wholesale: that is centrality, not
53
+ mis-homing.
54
+ - **Fatness alarm, not fatness rule.** A fat area (≳20–25 concepts) wants
55
+ **heading sections inside its `index.md`** first — the same prune as
56
+ sub-directories, for zero extra hops and no new enumeration to keep
57
+ sound. Directory nesting pays only at hundreds of concepts, and only
58
+ where the index's own headings already form separable, nameable
59
+ sub-groups — fatness alone never justifies depth; the sections that
60
+ formed are the evidence the split exists.
61
+ - **Duplication.** Read the area overviews for a fact re-explained in
62
+ several (drifting tables are the tell); capture-once-link-many says
63
+ extract it into one concept and link from the rest. Extraction is the one
64
+ refine move that touches bodies, and it redistributes rather than
65
+ rewrites: assemble the canonical concept from the copies, keep every
66
+ copy's unique domain-specific detail (in the extract, or in the one-line
67
+ note left beside each link), and where the copies *disagree*, which is
68
+ true is a [maintain](maintain.md) question — verify against reality or
69
+ flag the conflict in the proposal, never silently pick a winner while
70
+ merging. <!-- rule:okf-extract-not-rewrite -->
71
+ - **Vocabulary twins.** The [maintain](maintain.md) tag-curation recipe
72
+ (twins, echoes, singletons) applies to `type` too — `okf types <dir>`.
73
+ 4. **Plan — tier by leverage ÷ churn, free levers first.** Tag curation,
74
+ index heading-sectioning, and extraction before any file move; a move only
75
+ when the origin test demands one, and each gated by **do-nothing**: skip it
76
+ unless its value beats its churn. Record what you *declined* and why — the
77
+ decline list is what stops the next pass from re-proposing it.
78
+ 5. **Propose — never auto-apply.** Refine's output is a short report (the
79
+ evidence, the tiers, the declines) plus a ready-to-run execution prompt the
80
+ user can hand back later: a scope line, the governing principles above, the
81
+ explicit prohibitions, the closeout gate as acceptance. Analysis and
82
+ execution are separate on purpose — the judgment is spent once, frozen, and
83
+ then executed without re-derivation. <!-- rule:okf-refine-proposes -->
84
+ 6. **Execute only on approval**, then walk the
85
+ [Closeout gate](../reference/authoring.md#closeout--the-finishing-gate):
86
+ every touched `index.md` re-enumerated, links absolute bundle-relative so
87
+ they survived the moves, a dated `log.md` entry carrying the *why*,
88
+ validate/lint/loose clean, and a before/after of step 2's evidence.
89
+
90
+ Two traps: never split a cohesive cluster (a deliberately paired mirror flow)
91
+ to hit a size band, and never tag a concept with its own directory's name — a
92
+ group-name echo adds no edge.
@@ -6,9 +6,13 @@ can query cheaply is dead weight. The discipline is progressive disclosure
6
6
  (spec §6): every step pays a few hundred bytes to decide what the next step
7
7
  reads, and full bodies are read last, and only the winners.
8
8
 
9
- 1. **Guard once**: `command -v okf`. Missing → [doctor](doctor.md). No CLI at
10
- all → read the root `index.md`, then each relevant area's `index.md`, by hand.
11
- 2. **Ingest the map and decide where to look.** `okf index <dir> --no-body` is
9
+ 1. **Just run it — no presence probe.** Point the finder at a path or an `@slug`
10
+ (a registered bundle; bare `@` = the default). Only a shell `okf: command not
11
+ found` means install (→ [doctor](doctor.md)); with no CLI possible at all, read
12
+ the root `index.md` then each relevant area's `index.md` by hand. No bundle in
13
+ the cwd? `okf registry list` names the registered ones — address them by
14
+ `@slug`, don't hunt sibling directories.
15
+ 2. **Ingest the map and decide where to look.** `okf index <dir|@slug> --no-body` is
12
16
  the skeleton: every directory with its concept count, types, tags, children.
13
17
  *You* do the semantic matching here — the question names a meaning, the map
14
18
  names areas; connect them by judgment, not string equality. When an area
@@ -17,8 +21,37 @@ reads, and full bodies are read last, and only the winners.
17
21
  <!-- rule:okf-search-map-first -->
18
22
  3. **Cut across with the finder when the question is lexical.** An exact
19
23
  symbol, an error code, a column name, a phrase — things structure won't
20
- surface — go to `okf search <dir> <terms>` (terms AND together; `--regexp`
21
- for patterns like `err_[a-z]+_409`). Scope it with what the map taught you:
24
+ surface — go to `okf search <dir> <terms>`. Terms AND together and are matched
25
+ **literally against raw text**, so an exact query means what it looks like: a
26
+ phrase, a dotted version (`7.2.0`), an underscored identifier (`customer_id`),
27
+ a mid-word fragment (`ustomer`) and a word written in `backticks` all match
28
+ the way you typed them. <!-- rule:okf-search-exact-identifiers -->
29
+
30
+ **Match the engine to the shape of the query, not to habit** — the default
31
+ answers most of them, and the two engines fail in opposite directions:
32
+ <!-- rule:okf-search-engine-choice -->
33
+
34
+ | Your query is | Reach for | Because |
35
+ |---|---|---|
36
+ | an identifier, version, path, phrase, or anything in `` `backticks` `` | *nothing — the default* | matched literally; the index shatters all of these |
37
+ | a mid-word fragment (`ustomer`) | *nothing — the default* | an infix is not a token, so the index cannot reach it |
38
+ | a pattern (`err_[a-z]+_409`) | `-e` | Ruby regexp over raw text; still the scan |
39
+ | a partial word (`dedup` → `deduplication`) | *nothing — the default* | substring covers prefixes, and suffixes and infixes too |
40
+ | a theme, where you want the best match to lead | `--engine index` | BM25+ ranks by relevance, not by summed field weight |
41
+ | possibly mistyped | `--fuzzy` | edit distance 0.2 × term length — the index's alone |
42
+ | being reconciled with the browser page | `--engine index` | same MiniSearch build, so the two rank alike |
43
+
44
+ The index has exactly **three** things the default lacks: relevance ranking,
45
+ typo tolerance, and page parity. Its `prefix` capability is not a fourth — a
46
+ substring match already reaches every prefix, so `prefix` is what the index
47
+ needs to *catch up*, not a reason to choose it.
48
+
49
+ **`--fuzzy` is an engine switch, not a mode.** It routes to the index, so a
50
+ run that only wanted a typo forgiven also gets token matching, shattered
51
+ identifiers and unfindable code spans. Fix the spelling and stay on the
52
+ default when you can. <!-- rule:okf-search-fuzzy-is-a-switch -->
53
+
54
+ Scope any of them with what the map taught you:
22
55
  `--area billing`, `--type Decision`, `--tag idempotency`, `--in body`.
23
56
  Matches rank by where they hit, and the snippet often *is* the answer.
24
57
  When the answer may live in another registered bundle, span them — leading
@@ -40,6 +73,13 @@ Anti-patterns, each a real token bill:
40
73
  - **Grep before map.** Grep cannot find the entry that is *missing*, and it
41
74
  returns line noise where `search` returns ranked concepts. Grep is the
42
75
  fallback when the CLI is absent, not the first move.
43
- - **Mechanical synonym retries.** The finder is exact by design; *you* are the
76
+ - **Mechanical synonym retries.** The finder is exact by default; *you* are the
44
77
  fuzzy layer. When terms miss, learn the bundle's vocabulary — `okf tags
45
- <dir>`, `okf types <dir>` — and re-ask in its own words.
78
+ <dir>`, `okf types <dir>` — and re-ask in its own words. `--fuzzy` forgives a
79
+ *typo*, not a wrong vocabulary, so it is the wrong reach for this.
80
+ - **Flag-shopping a query that found nothing.** Cycling `--fuzzy`, then
81
+ `--engine index`, then `-e` over the same terms is guessing, and each engine
82
+ fails differently enough that one of them eventually returns *something* —
83
+ which is how a wrong answer gets found. Zero matches is usually a vocabulary
84
+ result, not an engine result: go back to the map and the tag list. Reach for a
85
+ different engine when you can say which property of the query needs it.
@@ -115,8 +115,9 @@ bundle-root [root-index](../templates/root-index.md), [log](../templates/log.md)
115
115
  ## Playbooks
116
116
 
117
117
  The step-by-step playbooks live in [../playbooks/](../playbooks/), one file per
118
- verb (produce, maintain, consume, curate, doctor), routed by the Commands table
119
- in [SKILL.md](../SKILL.md). The Closeout below is their shared finishing gate.
118
+ verb (search, produce, migrate, maintain, consume, curate, doctor), routed by the
119
+ Commands table in [SKILL.md](../SKILL.md). The Closeout below is their shared
120
+ finishing gate.
120
121
 
121
122
  ## Closeout — the finishing gate
122
123
 
@@ -6,14 +6,13 @@ reimplemented in this skill. They run the deterministic `okf` executable shipped
6
6
  the companion gem — the single source of truth for OKF mechanics. Your job is to
7
7
  invoke it correctly and interpret the result, not to reason out conformance by hand.
8
8
 
9
- ## Presence guard
9
+ ## When it isn't installed
10
10
 
11
- Check the tool exists before relying on it. If it is missing, the gem is not
12
- installed — say so and stop; never fabricate a result:
13
-
14
- ```bash
15
- command -v okf >/dev/null || echo "okf CLI not found — install it: 'gem install okf' (or from a checkout: 'cd gem && bundle exec rake install')"
16
- ```
11
+ Don't probe for the tool before using it — just run the verb. A shell `okf:
12
+ command not found` is the only thing that means the gem isn't installed: say so
13
+ and stop (`gem install okf`, or from a checkout `cd gem && bundle exec rake
14
+ install`); never fabricate a result. Any line that starts `error:` is the CLI
15
+ *answering* — a bundle or usage result to read, not a missing toolchain.
17
16
 
18
17
  ## Invocation
19
18
 
@@ -21,6 +20,12 @@ The surface is self-describing — `okf --help` maps every verb, `okf <verb> --h
21
20
  its flags. Ask the tool for what exists; this file carries only what `--help`
22
21
  cannot: each verb's semantics, its traps, and its JSON shape.
23
22
 
23
+ **The verb list is open, not closed.** An installed extension gem adds verbs of
24
+ its own, listed under `installed extensions:` in `okf help`. So a verb that
25
+ `--help` shows and this file does not document is **normal, not a documentation
26
+ error** — ask `okf <verb> --help` for it, and expect nothing here about its
27
+ semantics or JSON. Everything below documents the built-ins only.
28
+
24
29
  **`--json` is compact by design.** Every emitting verb prints single-line JSON —
25
30
  the token-efficient substrate you consume; `--pretty` (which implies `--json`)
26
31
  indents it for a human. The bytes differ, the JSON is identical, so parse either.
@@ -140,16 +145,74 @@ defect — a terminal leaf (a backlog item, a spec reference) can be loose by de
140
145
  The browser page's search brought to the CLI and extended to bodies, so "which
141
146
  concept covers X?" costs rows, not body reads. `okf search <dir> <term…>`:
142
147
  terms AND together — every term must hit at least one searched field, not
143
- necessarily the same one — as case-insensitive substrings, or as Ruby regular
144
- expressions with `--regexp`/`-e` (an invalid pattern is a usage error, exit 2).
148
+ necessarily the same one — matched **literally against raw text** by default, or
149
+ as Ruby regular expressions with `--regexp`/`-e` (an invalid pattern is a usage
150
+ error, exit 2). `--fuzzy` forgives typos; pairing it with `-e` is a usage error,
151
+ since a pattern is matched literally rather than by edit distance.
145
152
  `--in a,b` restricts the searched fields (title, id, tags, type, description,
146
153
  body); the shared `--type/--area/--tag` filters narrow the candidates *first*,
147
154
  so a search scoped by what `index` taught you stays surgical.
148
155
 
156
+ **The default is exact, so an exact query means what it looks like.** A phrase in
157
+ one argument (`"dedup key"`), a dotted version (`7.2.0`), an underscored
158
+ identifier (`customer_id`), a mid-word fragment (`ustomer`) and a word written in
159
+ `backticks` all match literally. This is what the scan engine buys, and it is the
160
+ default precisely because those queries are the common ones and the alternative
161
+ loses them silently. <!-- rule:okf-search-exact-identifiers -->
162
+
163
+ **`--engine index` is the other engine, and the one to reach for when ranking
164
+ matters more than exactness.** The engine is normally chosen by what the query
165
+ needs — `--fuzzy` routes to the index, anything else stays on the default scan —
166
+ and nothing is printed about the choice. `--engine NAME` overrides that for the
167
+ case the flags cannot express: a matching *model* requires no capability, so no
168
+ flag selects one. Under the index, terms match whole tokens and their prefixes
169
+ (`dedup` finds `deduplication`), rows rank by BM25+, and it is the engine the
170
+ browser page runs — so name it when reconciling a CLI answer with the page. The
171
+ cost is real: its tokenizer splits on punctuation, so identifiers shatter
172
+ (`customer_id` → `customer` + `id`), an infix finds nothing, and a backtick is
173
+ never split off at all, so a word inside a code span is unfindable — a large
174
+ silent loss, since technical prose is full of them. **Do not count on ranking to
175
+ rescue it** — BM25 normalizes by field length, so a short concept dense in `7`,
176
+ `2` and `0` can outrank the one that actually says `7.2.0`. Naming an engine that
177
+ cannot do what you also asked (`--engine index -e`) is a usage error naming one
178
+ that can. <!-- rule:okf-search-engine-choice -->
179
+
180
+ **The capabilities, and which engine has them.** An engine is selected by what
181
+ the query *requires*; only a matching model has to be named, because requiring
182
+ nothing is not something a flag can express:
183
+
184
+ | Flag | Capability | Engine | What it does |
185
+ |---|---|---|---|
186
+ | *(none)* | — | scan | literal substring over raw text; scores by summed field weight |
187
+ | `-e` / `--regexp` | `regexp` | scan | each term is a Ruby regexp, case-insensitive; invalid → exit 2 |
188
+ | `--fuzzy` | `fuzzy` | **index** | edit distance 0.2 × term length — and switches engine |
189
+ | `--engine index` | — | index | whole-token + prefix matching, BM25+ ranking, browser parity |
190
+ | `--engine scan` | — | scan | the default, spelled out |
191
+
192
+ Two consequences worth holding. **`--fuzzy` is an engine switch, not a mode**: it
193
+ carries the whole index with it, so a run that wanted one typo forgiven also gets
194
+ shattered identifiers and unfindable code spans — fix the spelling and stay on
195
+ the default when you can. And **`-e` moves nothing** now, because the default
196
+ engine already offers `regexp`; it changes how a term is *read* (pattern rather
197
+ than literal), not where it is matched. <!-- rule:okf-search-fuzzy-is-a-switch -->
198
+
199
+ `prefix` is a capability the index declares but no flag selects — it is always on
200
+ there. **It is not a reason to reach for the index**: a substring match already
201
+ covers every prefix and then some, so `dedup` finds `deduplication` under both
202
+ engines, while `duplication` and `uplicat` find it under the default only. Prefix
203
+ is what the index needs to catch up to raw text, not a capability it adds on top.
204
+ The index's real advantages over the default are exactly three — relevance
205
+ ranking, typo tolerance, and page parity.
206
+
149
207
  **Search spans bundles.** Leading @refs pick several registered bundles
150
208
  (`okf search @handbook @notes auth`); **`@all`** is the ref that means every one.
151
- The per-bundle rankings merge — scores are absolute term weights, so they
152
- compare across bundles — and each row carries its bundle's slug. This is the
209
+ Rows from different bundles are ranked together and comparable, and each row
210
+ carries its bundle's slug. Under `--engine index` the bundles go into **one
211
+ corpus** — BM25 prices a term by how rare it is, so separately-ranked lists would
212
+ not compare — which makes a score relative to the whole answer: the same concept
213
+ scores lower searched beside others than searched alone. The default scan needs
214
+ no such trick — its score is absolute, so a row is worth the same either way.
215
+ This is the
153
216
  cross-bundle retrieval the in-page search does not have: one question, every
154
217
  bundle you keep. <!-- rule:okf-search-all -->
155
218
 
@@ -178,12 +241,14 @@ which has no slug to give. Two sharp edges: every *leading* @-arg is taken as a
178
241
  the CLI notes both traps on stderr — and any ref, even one, switches the JSON
179
242
  envelope (next paragraph).
180
243
 
181
- Rows rank by **where** they hit — title 5, id 4, tags 3, type/description 2,
182
- body 1, summed over matched fields — and carry one bounded context snippet from
183
- the strongest match that needs context (description or body). Deliberately not
184
- fuzzy: the consuming agent is the fuzzy layer — when terms miss, learn the
185
- bundle's vocabulary from `tags`/`types` and re-ask in its own words, rather
186
- than hammering synonyms. Advisory read: **exit 0 even with zero matches**.
244
+ Rows rank by where they hit — title 5, id 4, tags 3, type/description 2, body 1 —
245
+ summed as an absolute score by the default scan, and carried as per-field boost
246
+ into **BM25+** under `--engine index`. Each row carries one bounded context
247
+ snippet from the strongest match that needs context (description or body). Every row still names the fields that hit (`matched`), so a result stays
248
+ citable rather than being a bare relevance number. Exact by default: the
249
+ consuming agent is the fuzzy layer — when terms miss, learn the bundle's
250
+ vocabulary from `tags`/`types` and re-ask in its own words, rather than
251
+ hammering synonyms or reaching for `--fuzzy` before you have looked. Advisory read: **exit 0 even with zero matches**.
187
252
  JSON, plain-dir mode: `{ bundle, query, count, matches: [{ id, title, type,
188
253
  area, tags, matched, score, snippet }] }`. Registry mode — any leading @ref,
189
254
  `@all` among them — swaps the envelope: `{ bundles: [{ slug, dir }, …],
@@ -237,10 +302,14 @@ in/out link degree). Add `--json` to any for a machine substrate.
237
302
  - **`tags`** — every tag with the concepts that carry it, ordered by count
238
303
  descending. The "what themes dominate" view. JSON: `{ bundle, count, tags: [{ tag,
239
304
  count, concepts: [id, …] }] }`. `--by type|area` regroups the list per concept
240
- dimension with **within-group** counts (a tag spanning groups appears in each) —
241
- the substrate for tag curation; the judgment recipe lives in the
242
- [maintain playbook](../playbooks/maintain.md). JSON: `{ bundle, count, by,
243
- groups: [{ <dim>, count, tags: […] }] }`.
305
+ dimension with **within-group** counts (a tag spanning groups appears in each);
306
+ each row also carries the tag's **total** across the narrowed set, printed
307
+ `count/total` when they differ — so a tag's locality reads per row (a plain
308
+ count = wholly local; `2/7` = a cross-cutting spread). The substrate for tag
309
+ curation and for [refine](../playbooks/refine.md)'s domain-vs-concern read;
310
+ the judgment recipes live in the [maintain playbook](../playbooks/maintain.md)
311
+ and the [refine playbook](../playbooks/refine.md). JSON: `{ bundle, count, by,
312
+ groups: [{ <dim>, count, tags: [{ tag, count, total, concepts }] }] }`.
244
313
  - **`types`** — every type with the concepts that carry it, ordered by count
245
314
  descending. The "what kinds of knowledge" view. JSON: `{ bundle, count, types:
246
315
  [{ type, count, concepts: [id, …] }] }`.
@@ -368,3 +437,11 @@ drops each node's body, and `--minimal` ships only `id`/`title` plus the type/ta
368
437
  indexes — the lean shape the `server` page boots from. Reach for the full dump
369
438
  only when the task truly consumes every body; for one question, the
370
439
  [search verb](#search--ranked-text-retrieval-metadata--body) is orders cheaper.
440
+
441
+ `--hubs` swaps the dump for the **inbound ranking**: every concept with at
442
+ least one inbound link, ranked by inbound degree, each with its links grouped
443
+ by *source area* (`core/status ×3 flows 2, billing 1`) — the evidence for
444
+ [refine](../playbooks/refine.md)'s hub origin test ("is this hub well-homed?").
445
+ A source at the bundle root counts under `(root)`. JSON: `{ bundle, count,
446
+ hubs: [{ id, area, inbound, by_area: { <area>: n } }] }`. Advisory read, exit 0;
447
+ `--minimal`/`--no-body` shape node payloads and change nothing here.
data/lib/okf/version.rb CHANGED
@@ -1,5 +1,5 @@
1
1
  # frozen_string_literal: true
2
2
 
3
3
  module OKF
4
- VERSION = "1.8.0"
4
+ VERSION = "1.10.0"
5
5
  end
data/lib/okf.rb CHANGED
@@ -40,6 +40,14 @@ module OKF
40
40
  require "okf/bundle"
41
41
  require "okf/bundle/graph"
42
42
  require "okf/bundle/search"
43
+ # These two lines ARE the engine preference order. Each engine registers itself
44
+ # at load, `Search.engines` is registration order, and the router walks it after
45
+ # putting DEFAULT_ENGINE first — so reordering these requires reorders which
46
+ # engine answers a query two engines could both answer. `loading_test.rb` pins
47
+ # the result (`[:index, :scan]`) so the coupling cannot drift unnoticed, but the
48
+ # coupling is here, not there.
49
+ require "okf/bundle/search/index"
50
+ require "okf/bundle/search/scan"
43
51
  require "okf/bundle/validator"
44
52
  require "okf/bundle/validator/result"
45
53
  require "okf/bundle/linter"
metadata CHANGED
@@ -1,7 +1,7 @@
1
1
  --- !ruby/object:Gem::Specification
2
2
  name: okf
3
3
  version: !ruby/object:Gem::Version
4
- version: 1.8.0
4
+ version: 1.10.0
5
5
  platform: ruby
6
6
  authors:
7
7
  - Rodrigo Serradura
@@ -37,6 +37,20 @@ dependencies:
37
37
  - - ">="
38
38
  - !ruby/object:Gem::Version
39
39
  version: '1.4'
40
+ - !ruby/object:Gem::Dependency
41
+ name: minifts
42
+ requirement: !ruby/object:Gem::Requirement
43
+ requirements:
44
+ - - "~>"
45
+ - !ruby/object:Gem::Version
46
+ version: '1.0'
47
+ type: :runtime
48
+ prerelease: false
49
+ version_requirements: !ruby/object:Gem::Requirement
50
+ requirements:
51
+ - - "~>"
52
+ - !ruby/object:Gem::Version
53
+ version: '1.0'
40
54
  description: |
41
55
  OKF (Open Knowledge Format) is portable knowledge: Markdown files with YAML
42
56
  frontmatter that both humans and agents read from one source. This gem is the
@@ -66,10 +80,28 @@ files:
66
80
  - lib/okf/bundle/linter/report.rb
67
81
  - lib/okf/bundle/reader.rb
68
82
  - lib/okf/bundle/search.rb
83
+ - lib/okf/bundle/search/index.rb
84
+ - lib/okf/bundle/search/scan.rb
69
85
  - lib/okf/bundle/validator.rb
70
86
  - lib/okf/bundle/validator/result.rb
71
87
  - lib/okf/bundle/writer.rb
72
88
  - lib/okf/cli.rb
89
+ - lib/okf/cli/catalog.rb
90
+ - lib/okf/cli/command.rb
91
+ - lib/okf/cli/files.rb
92
+ - lib/okf/cli/graph.rb
93
+ - lib/okf/cli/index.rb
94
+ - lib/okf/cli/lint.rb
95
+ - lib/okf/cli/loose.rb
96
+ - lib/okf/cli/registry.rb
97
+ - lib/okf/cli/render.rb
98
+ - lib/okf/cli/search.rb
99
+ - lib/okf/cli/server.rb
100
+ - lib/okf/cli/skill.rb
101
+ - lib/okf/cli/stats.rb
102
+ - lib/okf/cli/tags.rb
103
+ - lib/okf/cli/types.rb
104
+ - lib/okf/cli/validate.rb
73
105
  - lib/okf/concept.rb
74
106
  - lib/okf/concept/file.rb
75
107
  - lib/okf/markdown/citations.rb
@@ -77,10 +109,11 @@ files:
77
109
  - lib/okf/markdown/links.rb
78
110
  - lib/okf/path.rb
79
111
  - lib/okf/registry.rb
112
+ - lib/okf/render/graph.rb
113
+ - lib/okf/render/graph/template.html.erb
80
114
  - lib/okf/server/app.rb
81
- - lib/okf/server/graph.rb
82
- - lib/okf/server/graph/template.html.erb
83
115
  - lib/okf/server/hub.rb
116
+ - lib/okf/server/hub/not_found.rb
84
117
  - lib/okf/server/runner.rb
85
118
  - lib/okf/skill.rb
86
119
  - lib/okf/skill/SKILL.md
@@ -91,6 +124,7 @@ files:
91
124
  - lib/okf/skill/playbooks/menu.md
92
125
  - lib/okf/skill/playbooks/migrate.md
93
126
  - lib/okf/skill/playbooks/produce.md
127
+ - lib/okf/skill/playbooks/refine.md
94
128
  - lib/okf/skill/playbooks/search.md
95
129
  - lib/okf/skill/reference/APACHE-2.0.txt
96
130
  - lib/okf/skill/reference/SPEC.md