@sayknow-cli/coding-agent 0.5.26 → 0.6.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +23 -0
- package/dist/types/config/settings-schema.d.ts +25 -5
- package/dist/types/decisions/keyword-learning.d.ts +61 -0
- package/dist/types/decisions/llm-backend.d.ts +13 -1
- package/dist/types/decisions/prompt-triage.d.ts +42 -0
- package/dist/types/decisions/skill-routing.d.ts +41 -6
- package/dist/types/hooks/native-prompt-routing.d.ts +21 -0
- package/dist/types/hooks/native-skill-hook.d.ts +3 -0
- package/dist/types/hooks/skill-keywords.d.ts +9 -0
- package/dist/types/hooks/skill-state.d.ts +20 -3
- package/dist/types/hooks/ui-skill-keywords.d.ts +15 -0
- package/dist/types/sdk/session.d.ts +3 -13
- package/dist/types/session/agent-session.d.ts +8 -0
- package/dist/types/session/auth-storage-discovery.d.ts +13 -0
- package/dist/types/tools/browser.d.ts +2 -2
- package/package.json +7 -7
- package/scripts/eval-skill-routing.ts +37 -12
- package/src/config/settings-schema.ts +27 -5
- package/src/decisions/index.ts +8 -2
- package/src/decisions/keyword-learning.ts +678 -0
- package/src/decisions/llm-backend.ts +213 -67
- package/src/decisions/prompt-triage.ts +163 -0
- package/src/decisions/skill-routing.ts +39 -56
- package/src/decisions/typesafe-backend.ts +3 -0
- package/src/hooks/native-prompt-routing.ts +190 -0
- package/src/hooks/native-skill-hook.ts +21 -12
- package/src/hooks/skill-keywords.ts +9 -0
- package/src/hooks/skill-state.ts +41 -10
- package/src/hooks/ui-skill-keywords.ts +67 -10
- package/src/internal-urls/docs-index.generated.ts +1 -1
- package/src/sdk/session.ts +5 -82
- package/src/session/agent-session.ts +137 -37
- package/src/session/auth-storage-discovery.ts +83 -0
package/CHANGELOG.md
CHANGED
|
@@ -2,6 +2,29 @@
|
|
|
2
2
|
|
|
3
3
|
## [Unreleased]
|
|
4
4
|
|
|
5
|
+
## [0.6.0] - 2026-09-24
|
|
6
|
+
|
|
7
|
+
### Added
|
|
8
|
+
|
|
9
|
+
- The 13 bundled UI skills are now part of the typed decision instead of 20 hand-written regexes. Measured recall of the pattern table on real Korean prompts was 5/8; a prompt like "이 부분 손보기 전에 어떻게 갈지부터 같이 정하고 넘어가자" matched nothing and the skill never activated. The picked skill still arrives as the same hidden `developer` reminder, now labelled `semantic match` when the model answered and `matched` when a pattern did.
|
|
10
|
+
- Workflow routing and UI skill selection are **one** decision call, not two. The triage asks both questions in a single request (`questions: ["workflow", "uiSkill"]`), and a prompt the pattern table already answered drops that question from the request instead of paying for it — so adding UI routing cost zero extra round trips.
|
|
11
|
+
- The skill keyword table now fills itself. Every answered routing decision mines ordered two-stem shingles from the prompt into `~/.skc/agent/learned-skill-keywords.json`. A shingle is promoted to a real keyword after 2 independent prompts when the backend reports calibrated confidence ≥ 0.9, or 3 when it only ranks; a single contradicting answer blocks that shingle permanently, so a coincidence cannot become a rule. Learned entries carry strictly negative priority — a hand-written keyword always wins — and a promoted pattern answers for free on the next matching prompt. `decisions.keywordLearning: false` turns the store off; `SKC_LEARNED_KEYWORDS_PATH` relocates it.
|
|
12
|
+
- Routing and learning work in whatever language the prompt is written in. Mining uses the platform Unicode segmenter, so Chinese, Japanese and Thai prompts with no spaces come apart into words (Han runs the ICU dictionary misses fall back to character bigrams; hiragana-only tokens are dropped as grammar); non-ASCII alphabets are kept instead of discarded, and a learned stem gets a Unicode word boundary in spaced scripts and a bare substring match in unspaced ones. The length floor that skips the model on "ok thanks" is measured in signal rather than code units, so a six-character Chinese request is no longer thrown away. Measured against the hosted `jev` backend: 40/40 across ko, en, zh, ja, es and ru, and a ralplan phrasing repeated twice in Chinese, Japanese or Russian promotes a keyword that fires on an unseen sentence in that language.
|
|
13
|
+
- The Codex host gets the same router. `skc codex-native-hook` ran the hand-written keyword table and nothing else: the semantic stage had a parameter waiting for it since it shipped, and learned keywords were never loaded there, so a phrasing learned in a CLI session did nothing under Codex. The hook now runs all three stages — hand-written keywords, learned keywords, then one typed decision for whatever those missed — on the same backend chain (TypeSafe key, else a running local runtime, else the small model of the provider you are logged into) and feeds the answer back into the learned table. Because the hook is a fresh process per prompt, credentials and the model registry are opened only on a prompt the tables did not fully answer and closed before it exits; a prompt they did answer still costs the same ~0.3s it did before. Measured through the real hook subprocess: a Chinese ralplan phrasing routed by the model on the first prompt fired from the learned table on the next one.
|
|
14
|
+
|
|
15
|
+
### Changed
|
|
16
|
+
|
|
17
|
+
- Typed decisions (`decisions.enabled`) are now **on by default**. The keyword stage of workflow routing still runs first and free; the model stage only spends a small-model call on turns the table did not answer. Set `decisions.enabled: false` to route on the keyword table alone.
|
|
18
|
+
- The small-model fallback now ranks the **provider you are chatting with** first. Measured on a registry with Anthropic, Codex and Z.ai keys: price order put four free `zai/glm-*` reasoning tiers ahead of `claude-haiku-4-5`, and the first took 8.9s — past the 8s service deadline, so the answer was nothing. Haiku next to Opus, mini next to GPT is also the model you would have picked by hand.
|
|
19
|
+
- The routing decision now goes through the session's own `streamFn` instead of a bare `completeSimple`, so it gets the same proxy, credential-invalidation and retry wrapping as the turn, and a host that injects a transport intercepts it. Hosts opt sessions in with `AgentSessionConfig.semanticWorkflowRouting` (`createAgentSession` does); a bare `AgentSession` never routes, so scripted transports are not consumed by a call they did not script. While the decision is in flight the session reports `isStreaming`, so a second `prompt()` queues or rejects as it would mid-turn, and `abort()` cancels the decision.
|
|
20
|
+
- `task.modelRouting.enabled` now defaults to **true**. The tier lists it reads are empty until you set them, so nothing changes for anyone who has not configured them — but shipping it off meant a user who *had* configured tiers still got no routing, which made the whole feature dead code.
|
|
21
|
+
|
|
22
|
+
### Fixed
|
|
23
|
+
|
|
24
|
+
- The small-model fallback answered **nothing** on any registry with an Anthropic key: the cheapest qualifying entry, `claude-3-haiku-20240307`, is retired and returns 404, and the backend logged it as "model did not emit the forced tool call" and gave up. The three retired Claude 3.x ids are now dropped from the catalog *and* from stale on-disk cache rows, the backend logs the real status, and a candidate that errors or blows a 3.5s attempt budget is skipped for the next one and remembered for ten minutes so the stall is paid once, not per prompt.
|
|
25
|
+
- A local runtime contributes **one** candidate — the smallest chat model — and embedding models (`nomic-embed-text`) are no longer offered a chat call. Measured with Ollama: five local models ranked ahead of every hosted one, each timed out on cold start, and the hosted model was never reached before the deadline.
|
|
26
|
+
- A decision request whose caller had already aborted — measured: `abort()` landing during the previous backend's credential lookup — ran a full attempt budget of silence, because an abort listener added to a signal that has already fired never runs. Every backend now checks the signal after each await.
|
|
27
|
+
|
|
5
28
|
## [0.5.26] - 2026-09-23
|
|
6
29
|
### Added
|
|
7
30
|
|
|
@@ -2419,11 +2419,29 @@ export declare const SETTINGS_SCHEMA: {
|
|
|
2419
2419
|
};
|
|
2420
2420
|
readonly "decisions.enabled": {
|
|
2421
2421
|
readonly type: "boolean";
|
|
2422
|
-
readonly default:
|
|
2422
|
+
readonly default: true;
|
|
2423
2423
|
readonly ui: {
|
|
2424
2424
|
readonly tab: "context";
|
|
2425
2425
|
readonly label: "Typed decisions";
|
|
2426
|
-
readonly description: "
|
|
2426
|
+
readonly description: "Model-backed second stage for workflow routing. The keyword table runs first on every turn and costs nothing; this handles the phrasings it cannot express, which is most wording that is not a literal match. Costs one small model call, only on turns the keyword table did not already answer, on the cheapest backend available: TypeSafe when a key is stored, else a local runtime that is running, else the small model of the provider you are chatting with. Any failure falls back to keyword-only behaviour, but a successful answer can also select a different workflow than the deep-interview ambiguity detector would have. Turn off to route on the keyword table alone.";
|
|
2427
|
+
};
|
|
2428
|
+
};
|
|
2429
|
+
/**
|
|
2430
|
+
* Feed routing answers back into the deterministic keyword table.
|
|
2431
|
+
*
|
|
2432
|
+
* The hand-written table cannot be grown by hand for Korean — measured recall
|
|
2433
|
+
* was 0/9 — so it grows itself instead: a two-stem pattern that produced the
|
|
2434
|
+
* same routing answer on two distinct prompts, and was never contradicted, is
|
|
2435
|
+
* promoted and thereafter fires for free. One contradiction retracts it
|
|
2436
|
+
* permanently. Stems and hashes are stored; prompt text never is.
|
|
2437
|
+
*/
|
|
2438
|
+
readonly "decisions.keywordLearning": {
|
|
2439
|
+
readonly type: "boolean";
|
|
2440
|
+
readonly default: true;
|
|
2441
|
+
readonly ui: {
|
|
2442
|
+
readonly tab: "context";
|
|
2443
|
+
readonly label: "Learn routing keywords";
|
|
2444
|
+
readonly description: "Remember the phrasings the routing model resolves, so repeating one stops costing a model call. A pattern must give the same answer on two different prompts before it fires on its own, and a single disagreement removes it for good. Only word stems and hashes are written to disk, never your prompts. Turn off to keep the keyword table frozen at its built-in entries.";
|
|
2427
2445
|
};
|
|
2428
2446
|
};
|
|
2429
2447
|
/**
|
|
@@ -2435,12 +2453,14 @@ export declare const SETTINGS_SCHEMA: {
|
|
|
2435
2453
|
* mid-session invalidates the prompt cache, which on a long context costs
|
|
2436
2454
|
* more than the cheaper tier saves.
|
|
2437
2455
|
*
|
|
2438
|
-
* Needs `decisions.enabled` and at least two tiers configured.
|
|
2439
|
-
*
|
|
2456
|
+
* Needs `decisions.enabled` and at least two tiers configured. The tiers are
|
|
2457
|
+
* empty by default, so this is a no-op until the user sets them — which is why
|
|
2458
|
+
* it defaults on: shipping it off meant the routing code existed and never ran
|
|
2459
|
+
* even for users who had configured the tiers it needs.
|
|
2440
2460
|
*/
|
|
2441
2461
|
readonly "task.modelRouting.enabled": {
|
|
2442
2462
|
readonly type: "boolean";
|
|
2443
|
-
readonly default:
|
|
2463
|
+
readonly default: true;
|
|
2444
2464
|
readonly ui: {
|
|
2445
2465
|
readonly tab: "tasks";
|
|
2446
2466
|
readonly label: "Route subagent models per task";
|
|
@@ -0,0 +1,61 @@
|
|
|
1
|
+
import type { SkillKeywordDefinition } from "../hooks/skill-keywords";
|
|
2
|
+
import type { CanonicalSkcWorkflowSkill } from "../skill-state/active-state";
|
|
3
|
+
export declare const LEARNED_KEYWORD_STORE_VERSION = 1;
|
|
4
|
+
/** Test seam. Also lets a host keep the table out of the user's home. */
|
|
5
|
+
export declare function setLearnedKeywordStorePath(value: string | undefined): void;
|
|
6
|
+
export declare function resetLearnedKeywordCache(): void;
|
|
7
|
+
/**
|
|
8
|
+
* Where the learned table lives.
|
|
9
|
+
*
|
|
10
|
+
* User-global on purpose: a phrasing you use is a phrasing you use, and making it
|
|
11
|
+
* per-repository would mean relearning the same rule in every checkout.
|
|
12
|
+
* `SKC_LEARNED_KEYWORDS_PATH` redirects it, so a test driving a real session under
|
|
13
|
+
* a temp agent dir cannot write into the developer's actual home.
|
|
14
|
+
*/
|
|
15
|
+
export declare function getLearnedKeywordStorePath(): string;
|
|
16
|
+
/**
|
|
17
|
+
* Ordered stem pairs within a three-token window.
|
|
18
|
+
*
|
|
19
|
+
* The gap is what makes a learned pattern generalize where the literal table
|
|
20
|
+
* cannot: a prompt that says "make me a plan document" mines `plan … document`,
|
|
21
|
+
* which then also matches "make the plan into a document" — the phrasing that
|
|
22
|
+
* made the hand-written entry miss.
|
|
23
|
+
*/
|
|
24
|
+
export declare function mineShingles(text: string): string[];
|
|
25
|
+
/**
|
|
26
|
+
* Promoted patterns, in the shape the deterministic stage consumes.
|
|
27
|
+
*
|
|
28
|
+
* Priority sits one below the hand-written entry for the same workflow: when a
|
|
29
|
+
* learned pattern and an enumerated keyword disagree, the human wins.
|
|
30
|
+
*/
|
|
31
|
+
export declare function loadLearnedKeywordDefinitions(): Promise<SkillKeywordDefinition[]>;
|
|
32
|
+
export interface RoutingObservation {
|
|
33
|
+
text: string;
|
|
34
|
+
/** Null means the router deliberately chose no workflow — a negative example. */
|
|
35
|
+
skill: CanonicalSkcWorkflowSkill | null;
|
|
36
|
+
confidence?: number | undefined;
|
|
37
|
+
calibrated?: boolean | undefined;
|
|
38
|
+
}
|
|
39
|
+
export interface LearningOutcome {
|
|
40
|
+
promoted: SkillKeywordDefinition[];
|
|
41
|
+
retracted: number;
|
|
42
|
+
}
|
|
43
|
+
/**
|
|
44
|
+
* Feed one routing answer into the table.
|
|
45
|
+
*
|
|
46
|
+
* Best-effort by construction: it runs after the turn's routing decision is
|
|
47
|
+
* already made, so a failure here changes nothing the user can observe.
|
|
48
|
+
*/
|
|
49
|
+
export declare function observeRouting(observation: RoutingObservation): Promise<LearningOutcome>;
|
|
50
|
+
export interface LearnedKeywordSummary {
|
|
51
|
+
entries: {
|
|
52
|
+
keyword: string;
|
|
53
|
+
skill: CanonicalSkcWorkflowSkill;
|
|
54
|
+
positives: number;
|
|
55
|
+
lastSeen: number;
|
|
56
|
+
}[];
|
|
57
|
+
candidates: number;
|
|
58
|
+
blocked: number;
|
|
59
|
+
}
|
|
60
|
+
/** What the table has actually learned, for `skc` surfaces and tests. */
|
|
61
|
+
export declare function summarizeLearnedKeywords(): Promise<LearnedKeywordSummary>;
|
|
@@ -16,7 +16,7 @@
|
|
|
16
16
|
* 2. Every question goes in **one** call. Splitting them multiplies cost and latency
|
|
17
17
|
* while the enum constraint already keeps each field independent.
|
|
18
18
|
*/
|
|
19
|
-
import { type Api, type Model } from "@sayknow-cli/ai";
|
|
19
|
+
import { type Api, completeSimple, type Model } from "@sayknow-cli/ai";
|
|
20
20
|
import type { ModelRegistry } from "../config/model-registry";
|
|
21
21
|
import type { Settings } from "../config/settings";
|
|
22
22
|
import { type DecisionBackend } from "./types";
|
|
@@ -40,12 +40,24 @@ export interface LlmBackendDeps {
|
|
|
40
40
|
maxInputCostPerMTok?: number;
|
|
41
41
|
/** Injected in tests to make the local-runtime probe deterministic. */
|
|
42
42
|
fetchImpl?: typeof fetch;
|
|
43
|
+
/** Injected in tests to script provider answers without a network. */
|
|
44
|
+
completeImpl?: typeof completeSimple;
|
|
45
|
+
/** Per-candidate deadline; injected in tests so a hung provider does not cost real seconds. */
|
|
46
|
+
attemptTimeoutMs?: number;
|
|
43
47
|
registry: ModelRegistry;
|
|
44
48
|
settings: Settings;
|
|
45
49
|
sessionId?: string;
|
|
46
50
|
/** Overrides role resolution; used by callers that already picked a model. */
|
|
47
51
|
model?: Model<Api>;
|
|
52
|
+
/**
|
|
53
|
+
* Provider of the model the caller is already talking to. Its small model is tried
|
|
54
|
+
* first when nothing was configured; see `rankSmallModels` for why. A thunk is
|
|
55
|
+
* accepted so a long-lived backend follows the session when the user switches model.
|
|
56
|
+
*/
|
|
57
|
+
preferredProvider?: string | (() => string | undefined);
|
|
48
58
|
}
|
|
49
59
|
/** Reset between tests; also lets a caller force a fresh probe after starting a runtime. */
|
|
50
60
|
export declare function clearLocalRuntimeLivenessCache(): void;
|
|
61
|
+
/** Reset between tests. */
|
|
62
|
+
export declare function clearDeadModelCache(): void;
|
|
51
63
|
export declare function createLlmDecisionBackend(deps: LlmBackendDeps): DecisionBackend;
|
|
@@ -0,0 +1,42 @@
|
|
|
1
|
+
import { type BundledSkcUiSkillName } from "../defaults/skc-ui-skills";
|
|
2
|
+
import { type CanonicalSkcWorkflowSkill } from "../skill-state/active-state";
|
|
3
|
+
import type { DecisionService } from "./index";
|
|
4
|
+
/** Exported so tests can assert the contract the model is actually given. */
|
|
5
|
+
export declare function buildUiSkillCriteria(): Record<string, string>;
|
|
6
|
+
export interface PromptTriage {
|
|
7
|
+
/** Null means the router deliberately chose no workflow, not that it failed. */
|
|
8
|
+
workflow: CanonicalSkcWorkflowSkill | null;
|
|
9
|
+
uiSkill: BundledSkcUiSkillName | null;
|
|
10
|
+
/** Only meaningful when {@link calibrated} is true. */
|
|
11
|
+
workflowConfidence: number | undefined;
|
|
12
|
+
calibrated: boolean;
|
|
13
|
+
}
|
|
14
|
+
export interface PromptTriageRequest {
|
|
15
|
+
text: string;
|
|
16
|
+
/** Skip the workflow question — the keyword table already answered it. */
|
|
17
|
+
skipWorkflow?: boolean;
|
|
18
|
+
/** Skip the UI question — the regex table already matched. */
|
|
19
|
+
skipUiSkill?: boolean;
|
|
20
|
+
signal?: AbortSignal | undefined;
|
|
21
|
+
}
|
|
22
|
+
export type PromptTriager = (request: PromptTriageRequest) => Promise<PromptTriage | null>;
|
|
23
|
+
/**
|
|
24
|
+
* Build the per-turn triager.
|
|
25
|
+
*
|
|
26
|
+
* Returns null when the service is disabled, the prompt is too short to carry
|
|
27
|
+
* intent, every question was already answered for free, or the backend did not
|
|
28
|
+
* produce a usable result. Null means "no information", which is different from
|
|
29
|
+
* a result whose fields are all null — that one is the router saying "none", and
|
|
30
|
+
* the keyword learner treats it as a negative example.
|
|
31
|
+
*/
|
|
32
|
+
export declare function createPromptTriage(service: DecisionService): PromptTriager;
|
|
33
|
+
export type SkillRouter = (text: string, signal?: AbortSignal) => Promise<CanonicalSkcWorkflowSkill | null>;
|
|
34
|
+
/**
|
|
35
|
+
* Workflow-only view of the triager.
|
|
36
|
+
*
|
|
37
|
+
* A few lines over the same implementation rather than a second one: the eval
|
|
38
|
+
* harness in `scripts/eval-skill-routing.ts` and the routing tests want
|
|
39
|
+
* text-in/skill-out and have no UI question to ask, and a parallel router would
|
|
40
|
+
* be free to drift away from the thresholds this one enforces.
|
|
41
|
+
*/
|
|
42
|
+
export declare function createSemanticSkillRouter(service: DecisionService): SkillRouter;
|
|
@@ -1,10 +1,45 @@
|
|
|
1
|
-
|
|
2
|
-
|
|
1
|
+
/** Shared across every routing question so one answer shape covers them all. */
|
|
2
|
+
export declare const NONE_CHOICE = "none";
|
|
3
|
+
export declare const ROUTING_INSTRUCTIONS = "Which workflow should handle this user request? Choose none unless the request clearly calls for one of the workflows.";
|
|
3
4
|
/** Exported so tests can assert the contract the model is actually given. */
|
|
4
5
|
export declare function buildRoutingCriteria(): Record<string, string>;
|
|
5
|
-
export type SkillRouter = (text: string) => Promise<CanonicalSkcWorkflowSkill | null>;
|
|
6
6
|
/**
|
|
7
|
-
*
|
|
8
|
-
*
|
|
7
|
+
* Prompts below this length never carry enough signal to justify a model round-trip.
|
|
8
|
+
* Measured with {@link promptSignalLength}, not `String.length`: the floor was fitted
|
|
9
|
+
* to English and Korean, and in code units a complete Chinese request is shorter than
|
|
10
|
+
* "ok thanks".
|
|
9
11
|
*/
|
|
10
|
-
export declare
|
|
12
|
+
export declare const MIN_PROMPT_CHARS = 12;
|
|
13
|
+
/** Only the opening of a prompt decides its workflow; the rest is payload. */
|
|
14
|
+
export declare const MAX_PROMPT_CHARS = 4000;
|
|
15
|
+
/**
|
|
16
|
+
* Prompt length in Latin-letter equivalents.
|
|
17
|
+
*
|
|
18
|
+
* One Han character, kana or Hangul syllable carries what two or three Latin letters
|
|
19
|
+
* do, so each counts double. "先做架构设计" (6 code units) is a whole request; "谢谢" and
|
|
20
|
+
* "고마워요" still fall under the floor, which is the point of having one.
|
|
21
|
+
*/
|
|
22
|
+
export declare function promptSignalLength(text: string): number;
|
|
23
|
+
/**
|
|
24
|
+
* Minimum calibrated confidence required to activate a workflow.
|
|
25
|
+
*
|
|
26
|
+
* Activation is a strong move: it switches on the mutation guard, the Stop hook and the
|
|
27
|
+
* ask tool. Getting it wrong is worse than missing, because the user did not ask for any
|
|
28
|
+
* of that and has no obvious way to see why it appeared.
|
|
29
|
+
*
|
|
30
|
+
* Measured over ten routing prompts against the hosted model: every answer it reported
|
|
31
|
+
* at 1.00 was correct, and its single wrong answer reported 0.71. The lowest *correct*
|
|
32
|
+
* confidence was 0.67 — and that case was "none", so gating it out costs nothing. A
|
|
33
|
+
* floor here therefore removes the observed error without removing a real activation.
|
|
34
|
+
*
|
|
35
|
+
* One prompt sits close to this line. "추측하지 말고 모르는 건 다 물어봐" resolves to
|
|
36
|
+
* deep-interview in 8/8 samples but at 0.76-0.83, so the floor has roughly 0.01 of
|
|
37
|
+
* headroom on it. Raising the floor would drop a correct activation; lowering it would
|
|
38
|
+
* re-admit the 0.71 error. Treat 0.75 as fitted to a small sample and re-derive it from
|
|
39
|
+
* real usage rather than nudging it on a hunch.
|
|
40
|
+
*
|
|
41
|
+
* Only applied when the backend reports `calibrated: true`. An ordinary LLM answering
|
|
42
|
+
* through a forced enum has no meaningful confidence to compare against, so gating on a
|
|
43
|
+
* number it did not really produce would just be superstition.
|
|
44
|
+
*/
|
|
45
|
+
export declare const MIN_CALIBRATED_CONFIDENCE = 0.75;
|
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
import type { PromptTriager } from "../decisions/prompt-triage";
|
|
2
|
+
import { type SkillActiveState } from "./skill-state";
|
|
3
|
+
export interface NativePromptRoutingInput {
|
|
4
|
+
cwd: string;
|
|
5
|
+
text: string;
|
|
6
|
+
sessionId?: string;
|
|
7
|
+
threadId?: string;
|
|
8
|
+
turnId?: string;
|
|
9
|
+
stateDir?: string;
|
|
10
|
+
/** `config.yml` files in precedence order, lowest first; resolved by the hook. */
|
|
11
|
+
configPaths: readonly string[];
|
|
12
|
+
agentDir?: string;
|
|
13
|
+
/** Injected in tests so routing is exercised without credentials or a network. */
|
|
14
|
+
triager?: PromptTriager;
|
|
15
|
+
}
|
|
16
|
+
export interface NativePromptRoutingResult {
|
|
17
|
+
skillState: SkillActiveState | null;
|
|
18
|
+
/** Directive for a UI skill the model picked because the pattern table missed. */
|
|
19
|
+
uiSkillContext: string | null;
|
|
20
|
+
}
|
|
21
|
+
export declare function routeNativePrompt(input: NativePromptRoutingInput): Promise<NativePromptRoutingResult>;
|
|
@@ -1,3 +1,4 @@
|
|
|
1
|
+
import type { PromptTriager } from "../decisions/prompt-triage";
|
|
1
2
|
import { type EffectiveSkillConfigInput } from "./skill-state";
|
|
2
3
|
export type SkcNativeHookEventName = "UserPromptSubmit" | "Stop";
|
|
3
4
|
export interface SkcNativeHookDispatchResult {
|
|
@@ -11,6 +12,8 @@ interface SkcNativeHookDispatchOptions {
|
|
|
11
12
|
stateDir?: string;
|
|
12
13
|
effectiveSkillConfig?: EffectiveSkillConfigInput;
|
|
13
14
|
configPaths?: string[];
|
|
15
|
+
/** Injected in tests so the model stage runs without credentials or a network. */
|
|
16
|
+
triager?: PromptTriager;
|
|
14
17
|
}
|
|
15
18
|
export declare function clearSkcNativeSkillHookCachesForTesting(): void;
|
|
16
19
|
export declare function getSkcNativeSkillHookCacheStatsForTesting(): {
|
|
@@ -4,6 +4,15 @@ export interface SkillKeywordDefinition {
|
|
|
4
4
|
skill: SkcWorkflowSkill;
|
|
5
5
|
priority: number;
|
|
6
6
|
guidance: string;
|
|
7
|
+
/**
|
|
8
|
+
* Pre-compiled matcher, replacing the literal-substring compilation of
|
|
9
|
+
* `keyword`. Only learned entries set it: they are two stems with a bounded
|
|
10
|
+
* gap, which a literal string cannot express. `keyword` stays human-readable
|
|
11
|
+
* for logs and for the guidance line.
|
|
12
|
+
*/
|
|
13
|
+
pattern?: RegExp;
|
|
14
|
+
/** Mined from routing answers rather than written by hand. */
|
|
15
|
+
learned?: boolean;
|
|
7
16
|
}
|
|
8
17
|
export declare const SKC_WORKFLOW_SKILLS: readonly ["deep-interview", "ralplan", "ultragoal", "team"];
|
|
9
18
|
export type SkcWorkflowSkill = CanonicalSkcWorkflowSkill;
|
|
@@ -2,7 +2,7 @@ import type { SkillDiscoverySettings } from "../config/skill-settings-defaults";
|
|
|
2
2
|
import { type SkillActiveState } from "../skill-state/active-state";
|
|
3
3
|
import { initialPhaseForSkill } from "../skill-state/initial-phase";
|
|
4
4
|
export { initialPhaseForSkill };
|
|
5
|
-
import { type SkcWorkflowSkill } from "./skill-keywords";
|
|
5
|
+
import { type SkcWorkflowSkill, type SkillKeywordDefinition } from "./skill-keywords";
|
|
6
6
|
export declare const SKC_STATE_DIR = ".skc";
|
|
7
7
|
export declare const SKILL_ACTIVE_STATE_FILE = "skill-active-state.json";
|
|
8
8
|
export interface EffectiveSkillConfigInput {
|
|
@@ -15,6 +15,8 @@ export interface SkillKeywordMatch {
|
|
|
15
15
|
keyword: string;
|
|
16
16
|
skill: SkcWorkflowSkill;
|
|
17
17
|
priority: number;
|
|
18
|
+
/** Matched a pattern mined from routing answers, not a hand-written keyword. */
|
|
19
|
+
learned?: boolean;
|
|
18
20
|
}
|
|
19
21
|
export type { SkillActiveEntry, SkillActiveState } from "../skill-state/active-state";
|
|
20
22
|
export interface ModeState {
|
|
@@ -38,6 +40,12 @@ export interface RecordSkillActivationInput {
|
|
|
38
40
|
turnId?: string;
|
|
39
41
|
nowIso?: string;
|
|
40
42
|
stateDir?: string;
|
|
43
|
+
/**
|
|
44
|
+
* Patterns promoted by the keyword learner, appended after the hand-written
|
|
45
|
+
* table. The caller loads them because it knows whether learning is on; see
|
|
46
|
+
* `detectSkillKeywords`.
|
|
47
|
+
*/
|
|
48
|
+
learned?: readonly SkillKeywordDefinition[];
|
|
41
49
|
/**
|
|
42
50
|
* Semantic fallback, consulted only when no keyword matched. Supplying it turns the
|
|
43
51
|
* literal keyword table into a two-stage router; omitting it keeps the historical
|
|
@@ -60,8 +68,17 @@ export interface UserPromptSubmitStateInput {
|
|
|
60
68
|
prompt?: string;
|
|
61
69
|
sessionFile?: string;
|
|
62
70
|
}
|
|
63
|
-
|
|
64
|
-
|
|
71
|
+
/**
|
|
72
|
+
* Match a prompt against the keyword table.
|
|
73
|
+
*
|
|
74
|
+
* `learned` carries patterns mined from semantic routing answers (see
|
|
75
|
+
* `decisions/keyword-learning.ts`). It is a parameter rather than a module-level
|
|
76
|
+
* load because this file is imported by the hook process, where a synchronous
|
|
77
|
+
* disk read on every prompt is not acceptable and the caller already knows
|
|
78
|
+
* whether learning is enabled.
|
|
79
|
+
*/
|
|
80
|
+
export declare function detectSkillKeywords(text: string, learned?: readonly SkillKeywordDefinition[]): SkillKeywordMatch[];
|
|
81
|
+
export declare function detectPrimarySkillKeyword(text: string, learned?: readonly SkillKeywordDefinition[]): SkillKeywordMatch | null;
|
|
65
82
|
export declare function resolveSkcStateDir(cwd: string, stateDir?: string): string;
|
|
66
83
|
export interface StateRecoveryDiagnostic {
|
|
67
84
|
kind: "skill-active-state" | "mode-state";
|
|
@@ -43,6 +43,21 @@ export declare function detectUiSkillKeywords(text: string): UiSkillKeywordMatch
|
|
|
43
43
|
* UI/UX work.
|
|
44
44
|
*/
|
|
45
45
|
export declare function buildUiSkillActivationContext(text: string): string | null;
|
|
46
|
+
/**
|
|
47
|
+
* Same directive, for a skill chosen by the semantic stage rather than by a
|
|
48
|
+
* pattern. The regex table caught 5 of 8 real frontend prompts; this is the path
|
|
49
|
+
* for the other three, and it names its source so a wrong pick is traceable to
|
|
50
|
+
* the model rather than to a pattern nobody can find.
|
|
51
|
+
*/
|
|
52
|
+
export declare function buildUiSkillDirectiveForSkill(skill: BundledSkcUiSkillName): string;
|
|
53
|
+
/**
|
|
54
|
+
* What each bundled skill is *for*, in the words a user would recognise.
|
|
55
|
+
*
|
|
56
|
+
* This is the whole contract with the routing model — the skill ids alone carry
|
|
57
|
+
* almost no signal, and `appllama-app-design-skill` carries actively misleading
|
|
58
|
+
* signal. Kept next to the patterns so the two cannot drift apart.
|
|
59
|
+
*/
|
|
60
|
+
export declare const BUNDLED_UI_SKILL_MEANINGS: Record<BundledSkcUiSkillName, string>;
|
|
46
61
|
/**
|
|
47
62
|
* Frontend skills SKC routes to but deliberately does NOT vendor.
|
|
48
63
|
*
|
|
@@ -19,7 +19,8 @@ import { type LocalProtocolOptions } from "../internal-urls";
|
|
|
19
19
|
import { AgentRegistry } from "../registry/agent-registry";
|
|
20
20
|
import { MCPManager } from "../runtime-mcp";
|
|
21
21
|
import { AgentSession, type ForkContextSeed } from "../session/agent-session";
|
|
22
|
-
import { AuthStorage } from "../session/auth-storage";
|
|
22
|
+
import type { AuthStorage } from "../session/auth-storage";
|
|
23
|
+
import { discoverAuthStorage } from "../session/auth-storage-discovery";
|
|
23
24
|
import { SessionManager } from "../session/session-manager";
|
|
24
25
|
import { type BuildSystemPromptResult } from "../system-prompt";
|
|
25
26
|
import { BashTool, BUILTIN_TOOLS, createTools, EditTool, EvalTool, FindTool, HIDDEN_TOOLS, type LspStartupServerInfo, loadSshTool, ReadTool, ResolveTool, SearchTool, type Tool, type ToolSession, WebSearchTool, WriteTool } from "../tools";
|
|
@@ -218,18 +219,7 @@ export type { FileSlashCommand } from "../extensibility/slash-commands";
|
|
|
218
219
|
export type { Tool } from "../tools";
|
|
219
220
|
export { buildDirectoryTree, buildWorkspaceTree, type DirectoryTree, type WorkspaceTree } from "../workspace-tree";
|
|
220
221
|
export { BashTool, BUILTIN_TOOLS, createTools, EditTool, EvalTool, FindTool, HIDDEN_TOOLS, loadSshTool, ReadTool, ResolveTool, SearchTool, type ToolSession, WebSearchTool, WriteTool, };
|
|
221
|
-
|
|
222
|
-
* Create an AuthStorage instance.
|
|
223
|
-
*
|
|
224
|
-
* Default: local SQLite store at `<agentDir>/agent.db`.
|
|
225
|
-
*
|
|
226
|
-
* Broker mode: when `SKC_AUTH_BROKER_URL` is set, credentials are pulled from
|
|
227
|
-
* a remote auth-broker over the wire. Refresh tokens never leave the broker;
|
|
228
|
-
* the client receives access tokens with `refresh = "__remote__"` and calls
|
|
229
|
-
* back into the broker through the {@link AuthStorageOptions.refreshOAuthCredential}
|
|
230
|
-
* override to re-mint access tokens when needed.
|
|
231
|
-
*/
|
|
232
|
-
export declare function discoverAuthStorage(agentDir?: string): Promise<AuthStorage>;
|
|
222
|
+
export { discoverAuthStorage };
|
|
233
223
|
/**
|
|
234
224
|
* Discover extensions from cwd.
|
|
235
225
|
*/
|
|
@@ -223,6 +223,14 @@ export interface AgentSessionConfig {
|
|
|
223
223
|
toolRegistry?: Map<string, AgentTool>;
|
|
224
224
|
/** Tool-session factory context used to lazily attach workflow-gate-only tools. */
|
|
225
225
|
workflowGateToolSession?: ToolSession;
|
|
226
|
+
/**
|
|
227
|
+
* Stage two of workflow routing: before a genuine user turn, ask a small model
|
|
228
|
+
* through the session's own transport which SKC workflow the prompt calls for.
|
|
229
|
+
* Hosts that own an interactive user (`createAgentSession`) turn this on; a bare
|
|
230
|
+
* session leaves it off so a scripted or injected transport is never consumed by a
|
|
231
|
+
* call the host did not script. `decisions.enabled` still governs the user side.
|
|
232
|
+
*/
|
|
233
|
+
semanticWorkflowRouting?: boolean;
|
|
226
234
|
/** Current session pre-LLM message transform pipeline */
|
|
227
235
|
transformContext?: (messages: AgentMessage[], signal?: AbortSignal) => AgentMessage[] | Promise<AgentMessage[]>;
|
|
228
236
|
/** Provider payload hook used by the active session request path */
|
|
@@ -0,0 +1,13 @@
|
|
|
1
|
+
import { AuthStorage } from "./auth-storage";
|
|
2
|
+
/**
|
|
3
|
+
* Create an AuthStorage instance.
|
|
4
|
+
*
|
|
5
|
+
* Default: local SQLite store at `<agentDir>/agent.db`.
|
|
6
|
+
*
|
|
7
|
+
* Broker mode: when `SKC_AUTH_BROKER_URL` is set, credentials are pulled from
|
|
8
|
+
* a remote auth-broker over the wire. Refresh tokens never leave the broker;
|
|
9
|
+
* the client receives access tokens with `refresh = "__remote__"` and calls
|
|
10
|
+
* back into the broker through the {@link AuthStorageOptions.refreshOAuthCredential}
|
|
11
|
+
* override to re-mint access tokens when needed.
|
|
12
|
+
*/
|
|
13
|
+
export declare function discoverAuthStorage(agentDir?: string): Promise<AuthStorage>;
|
|
@@ -8,9 +8,9 @@ export { extractReadableFromHtml, type ReadableFormat, type ReadableResult } fro
|
|
|
8
8
|
export type { Observation, ObservationEntry } from "./browser/tab-protocol";
|
|
9
9
|
declare const browserSchema: z.ZodObject<{
|
|
10
10
|
action: z.ZodEnum<{
|
|
11
|
+
run: "run";
|
|
11
12
|
open: "open";
|
|
12
13
|
close: "close";
|
|
13
|
-
run: "run";
|
|
14
14
|
act: "act";
|
|
15
15
|
}>;
|
|
16
16
|
name: z.ZodOptional<z.ZodString>;
|
|
@@ -122,9 +122,9 @@ export declare class BrowserTool implements AgentTool<typeof browserSchema, Brow
|
|
|
122
122
|
readonly summary = "Control a headless browser to navigate and interact with web pages";
|
|
123
123
|
readonly parameters: z.ZodObject<{
|
|
124
124
|
action: z.ZodEnum<{
|
|
125
|
+
run: "run";
|
|
125
126
|
open: "open";
|
|
126
127
|
close: "close";
|
|
127
|
-
run: "run";
|
|
128
128
|
act: "act";
|
|
129
129
|
}>;
|
|
130
130
|
name: z.ZodOptional<z.ZodString>;
|
package/package.json
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"type": "module",
|
|
3
3
|
"name": "@sayknow-cli/coding-agent",
|
|
4
|
-
"version": "0.
|
|
4
|
+
"version": "0.6.0",
|
|
5
5
|
"description": "Sayknow-CLI CLI with read, bash, edit, write tools and session management",
|
|
6
6
|
"homepage": "https://sayknow-cli.com",
|
|
7
7
|
"author": "jaybeyond",
|
|
@@ -54,12 +54,12 @@
|
|
|
54
54
|
"@agentclientprotocol/sdk": "1.3.0",
|
|
55
55
|
"@babel/parser": "^7.29.3",
|
|
56
56
|
"@mozilla/readability": "^0.6.0",
|
|
57
|
-
"@sayknow-cli/stats": "0.
|
|
58
|
-
"@sayknow-cli/agent-core": "0.
|
|
59
|
-
"@sayknow-cli/ai": "0.
|
|
60
|
-
"@sayknow-cli/natives": "0.
|
|
61
|
-
"@sayknow-cli/tui": "0.
|
|
62
|
-
"@sayknow-cli/utils": "0.
|
|
57
|
+
"@sayknow-cli/stats": "0.6.0",
|
|
58
|
+
"@sayknow-cli/agent-core": "0.6.0",
|
|
59
|
+
"@sayknow-cli/ai": "0.6.0",
|
|
60
|
+
"@sayknow-cli/natives": "0.6.0",
|
|
61
|
+
"@sayknow-cli/tui": "0.6.0",
|
|
62
|
+
"@sayknow-cli/utils": "0.6.0",
|
|
63
63
|
"@puppeteer/browsers": "^2.13.0",
|
|
64
64
|
"@types/turndown": "5.0.6",
|
|
65
65
|
"@xterm/headless": "^6.0.0",
|
|
@@ -19,7 +19,7 @@ import { ModelRegistry } from "../src/config/model-registry";
|
|
|
19
19
|
import { resolveRoleSelection } from "../src/config/model-resolver";
|
|
20
20
|
import { Settings } from "../src/config/settings";
|
|
21
21
|
import { createDecisionService, createLlmDecisionBackend, createTypeSafeDecisionBackend } from "../src/decisions";
|
|
22
|
-
import { createSemanticSkillRouter } from "../src/decisions/
|
|
22
|
+
import { createSemanticSkillRouter } from "../src/decisions/prompt-triage";
|
|
23
23
|
import { detectPrimarySkillKeyword } from "../src/hooks/skill-state";
|
|
24
24
|
import { discoverAuthStorage } from "../src/sdk";
|
|
25
25
|
|
|
@@ -27,7 +27,8 @@ type Expected = "deep-interview" | "ralplan" | "ultragoal" | "team" | null;
|
|
|
27
27
|
interface Case {
|
|
28
28
|
prompt: string;
|
|
29
29
|
expect: Expected;
|
|
30
|
-
|
|
30
|
+
/** BCP 47 primary subtag. The keyword table only knows ko and en; every other row measures the semantic stage alone. */
|
|
31
|
+
lang: string;
|
|
31
32
|
}
|
|
32
33
|
|
|
33
34
|
const CASES: Case[] = [
|
|
@@ -54,6 +55,25 @@ const CASES: Case[] = [
|
|
|
54
55
|
{ prompt: "우리 서비스에 이 모델 붙이면 뭐가 좋아?", expect: null, lang: "ko" },
|
|
55
56
|
{ prompt: "fix the failing lint rule in src/utils.ts", expect: null, lang: "en" },
|
|
56
57
|
{ prompt: "what does this regex do?", expect: null, lang: "en" },
|
|
58
|
+
// Languages the hand-written table has no entries for. Routing here is the
|
|
59
|
+
// semantic stage or nothing, which is what the per-language column shows.
|
|
60
|
+
{ prompt: "需求还不清楚,先通过提问把规格问出来", expect: "deep-interview", lang: "zh" },
|
|
61
|
+
{ prompt: "架构风险很大,先给我一个需要审批的详细计划", expect: "ralplan", lang: "zh" },
|
|
62
|
+
{ prompt: "把这个目标登记下来,持续跟踪直到全部交付验证完", expect: "ultragoal", lang: "zh" },
|
|
63
|
+
{ prompt: "任务太大了,拆成几个并行的工作者一起做", expect: "team", lang: "zh" },
|
|
64
|
+
{ prompt: "这个测试为什么会挂?", expect: null, lang: "zh" },
|
|
65
|
+
{ prompt: "修一下 README 里的错别字", expect: null, lang: "zh" },
|
|
66
|
+
{ prompt: "要件がまだ曖昧なので、質問して仕様を引き出して", expect: "deep-interview", lang: "ja" },
|
|
67
|
+
{ prompt: "実装前に設計案を比較した計画書を作って承認を待って", expect: "ralplan", lang: "ja" },
|
|
68
|
+
{ prompt: "この目標を最後まで追跡して、途中で忘れないで", expect: "ultragoal", lang: "ja" },
|
|
69
|
+
{ prompt: "作業が大きいのでワーカーを複数立てて並列で進めて", expect: "team", lang: "ja" },
|
|
70
|
+
{ prompt: "この関数は何をしているか説明して", expect: null, lang: "ja" },
|
|
71
|
+
{ prompt: "Hazme preguntas hasta que los requisitos estén claros", expect: "deep-interview", lang: "es" },
|
|
72
|
+
{ prompt: "Prepara un plan detallado y espera mi aprobación antes de tocar código", expect: "ralplan", lang: "es" },
|
|
73
|
+
{ prompt: "Arregla el error de lint en src/utils.ts", expect: null, lang: "es" },
|
|
74
|
+
{ prompt: "Составь согласованный план миграции и жди моего одобрения", expect: "ralplan", lang: "ru" },
|
|
75
|
+
{ prompt: "Разбей работу на несколько параллельных воркеров", expect: "team", lang: "ru" },
|
|
76
|
+
{ prompt: "Что делает эта регулярка?", expect: null, lang: "ru" },
|
|
57
77
|
];
|
|
58
78
|
|
|
59
79
|
function pct(hit: number, total: number): string {
|
|
@@ -120,18 +140,23 @@ async function main(): Promise<void> {
|
|
|
120
140
|
};
|
|
121
141
|
const positives = (row: (typeof rows)[number]) => row.case.expect !== null;
|
|
122
142
|
const negatives = (row: (typeof rows)[number]) => row.case.expect === null;
|
|
123
|
-
const
|
|
124
|
-
const en = (row: (typeof rows)[number]) => positives(row) && row.case.lang === "en";
|
|
143
|
+
const langs = [...new Set(CASES.map(testCase => testCase.lang))];
|
|
125
144
|
const hybrid = (row: (typeof rows)[number]) => row.keyword ?? row.semantic;
|
|
126
145
|
|
|
127
146
|
console.log("\n=== stage comparison ===");
|
|
147
|
+
// Both hosts ship the hybrid: the session in `AgentSession#routeWorkflowSemantically`,
|
|
148
|
+
// the Codex hook in `hooks/native-prompt-routing.ts`. The single-stage rows show
|
|
149
|
+
// what each stage contributes on its own.
|
|
128
150
|
for (const [label, pick] of [
|
|
129
|
-
["keyword only
|
|
130
|
-
["semantic only
|
|
131
|
-
["keyword+semantic (
|
|
151
|
+
["keyword only", (row: (typeof rows)[number]) => row.keyword],
|
|
152
|
+
["semantic only", (row: (typeof rows)[number]) => row.semantic],
|
|
153
|
+
["keyword+semantic (SHIPPED)", hybrid],
|
|
132
154
|
] as const) {
|
|
155
|
+
const perLang = langs
|
|
156
|
+
.map(lang => `${lang} ${pct(...score(pick, row => positives(row) && row.case.lang === lang))}`)
|
|
157
|
+
.join(" ");
|
|
133
158
|
console.log(
|
|
134
|
-
`${label.padEnd(32)} all ${pct(...score(pick, () => true))}
|
|
159
|
+
`${label.padEnd(32)} all ${pct(...score(pick, () => true))} ${perLang} clean-negatives ${pct(...score(pick, negatives))}`,
|
|
135
160
|
);
|
|
136
161
|
}
|
|
137
162
|
|
|
@@ -140,13 +165,13 @@ async function main(): Promise<void> {
|
|
|
140
165
|
`\nlatency p50 ${latencies[Math.floor(latencies.length / 2)]}ms p95 ${latencies[Math.max(0, Math.ceil(latencies.length * 0.95) - 1)]}ms max ${latencies.at(-1)}ms`,
|
|
141
166
|
);
|
|
142
167
|
|
|
143
|
-
// Report against what
|
|
144
|
-
const misses = rows.filter(row => row
|
|
168
|
+
// Report against what ships: keyword first, the model for what it missed.
|
|
169
|
+
const misses = rows.filter(row => hybrid(row) !== row.case.expect);
|
|
145
170
|
if (misses.length > 0) {
|
|
146
|
-
console.log("\n=== misses in shipped configuration (semantic
|
|
171
|
+
console.log("\n=== misses in shipped configuration (keyword+semantic) ===");
|
|
147
172
|
for (const row of misses)
|
|
148
173
|
console.log(
|
|
149
|
-
` [${row.case.lang}] want=${row.case.expect ?? "none"} got=${row
|
|
174
|
+
` [${row.case.lang}] want=${row.case.expect ?? "none"} got=${hybrid(row) ?? "none"} :: ${row.case.prompt}`,
|
|
150
175
|
);
|
|
151
176
|
}
|
|
152
177
|
|