kenaz 0.3.1 → 0.3.3
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/core/.agents/skills/typesafe-ai/LICENSE +21 -21
- package/core/.agents/skills/typesafe-ai/SKILL.md +149 -149
- package/core/.claude/commands/Ansuz.md +19 -30
- package/core/.claude/hooks/_ansuz-conversation.js +1 -0
- package/core/.claude/hooks/_ansuz-conversation.test.js +6 -0
- package/core/.claude/hooks/_run-user-prompt.js +0 -4
- package/core/.claude/hooks/_run-user-prompt.test.js +2 -3
- package/core/.claude/hooks/session-init.ps1 +49 -49
- package/core/.claude/settings.json +0 -37
- package/core/.claude/skills/ui-ux-pro-max/data/charts.csv +26 -26
- package/core/.claude/skills/ui-ux-pro-max/data/colors.csv +97 -97
- package/core/.claude/skills/ui-ux-pro-max/data/icons.csv +101 -101
- package/core/.claude/skills/ui-ux-pro-max/data/landing.csv +31 -31
- package/core/.claude/skills/ui-ux-pro-max/data/products.csv +96 -96
- package/core/.claude/skills/ui-ux-pro-max/data/react-performance.csv +45 -45
- package/core/.claude/skills/ui-ux-pro-max/data/stacks/astro.csv +54 -54
- package/core/.claude/skills/ui-ux-pro-max/data/stacks/flutter.csv +53 -53
- package/core/.claude/skills/ui-ux-pro-max/data/stacks/html-tailwind.csv +56 -56
- package/core/.claude/skills/ui-ux-pro-max/data/stacks/jetpack-compose.csv +53 -53
- package/core/.claude/skills/ui-ux-pro-max/data/stacks/nextjs.csv +53 -53
- package/core/.claude/skills/ui-ux-pro-max/data/stacks/nuxt-ui.csv +51 -51
- package/core/.claude/skills/ui-ux-pro-max/data/stacks/nuxtjs.csv +59 -59
- package/core/.claude/skills/ui-ux-pro-max/data/stacks/react-native.csv +52 -52
- package/core/.claude/skills/ui-ux-pro-max/data/stacks/react.csv +54 -54
- package/core/.claude/skills/ui-ux-pro-max/data/stacks/shadcn.csv +61 -61
- package/core/.claude/skills/ui-ux-pro-max/data/stacks/svelte.csv +54 -54
- package/core/.claude/skills/ui-ux-pro-max/data/stacks/swiftui.csv +51 -51
- package/core/.claude/skills/ui-ux-pro-max/data/stacks/vue.csv +50 -50
- package/core/.claude/skills/ui-ux-pro-max/data/styles.csv +68 -68
- package/core/.claude/skills/ui-ux-pro-max/data/typography.csv +57 -57
- package/core/.claude/skills/ui-ux-pro-max/data/ui-reasoning.csv +101 -101
- package/core/.claude/skills/ui-ux-pro-max/data/ux-guidelines.csv +99 -99
- package/core/.claude/skills/ui-ux-pro-max/data/web-interface.csv +31 -31
- package/core/.claude/skills/ui-ux-pro-max/scripts/core.py +253 -253
- package/core/.claude/skills/ui-ux-pro-max/scripts/design_system.py +1067 -1067
- package/core/.claude/skills/ui-ux-pro-max/scripts/search.py +106 -106
- package/core/.collaboration/core/task-api.js +3 -8
- package/core/.formation/bifrost/codex-doctor.js +1 -0
- package/core/.formation/bifrost/setup-codex.js +14 -5
- package/core/.formation/bifrost/setup-codex.test.js +20 -0
- package/core/.formation/tests/hello.txt +2 -2
- package/core/.formation/tools/prompt.js +58 -5
- package/core/.formation/tools/prompt.test.js +29 -0
- package/core/.kenaz/claude-hooks-manifest.json +16 -16
- package/core/.kenaz/knowledge/ladybug_2c51da1e_shadow_file_frame_group.patch +163 -163
- package/core/.kenaz/knowledge/ladybug_30cf85690e_shadow_page_read_during_checkpoint.patch +393 -393
- package/core/.kenaz/knowledge/ladybug_7a17858d3a_node_group_delete_lock.patch +102 -102
- package/core/.kenaz/knowledge/ladybug_926a0ba1_column_checkpoint.patch +412 -412
- package/core/.kenaz/knowledge/ladybug_a346c4c4b6_pk_index_header_pages.patch +144 -144
- package/core/.kenaz/knowledge/ladybug_ba5f38815b_csr_node_group.patch +86 -86
- package/core/.kenaz/muninn_episode_write.ps1 +17 -17
- package/core/.kenaz/specialist/field_report_watermark.json +1 -1
- package/core/.kenaz/specialist/memory_guard_fired.json +1 -1
- package/core/.kenaz/specialist/memory_guard_trigger_cache.json +1 -1
- package/core/.kenaz/specialist/t1777_probe.txt +1 -1
- package/core/.matrix/index.matrix +5 -5
- package/core/.specialist/Ansuz/RULES.md +6 -10
- package/core/.specialist/tools/__tests__/task-2218-boot-session-id.test.js +17 -1
- package/core/.specialist/tools/__tests__/task-2222-session-identity-interleave.test.js +17 -1
- package/core/.specialist/tools/__tests__/task-2254-handoff-archive-contract.test.js +6 -5
- package/core/.specialist/tools/ansuz-boot-fast.js +44 -100
- package/core/.specialist/tools/ansuz-boot-fast.test.js +45 -1
- package/core/KENAZ_CORE_VERSION +1 -1
- package/core/dev/ANSUZ_STARTUP_INJECTION_REVIEW.md +355 -355
- package/core/dev/JEV_INTEGRATION_PLAN.md +189 -189
- package/core/dev/design/matrix_format_EXAMPLE_memory.matrix +13 -13
- package/core/dev/design/matrix_format_EXAMPLE_project_status.matrix +9 -9
- package/core/dev/design/matrix_format_EXAMPLE_tasks.matrix +14 -14
- package/core/dev/design/matrix_v2_EXAMPLE.matrix +67 -67
- package/core/dev/jev-eval/ACTIVE_CORRECTION.md +52 -52
- package/core/dev/jev-eval/CONTEXT_REVIEW.md +31 -31
- package/core/dev/jev-eval/EVALUATION_STATUS.md +39 -39
- package/core/dev/jev-eval/HISTORICAL_CORPUS.md +33 -33
- package/core/dev/jev-eval/README.md +99 -99
- package/core/dev/jev-eval/THALAMUS_RANKING.md +58 -58
- package/core/dev/jev-eval/reports/annotation-pilot-2026-09-20.json +32 -32
- package/core/dev/jev-eval/reports/context-review-2026-09-20.json +48 -48
- package/core/dev/jev-eval/reports/history-corpus-2026-09-20.json +52 -52
- package/core/dev/pipeline/SELF_ENHANCEMENT_STATUS.json.bak +47 -47
- package/core/dev/pipeline/tasks/.gitkeep +1 -1
- package/core/dev/pipeline/test_results/TASK_TEST_PARALLEL_002_result.txt +27 -27
- package/core/dev/plans/ANSUZ_SLIM_PLAN.md +68 -0
- package/core/dev/scripts/init_git.sh +0 -0
- package/core/dev/scripts/init_project.sh +0 -0
- package/core/dev/scripts/legacy/dev_agent_scheduler.ps1.bak +536 -536
- package/core/dev/scripts/legacy/multi_agent_scheduler.ps1.bak +1978 -1978
- package/core/dev/scripts/matrix/src/cli/commands/list-tags.ts.wip +140 -140
- package/core/dev/scripts/matrix/src/cli/commands/manage-tags.ts.wip +173 -173
- package/core/dev/scripts/matrix/src/cli/commands/stats.ts.wip +168 -168
- package/core/dev/scripts/matrix/src/cli/commands/tokens.ts.wip +126 -126
- package/core/dev/scripts/multi_agent_scheduler.sh +0 -0
- package/core/dev/scripts/run_once.sh +0 -0
- package/core/dev/scripts/setup_mcp.sh +0 -0
- package/core/dev/scripts/setup_rag.sh +0 -0
- package/core/dev/scripts/start_dev_agent.sh +0 -0
- package/core/dev/scripts/start_multi_agent.sh +0 -0
- package/core/dev/scripts/sync_task_queue.sh +0 -0
- package/core/k-cli/dist/alaya/run.js +91 -0
- package/core/k-cli/dist/core/output.js +24 -0
- package/core/k-cli/dist/core/process.js +38 -0
- package/core/k-cli/dist/core/project.js +80 -0
- package/core/k-cli/dist/core/spawn.js +268 -0
- package/core/k-cli/dist/core/types.js +2 -0
- package/core/k-cli/dist/dashboard/update.js +53 -0
- package/core/k-cli/dist/dashboard/version.js +20 -0
- package/core/k-cli/dist/dispatch/run.js +241 -0
- package/core/k-cli/dist/dispatch/status.js +103 -0
- package/core/k-cli/dist/index.js +283 -0
- package/core/k-cli/dist/project/initialize.js +301 -0
- package/core/k-cli/dist/project/install.js +89 -0
- package/core/k-cli/dist/project/status.js +138 -0
- package/core/k-cli/dist/specialist/list.js +100 -0
- package/core/k-cli/dist/specialist/run.js +93 -0
- package/core/k-cli/dist/task/create.js +194 -0
- package/core/k-cli/dist/task/delete.js +98 -0
- package/core/k-cli/dist/task/get.js +55 -0
- package/core/k-cli/dist/task/list.js +58 -0
- package/core/k-cli/dist/task/submit.js +156 -0
- package/core/k-cli/dist/ygg/run.js +51 -0
- package/core/k-cli/node_modules/.bin/yaml +16 -0
- package/core/k-cli/node_modules/.bin/yaml.cmd +17 -0
- package/core/k-cli/node_modules/.bin/yaml.ps1 +28 -0
- package/core/k-cli/node_modules/.package-lock.json +23 -0
- package/core/k-cli/node_modules/yaml/LICENSE +13 -0
- package/core/k-cli/node_modules/yaml/README.md +172 -0
- package/core/k-cli/node_modules/yaml/bin.mjs +11 -0
- package/core/k-cli/node_modules/yaml/browser/dist/compose/compose-collection.js +88 -0
- package/core/k-cli/node_modules/yaml/browser/dist/compose/compose-doc.js +43 -0
- package/core/k-cli/node_modules/yaml/browser/dist/compose/compose-node.js +109 -0
- package/core/k-cli/node_modules/yaml/browser/dist/compose/compose-scalar.js +86 -0
- package/core/k-cli/node_modules/yaml/browser/dist/compose/composer.js +217 -0
- package/core/k-cli/node_modules/yaml/browser/dist/compose/resolve-block-map.js +115 -0
- package/core/k-cli/node_modules/yaml/browser/dist/compose/resolve-block-scalar.js +198 -0
- package/core/k-cli/node_modules/yaml/browser/dist/compose/resolve-block-seq.js +49 -0
- package/core/k-cli/node_modules/yaml/browser/dist/compose/resolve-end.js +37 -0
- package/core/k-cli/node_modules/yaml/browser/dist/compose/resolve-flow-collection.js +207 -0
- package/core/k-cli/node_modules/yaml/browser/dist/compose/resolve-flow-scalar.js +223 -0
- package/core/k-cli/node_modules/yaml/browser/dist/compose/resolve-props.js +146 -0
- package/core/k-cli/node_modules/yaml/browser/dist/compose/util-contains-newline.js +34 -0
- package/core/k-cli/node_modules/yaml/browser/dist/compose/util-empty-scalar-position.js +26 -0
- package/core/k-cli/node_modules/yaml/browser/dist/compose/util-flow-indent-check.js +15 -0
- package/core/k-cli/node_modules/yaml/browser/dist/compose/util-map-includes.js +13 -0
- package/core/k-cli/node_modules/yaml/browser/dist/doc/Document.js +335 -0
- package/core/k-cli/node_modules/yaml/browser/dist/doc/anchors.js +71 -0
- package/core/k-cli/node_modules/yaml/browser/dist/doc/applyReviver.js +55 -0
- package/core/k-cli/node_modules/yaml/browser/dist/doc/createNode.js +88 -0
- package/core/k-cli/node_modules/yaml/browser/dist/doc/directives.js +176 -0
- package/core/k-cli/node_modules/yaml/browser/dist/errors.js +57 -0
- package/core/k-cli/node_modules/yaml/browser/dist/index.js +17 -0
- package/core/k-cli/node_modules/yaml/browser/dist/log.js +11 -0
- package/core/k-cli/node_modules/yaml/browser/dist/nodes/Alias.js +114 -0
- package/core/k-cli/node_modules/yaml/browser/dist/nodes/Collection.js +147 -0
- package/core/k-cli/node_modules/yaml/browser/dist/nodes/Node.js +38 -0
- package/core/k-cli/node_modules/yaml/browser/dist/nodes/Pair.js +36 -0
- package/core/k-cli/node_modules/yaml/browser/dist/nodes/Scalar.js +24 -0
- package/core/k-cli/node_modules/yaml/browser/dist/nodes/YAMLMap.js +144 -0
- package/core/k-cli/node_modules/yaml/browser/dist/nodes/YAMLSeq.js +113 -0
- package/core/k-cli/node_modules/yaml/browser/dist/nodes/addPairToJSMap.js +63 -0
- package/core/k-cli/node_modules/yaml/browser/dist/nodes/identity.js +36 -0
- package/core/k-cli/node_modules/yaml/browser/dist/nodes/toJS.js +37 -0
- package/core/k-cli/node_modules/yaml/browser/dist/parse/cst-scalar.js +214 -0
- package/core/k-cli/node_modules/yaml/browser/dist/parse/cst-stringify.js +61 -0
- package/core/k-cli/node_modules/yaml/browser/dist/parse/cst-visit.js +97 -0
- package/core/k-cli/node_modules/yaml/browser/dist/parse/cst.js +98 -0
- package/core/k-cli/node_modules/yaml/browser/dist/parse/lexer.js +717 -0
- package/core/k-cli/node_modules/yaml/browser/dist/parse/line-counter.js +39 -0
- package/core/k-cli/node_modules/yaml/browser/dist/parse/parser.js +967 -0
- package/core/k-cli/node_modules/yaml/browser/dist/public-api.js +102 -0
- package/core/k-cli/node_modules/yaml/browser/dist/schema/Schema.js +37 -0
- package/core/k-cli/node_modules/yaml/browser/dist/schema/common/map.js +17 -0
- package/core/k-cli/node_modules/yaml/browser/dist/schema/common/null.js +15 -0
- package/core/k-cli/node_modules/yaml/browser/dist/schema/common/seq.js +17 -0
- package/core/k-cli/node_modules/yaml/browser/dist/schema/common/string.js +14 -0
- package/core/k-cli/node_modules/yaml/browser/dist/schema/core/bool.js +19 -0
- package/core/k-cli/node_modules/yaml/browser/dist/schema/core/float.js +43 -0
- package/core/k-cli/node_modules/yaml/browser/dist/schema/core/int.js +38 -0
- package/core/k-cli/node_modules/yaml/browser/dist/schema/core/schema.js +23 -0
- package/core/k-cli/node_modules/yaml/browser/dist/schema/json/schema.js +62 -0
- package/core/k-cli/node_modules/yaml/browser/dist/schema/tags.js +96 -0
- package/core/k-cli/node_modules/yaml/browser/dist/schema/yaml-1.1/binary.js +58 -0
- package/core/k-cli/node_modules/yaml/browser/dist/schema/yaml-1.1/bool.js +26 -0
- package/core/k-cli/node_modules/yaml/browser/dist/schema/yaml-1.1/float.js +46 -0
- package/core/k-cli/node_modules/yaml/browser/dist/schema/yaml-1.1/int.js +71 -0
- package/core/k-cli/node_modules/yaml/browser/dist/schema/yaml-1.1/merge.js +64 -0
- package/core/k-cli/node_modules/yaml/browser/dist/schema/yaml-1.1/omap.js +74 -0
- package/core/k-cli/node_modules/yaml/browser/dist/schema/yaml-1.1/pairs.js +78 -0
- package/core/k-cli/node_modules/yaml/browser/dist/schema/yaml-1.1/schema.js +39 -0
- package/core/k-cli/node_modules/yaml/browser/dist/schema/yaml-1.1/set.js +93 -0
- package/core/k-cli/node_modules/yaml/browser/dist/schema/yaml-1.1/timestamp.js +101 -0
- package/core/k-cli/node_modules/yaml/browser/dist/stringify/foldFlowLines.js +146 -0
- package/core/k-cli/node_modules/yaml/browser/dist/stringify/stringify.js +129 -0
- package/core/k-cli/node_modules/yaml/browser/dist/stringify/stringifyCollection.js +153 -0
- package/core/k-cli/node_modules/yaml/browser/dist/stringify/stringifyComment.js +20 -0
- package/core/k-cli/node_modules/yaml/browser/dist/stringify/stringifyDocument.js +85 -0
- package/core/k-cli/node_modules/yaml/browser/dist/stringify/stringifyNumber.js +24 -0
- package/core/k-cli/node_modules/yaml/browser/dist/stringify/stringifyPair.js +150 -0
- package/core/k-cli/node_modules/yaml/browser/dist/stringify/stringifyString.js +336 -0
- package/core/k-cli/node_modules/yaml/browser/dist/util.js +11 -0
- package/core/k-cli/node_modules/yaml/browser/dist/visit.js +233 -0
- package/core/k-cli/node_modules/yaml/browser/index.js +5 -0
- package/core/k-cli/node_modules/yaml/browser/package.json +3 -0
- package/core/k-cli/node_modules/yaml/dist/cli.d.ts +8 -0
- package/core/k-cli/node_modules/yaml/dist/cli.mjs +201 -0
- package/core/k-cli/node_modules/yaml/dist/compose/compose-collection.d.ts +11 -0
- package/core/k-cli/node_modules/yaml/dist/compose/compose-collection.js +90 -0
- package/core/k-cli/node_modules/yaml/dist/compose/compose-doc.d.ts +7 -0
- package/core/k-cli/node_modules/yaml/dist/compose/compose-doc.js +45 -0
- package/core/k-cli/node_modules/yaml/dist/compose/compose-node.d.ts +29 -0
- package/core/k-cli/node_modules/yaml/dist/compose/compose-node.js +112 -0
- package/core/k-cli/node_modules/yaml/dist/compose/compose-scalar.d.ts +5 -0
- package/core/k-cli/node_modules/yaml/dist/compose/compose-scalar.js +88 -0
- package/core/k-cli/node_modules/yaml/dist/compose/composer.d.ts +63 -0
- package/core/k-cli/node_modules/yaml/dist/compose/composer.js +222 -0
- package/core/k-cli/node_modules/yaml/dist/compose/resolve-block-map.d.ts +6 -0
- package/core/k-cli/node_modules/yaml/dist/compose/resolve-block-map.js +117 -0
- package/core/k-cli/node_modules/yaml/dist/compose/resolve-block-scalar.d.ts +11 -0
- package/core/k-cli/node_modules/yaml/dist/compose/resolve-block-scalar.js +200 -0
- package/core/k-cli/node_modules/yaml/dist/compose/resolve-block-seq.d.ts +6 -0
- package/core/k-cli/node_modules/yaml/dist/compose/resolve-block-seq.js +51 -0
- package/core/k-cli/node_modules/yaml/dist/compose/resolve-end.d.ts +6 -0
- package/core/k-cli/node_modules/yaml/dist/compose/resolve-end.js +39 -0
- package/core/k-cli/node_modules/yaml/dist/compose/resolve-flow-collection.d.ts +7 -0
- package/core/k-cli/node_modules/yaml/dist/compose/resolve-flow-collection.js +209 -0
- package/core/k-cli/node_modules/yaml/dist/compose/resolve-flow-scalar.d.ts +10 -0
- package/core/k-cli/node_modules/yaml/dist/compose/resolve-flow-scalar.js +225 -0
- package/core/k-cli/node_modules/yaml/dist/compose/resolve-props.d.ts +23 -0
- package/core/k-cli/node_modules/yaml/dist/compose/resolve-props.js +148 -0
- package/core/k-cli/node_modules/yaml/dist/compose/util-contains-newline.d.ts +2 -0
- package/core/k-cli/node_modules/yaml/dist/compose/util-contains-newline.js +36 -0
- package/core/k-cli/node_modules/yaml/dist/compose/util-empty-scalar-position.d.ts +2 -0
- package/core/k-cli/node_modules/yaml/dist/compose/util-empty-scalar-position.js +28 -0
- package/core/k-cli/node_modules/yaml/dist/compose/util-flow-indent-check.d.ts +3 -0
- package/core/k-cli/node_modules/yaml/dist/compose/util-flow-indent-check.js +17 -0
- package/core/k-cli/node_modules/yaml/dist/compose/util-map-includes.d.ts +4 -0
- package/core/k-cli/node_modules/yaml/dist/compose/util-map-includes.js +15 -0
- package/core/k-cli/node_modules/yaml/dist/doc/Document.d.ts +141 -0
- package/core/k-cli/node_modules/yaml/dist/doc/Document.js +337 -0
- package/core/k-cli/node_modules/yaml/dist/doc/anchors.d.ts +24 -0
- package/core/k-cli/node_modules/yaml/dist/doc/anchors.js +76 -0
- package/core/k-cli/node_modules/yaml/dist/doc/applyReviver.d.ts +9 -0
- package/core/k-cli/node_modules/yaml/dist/doc/applyReviver.js +57 -0
- package/core/k-cli/node_modules/yaml/dist/doc/createNode.d.ts +17 -0
- package/core/k-cli/node_modules/yaml/dist/doc/createNode.js +90 -0
- package/core/k-cli/node_modules/yaml/dist/doc/directives.d.ts +49 -0
- package/core/k-cli/node_modules/yaml/dist/doc/directives.js +178 -0
- package/core/k-cli/node_modules/yaml/dist/errors.d.ts +21 -0
- package/core/k-cli/node_modules/yaml/dist/errors.js +62 -0
- package/core/k-cli/node_modules/yaml/dist/index.d.ts +25 -0
- package/core/k-cli/node_modules/yaml/dist/index.js +50 -0
- package/core/k-cli/node_modules/yaml/dist/log.d.ts +3 -0
- package/core/k-cli/node_modules/yaml/dist/log.js +19 -0
- package/core/k-cli/node_modules/yaml/dist/nodes/Alias.d.ts +29 -0
- package/core/k-cli/node_modules/yaml/dist/nodes/Alias.js +116 -0
- package/core/k-cli/node_modules/yaml/dist/nodes/Collection.d.ts +73 -0
- package/core/k-cli/node_modules/yaml/dist/nodes/Collection.js +151 -0
- package/core/k-cli/node_modules/yaml/dist/nodes/Node.d.ts +53 -0
- package/core/k-cli/node_modules/yaml/dist/nodes/Node.js +40 -0
- package/core/k-cli/node_modules/yaml/dist/nodes/Pair.d.ts +22 -0
- package/core/k-cli/node_modules/yaml/dist/nodes/Pair.js +39 -0
- package/core/k-cli/node_modules/yaml/dist/nodes/Scalar.d.ts +43 -0
- package/core/k-cli/node_modules/yaml/dist/nodes/Scalar.js +27 -0
- package/core/k-cli/node_modules/yaml/dist/nodes/YAMLMap.d.ts +53 -0
- package/core/k-cli/node_modules/yaml/dist/nodes/YAMLMap.js +147 -0
- package/core/k-cli/node_modules/yaml/dist/nodes/YAMLSeq.d.ts +60 -0
- package/core/k-cli/node_modules/yaml/dist/nodes/YAMLSeq.js +115 -0
- package/core/k-cli/node_modules/yaml/dist/nodes/addPairToJSMap.d.ts +4 -0
- package/core/k-cli/node_modules/yaml/dist/nodes/addPairToJSMap.js +65 -0
- package/core/k-cli/node_modules/yaml/dist/nodes/identity.d.ts +23 -0
- package/core/k-cli/node_modules/yaml/dist/nodes/identity.js +53 -0
- package/core/k-cli/node_modules/yaml/dist/nodes/toJS.d.ts +29 -0
- package/core/k-cli/node_modules/yaml/dist/nodes/toJS.js +39 -0
- package/core/k-cli/node_modules/yaml/dist/options.d.ts +350 -0
- package/core/k-cli/node_modules/yaml/dist/parse/cst-scalar.d.ts +64 -0
- package/core/k-cli/node_modules/yaml/dist/parse/cst-scalar.js +218 -0
- package/core/k-cli/node_modules/yaml/dist/parse/cst-stringify.d.ts +8 -0
- package/core/k-cli/node_modules/yaml/dist/parse/cst-stringify.js +63 -0
- package/core/k-cli/node_modules/yaml/dist/parse/cst-visit.d.ts +39 -0
- package/core/k-cli/node_modules/yaml/dist/parse/cst-visit.js +99 -0
- package/core/k-cli/node_modules/yaml/dist/parse/cst.d.ts +109 -0
- package/core/k-cli/node_modules/yaml/dist/parse/cst.js +112 -0
- package/core/k-cli/node_modules/yaml/dist/parse/lexer.d.ts +87 -0
- package/core/k-cli/node_modules/yaml/dist/parse/lexer.js +719 -0
- package/core/k-cli/node_modules/yaml/dist/parse/line-counter.d.ts +22 -0
- package/core/k-cli/node_modules/yaml/dist/parse/line-counter.js +41 -0
- package/core/k-cli/node_modules/yaml/dist/parse/parser.d.ts +84 -0
- package/core/k-cli/node_modules/yaml/dist/parse/parser.js +972 -0
- package/core/k-cli/node_modules/yaml/dist/public-api.d.ts +44 -0
- package/core/k-cli/node_modules/yaml/dist/public-api.js +107 -0
- package/core/k-cli/node_modules/yaml/dist/schema/Schema.d.ts +17 -0
- package/core/k-cli/node_modules/yaml/dist/schema/Schema.js +39 -0
- package/core/k-cli/node_modules/yaml/dist/schema/common/map.d.ts +2 -0
- package/core/k-cli/node_modules/yaml/dist/schema/common/map.js +19 -0
- package/core/k-cli/node_modules/yaml/dist/schema/common/null.d.ts +4 -0
- package/core/k-cli/node_modules/yaml/dist/schema/common/null.js +17 -0
- package/core/k-cli/node_modules/yaml/dist/schema/common/seq.d.ts +2 -0
- package/core/k-cli/node_modules/yaml/dist/schema/common/seq.js +19 -0
- package/core/k-cli/node_modules/yaml/dist/schema/common/string.d.ts +2 -0
- package/core/k-cli/node_modules/yaml/dist/schema/common/string.js +16 -0
- package/core/k-cli/node_modules/yaml/dist/schema/core/bool.d.ts +4 -0
- package/core/k-cli/node_modules/yaml/dist/schema/core/bool.js +21 -0
- package/core/k-cli/node_modules/yaml/dist/schema/core/float.d.ts +4 -0
- package/core/k-cli/node_modules/yaml/dist/schema/core/float.js +47 -0
- package/core/k-cli/node_modules/yaml/dist/schema/core/int.d.ts +4 -0
- package/core/k-cli/node_modules/yaml/dist/schema/core/int.js +42 -0
- package/core/k-cli/node_modules/yaml/dist/schema/core/schema.d.ts +1 -0
- package/core/k-cli/node_modules/yaml/dist/schema/core/schema.js +25 -0
- package/core/k-cli/node_modules/yaml/dist/schema/json/schema.d.ts +2 -0
- package/core/k-cli/node_modules/yaml/dist/schema/json/schema.js +64 -0
- package/core/k-cli/node_modules/yaml/dist/schema/json-schema.d.ts +69 -0
- package/core/k-cli/node_modules/yaml/dist/schema/tags.d.ts +48 -0
- package/core/k-cli/node_modules/yaml/dist/schema/tags.js +99 -0
- package/core/k-cli/node_modules/yaml/dist/schema/types.d.ts +92 -0
- package/core/k-cli/node_modules/yaml/dist/schema/yaml-1.1/binary.d.ts +2 -0
- package/core/k-cli/node_modules/yaml/dist/schema/yaml-1.1/binary.js +70 -0
- package/core/k-cli/node_modules/yaml/dist/schema/yaml-1.1/bool.d.ts +7 -0
- package/core/k-cli/node_modules/yaml/dist/schema/yaml-1.1/bool.js +29 -0
- package/core/k-cli/node_modules/yaml/dist/schema/yaml-1.1/float.d.ts +4 -0
- package/core/k-cli/node_modules/yaml/dist/schema/yaml-1.1/float.js +50 -0
- package/core/k-cli/node_modules/yaml/dist/schema/yaml-1.1/int.d.ts +5 -0
- package/core/k-cli/node_modules/yaml/dist/schema/yaml-1.1/int.js +76 -0
- package/core/k-cli/node_modules/yaml/dist/schema/yaml-1.1/merge.d.ts +9 -0
- package/core/k-cli/node_modules/yaml/dist/schema/yaml-1.1/merge.js +68 -0
- package/core/k-cli/node_modules/yaml/dist/schema/yaml-1.1/omap.d.ts +22 -0
- package/core/k-cli/node_modules/yaml/dist/schema/yaml-1.1/omap.js +77 -0
- package/core/k-cli/node_modules/yaml/dist/schema/yaml-1.1/pairs.d.ts +10 -0
- package/core/k-cli/node_modules/yaml/dist/schema/yaml-1.1/pairs.js +82 -0
- package/core/k-cli/node_modules/yaml/dist/schema/yaml-1.1/schema.d.ts +1 -0
- package/core/k-cli/node_modules/yaml/dist/schema/yaml-1.1/schema.js +41 -0
- package/core/k-cli/node_modules/yaml/dist/schema/yaml-1.1/set.d.ts +28 -0
- package/core/k-cli/node_modules/yaml/dist/schema/yaml-1.1/set.js +96 -0
- package/core/k-cli/node_modules/yaml/dist/schema/yaml-1.1/timestamp.d.ts +6 -0
- package/core/k-cli/node_modules/yaml/dist/schema/yaml-1.1/timestamp.js +105 -0
- package/core/k-cli/node_modules/yaml/dist/stringify/foldFlowLines.d.ts +34 -0
- package/core/k-cli/node_modules/yaml/dist/stringify/foldFlowLines.js +151 -0
- package/core/k-cli/node_modules/yaml/dist/stringify/stringify.d.ts +21 -0
- package/core/k-cli/node_modules/yaml/dist/stringify/stringify.js +132 -0
- package/core/k-cli/node_modules/yaml/dist/stringify/stringifyCollection.d.ts +17 -0
- package/core/k-cli/node_modules/yaml/dist/stringify/stringifyCollection.js +155 -0
- package/core/k-cli/node_modules/yaml/dist/stringify/stringifyComment.d.ts +10 -0
- package/core/k-cli/node_modules/yaml/dist/stringify/stringifyComment.js +24 -0
- package/core/k-cli/node_modules/yaml/dist/stringify/stringifyDocument.d.ts +4 -0
- package/core/k-cli/node_modules/yaml/dist/stringify/stringifyDocument.js +87 -0
- package/core/k-cli/node_modules/yaml/dist/stringify/stringifyNumber.d.ts +2 -0
- package/core/k-cli/node_modules/yaml/dist/stringify/stringifyNumber.js +26 -0
- package/core/k-cli/node_modules/yaml/dist/stringify/stringifyPair.d.ts +3 -0
- package/core/k-cli/node_modules/yaml/dist/stringify/stringifyPair.js +152 -0
- package/core/k-cli/node_modules/yaml/dist/stringify/stringifyString.d.ts +9 -0
- package/core/k-cli/node_modules/yaml/dist/stringify/stringifyString.js +338 -0
- package/core/k-cli/node_modules/yaml/dist/test-events.d.ts +4 -0
- package/core/k-cli/node_modules/yaml/dist/test-events.js +134 -0
- package/core/k-cli/node_modules/yaml/dist/util.d.ts +16 -0
- package/core/k-cli/node_modules/yaml/dist/util.js +28 -0
- package/core/k-cli/node_modules/yaml/dist/visit.d.ts +102 -0
- package/core/k-cli/node_modules/yaml/dist/visit.js +236 -0
- package/core/k-cli/node_modules/yaml/package.json +97 -0
- package/core/k-cli/node_modules/yaml/util.js +2 -0
- package/core/plugins/kenaz/commands/Ansuz.md +19 -30
- package/core/plugins/kenaz/commands/kenaz.md +1 -1
- package/core/plugins/kenaz/hooks/hooks.json +1 -5
- package/core/plugins/kenaz/hooks/kenaz-hook.js +0 -4
- package/core/plugins/kenaz/scripts/kenaz-mode.js +10 -4
- package/core/plugins/kenaz/scripts/kenaz.test.js +1 -1
- package/core/plugins/kenaz/skills/ui-ux-pro-max/data/charts.csv +26 -26
- package/core/plugins/kenaz/skills/ui-ux-pro-max/data/colors.csv +97 -97
- package/core/plugins/kenaz/skills/ui-ux-pro-max/data/icons.csv +101 -101
- package/core/plugins/kenaz/skills/ui-ux-pro-max/data/landing.csv +31 -31
- package/core/plugins/kenaz/skills/ui-ux-pro-max/data/products.csv +96 -96
- package/core/plugins/kenaz/skills/ui-ux-pro-max/data/react-performance.csv +45 -45
- package/core/plugins/kenaz/skills/ui-ux-pro-max/data/stacks/astro.csv +54 -54
- package/core/plugins/kenaz/skills/ui-ux-pro-max/data/stacks/flutter.csv +53 -53
- package/core/plugins/kenaz/skills/ui-ux-pro-max/data/stacks/html-tailwind.csv +56 -56
- package/core/plugins/kenaz/skills/ui-ux-pro-max/data/stacks/jetpack-compose.csv +53 -53
- package/core/plugins/kenaz/skills/ui-ux-pro-max/data/stacks/nextjs.csv +53 -53
- package/core/plugins/kenaz/skills/ui-ux-pro-max/data/stacks/nuxt-ui.csv +51 -51
- package/core/plugins/kenaz/skills/ui-ux-pro-max/data/stacks/nuxtjs.csv +59 -59
- package/core/plugins/kenaz/skills/ui-ux-pro-max/data/stacks/react-native.csv +52 -52
- package/core/plugins/kenaz/skills/ui-ux-pro-max/data/stacks/react.csv +54 -54
- package/core/plugins/kenaz/skills/ui-ux-pro-max/data/stacks/shadcn.csv +61 -61
- package/core/plugins/kenaz/skills/ui-ux-pro-max/data/stacks/svelte.csv +54 -54
- package/core/plugins/kenaz/skills/ui-ux-pro-max/data/stacks/swiftui.csv +51 -51
- package/core/plugins/kenaz/skills/ui-ux-pro-max/data/stacks/vue.csv +50 -50
- package/core/plugins/kenaz/skills/ui-ux-pro-max/data/styles.csv +68 -68
- package/core/plugins/kenaz/skills/ui-ux-pro-max/data/typography.csv +57 -57
- package/core/plugins/kenaz/skills/ui-ux-pro-max/data/ui-reasoning.csv +101 -101
- package/core/plugins/kenaz/skills/ui-ux-pro-max/data/ux-guidelines.csv +99 -99
- package/core/plugins/kenaz/skills/ui-ux-pro-max/data/web-interface.csv +31 -31
- package/core/plugins/kenaz/skills/ui-ux-pro-max/scripts/core.py +253 -253
- package/core/plugins/kenaz/skills/ui-ux-pro-max/scripts/design_system.py +1067 -1067
- package/core/plugins/kenaz/skills/ui-ux-pro-max/scripts/search.py +106 -106
- package/lib/cli.js +50 -3
- package/lib/cli.test.js +35 -0
- package/lib/dashboard.js +10 -0
- package/package.json +6 -6
- package/core/.specialist/tools/__tests__/task-1863-boot-lease-reuse.test.js +0 -99
- package/core/.specialist/tools/__tests__/task-1977-banner-cap.test.js +0 -129
|
@@ -1,67 +1,67 @@
|
|
|
1
|
-
# .matrix v2.0 - Delta Stream Format Example
|
|
2
|
-
# KenazAI Project Status (Feb 6-7, 2026)
|
|
3
|
-
|
|
4
|
-
@meta
|
|
5
|
-
version: 2.0
|
|
6
|
-
type: project_status
|
|
7
|
-
project: KenazAI
|
|
8
|
-
created: 2026-02-06T20:00:00Z
|
|
9
|
-
last_updated: 2026-02-07T08:00:00Z
|
|
10
|
-
|
|
11
|
-
@baseline 2026-02-06T20:00:00Z
|
|
12
|
-
total_tasks: 100
|
|
13
|
-
completed: 85
|
|
14
|
-
in_progress: 10
|
|
15
|
-
blocked: 5
|
|
16
|
-
team_members: 8
|
|
17
|
-
active_agents: [dev_agent, frontend_agent, backend_agent, pm_agent]
|
|
18
|
-
avg_velocity: 12_tasks_per_day
|
|
19
|
-
tech_stack: [Python, TypeScript, React, FastAPI, Supabase]
|
|
20
|
-
|
|
21
|
-
@delta 2026-02-06T21:00:15Z author:dev_agent
|
|
22
|
-
+task T121 completed commits:15 files:8 duration:2h15m
|
|
23
|
-
title:"Fix authentication session timeout bug"
|
|
24
|
-
tags:[auth, security, P0, bug_fix]
|
|
25
|
-
pr_url:github.com/kenazai/kenaz/pull/121
|
|
26
|
-
tests_added:12 coverage:95%
|
|
27
|
-
+task T122 completed commits:3 files:2 duration:45m
|
|
28
|
-
title:"Add metrics dashboard panel"
|
|
29
|
-
tags:[ui, dashboard, P1, feature]
|
|
30
|
-
pr_url:github.com/kenazai/kenaz/pull/122
|
|
31
|
-
+blocked T130 reason:"waiting_for_api_design_approval" blocker:backend_team
|
|
32
|
-
title:"Implement OAuth 2.0 support"
|
|
33
|
-
+metric total_completed:87 in_progress:9 blocked:6 velocity:14_tasks_per_day
|
|
34
|
-
|
|
35
|
-
@delta 2026-02-06T23:30:00Z author:backend_lead
|
|
36
|
-
+comment T130 message:"API design approved, ready to implement"
|
|
37
|
-
-blocked T130 resolution:"API design approved, OAuth specs finalized"
|
|
38
|
-
|
|
39
|
-
@delta 2026-02-07T05:30:42Z author:full_stack_dev
|
|
40
|
-
+task T130 in_progress assigned:full_stack_dev branch:feature/oauth-integration
|
|
41
|
-
started:2026-02-07T05:30:42Z
|
|
42
|
-
+metric active_branches:3 in_progress:10 blocked:5
|
|
43
|
-
|
|
44
|
-
@delta 2026-02-07T08:00:00Z author:full_stack_dev
|
|
45
|
-
+task T130 completed commits:12 files:15 duration:2h30m
|
|
46
|
-
title:"Implement OAuth 2.0 support"
|
|
47
|
-
tags:[auth, oauth, P0, feature]
|
|
48
|
-
pr_url:github.com/kenazai/kenaz/pull/130
|
|
49
|
-
tests_added:25 coverage:92%
|
|
50
|
-
breaking_changes:false
|
|
51
|
-
migration_required:false
|
|
52
|
-
+created T135 title:"Implement matrix v2.0 format" priority:P0
|
|
53
|
-
assigned:developer_team estimated:4_weeks tags:[infrastructure, format]
|
|
54
|
-
description:"Design and implement v2.0 delta stream format with 20-50:1 token efficiency"
|
|
55
|
-
+metric total_completed:88 in_progress:10 blocked:5 velocity:15_tasks_per_day
|
|
56
|
-
total_commits:30 total_files_changed:25 total_tests:37
|
|
57
|
-
|
|
58
|
-
# Query Examples:
|
|
59
|
-
# 1. Current state: parse all deltas from baseline
|
|
60
|
-
# 2. State at 2026-02-06 21:00: baseline + first delta
|
|
61
|
-
# 3. Task T130 history: filter deltas mentioning T130
|
|
62
|
-
# 4. Velocity trend: extract velocity metrics from deltas
|
|
63
|
-
|
|
64
|
-
# Token Analysis:
|
|
65
|
-
# This file: ~125 tokens (estimated)
|
|
66
|
-
# Equivalent JSON: ~2,500 tokens
|
|
67
|
-
# Compression ratio: 20:1 ✓
|
|
1
|
+
# .matrix v2.0 - Delta Stream Format Example
|
|
2
|
+
# KenazAI Project Status (Feb 6-7, 2026)
|
|
3
|
+
|
|
4
|
+
@meta
|
|
5
|
+
version: 2.0
|
|
6
|
+
type: project_status
|
|
7
|
+
project: KenazAI
|
|
8
|
+
created: 2026-02-06T20:00:00Z
|
|
9
|
+
last_updated: 2026-02-07T08:00:00Z
|
|
10
|
+
|
|
11
|
+
@baseline 2026-02-06T20:00:00Z
|
|
12
|
+
total_tasks: 100
|
|
13
|
+
completed: 85
|
|
14
|
+
in_progress: 10
|
|
15
|
+
blocked: 5
|
|
16
|
+
team_members: 8
|
|
17
|
+
active_agents: [dev_agent, frontend_agent, backend_agent, pm_agent]
|
|
18
|
+
avg_velocity: 12_tasks_per_day
|
|
19
|
+
tech_stack: [Python, TypeScript, React, FastAPI, Supabase]
|
|
20
|
+
|
|
21
|
+
@delta 2026-02-06T21:00:15Z author:dev_agent
|
|
22
|
+
+task T121 completed commits:15 files:8 duration:2h15m
|
|
23
|
+
title:"Fix authentication session timeout bug"
|
|
24
|
+
tags:[auth, security, P0, bug_fix]
|
|
25
|
+
pr_url:github.com/kenazai/kenaz/pull/121
|
|
26
|
+
tests_added:12 coverage:95%
|
|
27
|
+
+task T122 completed commits:3 files:2 duration:45m
|
|
28
|
+
title:"Add metrics dashboard panel"
|
|
29
|
+
tags:[ui, dashboard, P1, feature]
|
|
30
|
+
pr_url:github.com/kenazai/kenaz/pull/122
|
|
31
|
+
+blocked T130 reason:"waiting_for_api_design_approval" blocker:backend_team
|
|
32
|
+
title:"Implement OAuth 2.0 support"
|
|
33
|
+
+metric total_completed:87 in_progress:9 blocked:6 velocity:14_tasks_per_day
|
|
34
|
+
|
|
35
|
+
@delta 2026-02-06T23:30:00Z author:backend_lead
|
|
36
|
+
+comment T130 message:"API design approved, ready to implement"
|
|
37
|
+
-blocked T130 resolution:"API design approved, OAuth specs finalized"
|
|
38
|
+
|
|
39
|
+
@delta 2026-02-07T05:30:42Z author:full_stack_dev
|
|
40
|
+
+task T130 in_progress assigned:full_stack_dev branch:feature/oauth-integration
|
|
41
|
+
started:2026-02-07T05:30:42Z
|
|
42
|
+
+metric active_branches:3 in_progress:10 blocked:5
|
|
43
|
+
|
|
44
|
+
@delta 2026-02-07T08:00:00Z author:full_stack_dev
|
|
45
|
+
+task T130 completed commits:12 files:15 duration:2h30m
|
|
46
|
+
title:"Implement OAuth 2.0 support"
|
|
47
|
+
tags:[auth, oauth, P0, feature]
|
|
48
|
+
pr_url:github.com/kenazai/kenaz/pull/130
|
|
49
|
+
tests_added:25 coverage:92%
|
|
50
|
+
breaking_changes:false
|
|
51
|
+
migration_required:false
|
|
52
|
+
+created T135 title:"Implement matrix v2.0 format" priority:P0
|
|
53
|
+
assigned:developer_team estimated:4_weeks tags:[infrastructure, format]
|
|
54
|
+
description:"Design and implement v2.0 delta stream format with 20-50:1 token efficiency"
|
|
55
|
+
+metric total_completed:88 in_progress:10 blocked:5 velocity:15_tasks_per_day
|
|
56
|
+
total_commits:30 total_files_changed:25 total_tests:37
|
|
57
|
+
|
|
58
|
+
# Query Examples:
|
|
59
|
+
# 1. Current state: parse all deltas from baseline
|
|
60
|
+
# 2. State at 2026-02-06 21:00: baseline + first delta
|
|
61
|
+
# 3. Task T130 history: filter deltas mentioning T130
|
|
62
|
+
# 4. Velocity trend: extract velocity metrics from deltas
|
|
63
|
+
|
|
64
|
+
# Token Analysis:
|
|
65
|
+
# This file: ~125 tokens (estimated)
|
|
66
|
+
# Equivalent JSON: ~2,500 tokens
|
|
67
|
+
# Compression ratio: 20:1 ✓
|
|
@@ -1,52 +1,52 @@
|
|
|
1
|
-
# Active Jev correction experiment
|
|
2
|
-
|
|
3
|
-
Local implementation and verification: 2026-09-20. This is a user-authorized opt-in experiment, not a claim of validated model quality or an installer release.
|
|
4
|
-
|
|
5
|
-
## Enable and disable
|
|
6
|
-
|
|
7
|
-
In **Default Models**, use the single **Jev-assisted judgments** switch for the current project (TASK_1970). It controls correction detection and Thalamus context ranking together, defaults to OFF, and saves both `jevCorrectionEnabled` and `jevThalamusEnabled` in one atomic settings request. The runtime keeps these project-local fields for compatibility; it does not inherit them from the Kenaz core project. Partial legacy settings are displayed as mixed, and the next check enables both.
|
|
8
|
-
|
|
9
|
-
Provide your own `TYPESAFE_API_KEY` in the environment used to launch Kenaz, then restart Kenaz and its workers after changing that environment. The UI reports key presence, not a successful authentication check. There is no key input or key storage in these settings. On Windows, TASK_1971 startup reads the saved user environment key from HKCU only when the process variable is absent, so an older parent terminal does not prevent loading it. An explicitly present process value always wins; existing workers still need relaunching to inherit changes. Users without a key retain the existing workflow, even if the enabled preference was saved.
|
|
10
|
-
|
|
11
|
-
ON makes valid Jev results authoritative for correction classification on the selected messages: positives add correction and make it the primary signal; negatives remove correction. Other label values, their relative order, and existing translation text are preserved. Existing consumers that process only the primary signal consequently follow correction instead of another primary label on positive results. Jev is not a text-generation model and does not replace the Ansuz or dispatch model selection.
|
|
12
|
-
|
|
13
|
-
OFF restores the existing workflow for **future** judgments. It does not erase or undo previously accumulated user-profile signals, expertise changes, or feedback. In-flight results are discarded if the flag is OFF when they return.
|
|
14
|
-
|
|
15
|
-
The shared extraction function is called by ordinary Muninn observations, the incremental Muninn watcher, and manual `backfill-user-model.js` runs. Consequently, running a manual historical backfill with this project flag and a valid key can also send eligible historical messages to Jev. No such backfill was performed for this implementation. The flag is project-local, while the existing profile writer stores results in the cross-project `~/.kenaz/user_profile.json`. Jev correction classification is not currently connected to Thalamus stimulation or context ranking.
|
|
16
|
-
|
|
17
|
-
## Execution and fallback
|
|
18
|
-
|
|
19
|
-
The existing synchronous Muninn/Haiku translation and classification still runs first. Jev then overlays correction decisions using fixed model `jev-1.13.0` and experimental probability threshold **0.5**, which is not calibrated on human-reviewed historical cases. This integration does not establish a whole-workflow speedup.
|
|
20
|
-
|
|
21
|
-
Each batch selects at most eight eligible recent user messages from a bounded scan. Each request contains only its target message and up to four preceding visible conversation entries; no later message is shared with earlier targets. Known credential patterns and generated instruction blocks are filtered before transmission. The combined worker input is capped at 32 KiB. Opting in sends this limited conversation context to TypeSafe.
|
|
22
|
-
|
|
23
|
-
A hidden Node worker sends isolated requests concurrently, with a shared 800 ms network deadline including response bodies and a 1,500 ms outer process timeout. There are no retries or configurable endpoint redirects. Any selected request failure or invalid response retains the complete original classification. Missing configuration, OFF, or a missing/invalid key makes no Jev request. If Haiku fails but Jev succeeds, positive correction-only signals can be emitted without inventing a translation.
|
|
24
|
-
|
|
25
|
-
## Local feedback
|
|
26
|
-
|
|
27
|
-
- `<project>/.kenaz/jev/decisions.jsonl`: best-effort, size-limited decision metadata, including hashed message IDs, probabilities, original correction flags, changed flags, latency, and fallback reasons. It excludes message text, keys, headers, and provider error bodies.
|
|
28
|
-
- `<project>/.kenaz/jev/feedback.jsonl`: up to 200 **Helpful / Incorrect judgment** feedback events, with timestamp, model, enabled preference, verdict, and feature scope. The unified control records `feature: all`; historical correction/Thalamus classifications are preserved, and missing historical feature means correction. These describe overall experience, not adjudicated labels for individual cases.
|
|
29
|
-
|
|
30
|
-
These new records are not automatically uploaded. They do not change the pre-existing user-profile storage behavior. General feedback and disagreement with the baseline cannot be interpreted as accuracy measurements without case-level review.
|
|
31
|
-
|
|
32
|
-
Settings requests validate the expected project. Writes preserve omitted fields, serialize concurrent updates, and replace files atomically. The frontend only displays acknowledged changes, discards old project responses, imposes a five-second request/body deadline, and does not automatically retry mutations whose acknowledgement may have been lost.
|
|
33
|
-
|
|
34
|
-
## Verification and remaining limits
|
|
35
|
-
|
|
36
|
-
Verified locally:
|
|
37
|
-
|
|
38
|
-
- 16 runtime tests, including real `extractUserSignals` to `updateUserModel` integration against an isolated temporary profile; positive, negative, OFF, timeout, privacy, and fallback cases.
|
|
39
|
-
- 54 existing research tests; combined Node run: **70/70**.
|
|
40
|
-
- 13 frontend controller/SSR tests; full Dashboard TypeScript check passed.
|
|
41
|
-
- 8 Rust settings tests, including actual temporary-file persistence, concurrent partial updates, explicit false, project/stale-write rejection, and feedback retention.
|
|
42
|
-
- ESLint completed with zero errors and 31 style warnings; diff whitespace checks passed.
|
|
43
|
-
- A real isolated Node worker without a key returned `missing_key` without network access.
|
|
44
|
-
- Local HTTP rejection probes: status GET with a wrong project returned `409 project_mismatch`; models POST with a wrong project returned 409; invalid feedback verdict returned 422. These probes did not enable the feature or save feedback.
|
|
45
|
-
|
|
46
|
-
No real historical messages were submitted to Jev during this implementation. The production adapter has not undergone a live model-quality evaluation, browser DOM interaction acceptance, or installed-build validation. Previous synthetic latency results are evidence for the earlier smoke requests, not a latency guarantee for this integrated workflow.
|
|
47
|
-
|
|
48
|
-
Changes remain local and uncommitted; no installer was built or released. The packaging script `dashboard/scripts/kenaz-core-staging.js` stages Git-tracked files only. New runtime files must be included in the tracked release changes before packaging, or the optional adapter will be absent and Muninn will retain the baseline.
|
|
49
|
-
|
|
50
|
-
Follow-up [packaging verification](PACKAGING_VERIFICATION.md) exercised the real staging function in an isolated candidate and passed staged-worker no-key and mocked decision checks. The real index still omits the untracked helper; this verification does not remove that release prerequisite or establish installed-build acceptance.
|
|
51
|
-
|
|
52
|
-
See [evaluation status](EVALUATION_STATUS.md) for the unchanged quality evidence and historical-corpus limitations.
|
|
1
|
+
# Active Jev correction experiment
|
|
2
|
+
|
|
3
|
+
Local implementation and verification: 2026-09-20. This is a user-authorized opt-in experiment, not a claim of validated model quality or an installer release.
|
|
4
|
+
|
|
5
|
+
## Enable and disable
|
|
6
|
+
|
|
7
|
+
In **Default Models**, use the single **Jev-assisted judgments** switch for the current project (TASK_1970). It controls correction detection and Thalamus context ranking together, defaults to OFF, and saves both `jevCorrectionEnabled` and `jevThalamusEnabled` in one atomic settings request. The runtime keeps these project-local fields for compatibility; it does not inherit them from the Kenaz core project. Partial legacy settings are displayed as mixed, and the next check enables both.
|
|
8
|
+
|
|
9
|
+
Provide your own `TYPESAFE_API_KEY` in the environment used to launch Kenaz, then restart Kenaz and its workers after changing that environment. The UI reports key presence, not a successful authentication check. There is no key input or key storage in these settings. On Windows, TASK_1971 startup reads the saved user environment key from HKCU only when the process variable is absent, so an older parent terminal does not prevent loading it. An explicitly present process value always wins; existing workers still need relaunching to inherit changes. Users without a key retain the existing workflow, even if the enabled preference was saved.
|
|
10
|
+
|
|
11
|
+
ON makes valid Jev results authoritative for correction classification on the selected messages: positives add correction and make it the primary signal; negatives remove correction. Other label values, their relative order, and existing translation text are preserved. Existing consumers that process only the primary signal consequently follow correction instead of another primary label on positive results. Jev is not a text-generation model and does not replace the Ansuz or dispatch model selection.
|
|
12
|
+
|
|
13
|
+
OFF restores the existing workflow for **future** judgments. It does not erase or undo previously accumulated user-profile signals, expertise changes, or feedback. In-flight results are discarded if the flag is OFF when they return.
|
|
14
|
+
|
|
15
|
+
The shared extraction function is called by ordinary Muninn observations, the incremental Muninn watcher, and manual `backfill-user-model.js` runs. Consequently, running a manual historical backfill with this project flag and a valid key can also send eligible historical messages to Jev. No such backfill was performed for this implementation. The flag is project-local, while the existing profile writer stores results in the cross-project `~/.kenaz/user_profile.json`. Jev correction classification is not currently connected to Thalamus stimulation or context ranking.
|
|
16
|
+
|
|
17
|
+
## Execution and fallback
|
|
18
|
+
|
|
19
|
+
The existing synchronous Muninn/Haiku translation and classification still runs first. Jev then overlays correction decisions using fixed model `jev-1.13.0` and experimental probability threshold **0.5**, which is not calibrated on human-reviewed historical cases. This integration does not establish a whole-workflow speedup.
|
|
20
|
+
|
|
21
|
+
Each batch selects at most eight eligible recent user messages from a bounded scan. Each request contains only its target message and up to four preceding visible conversation entries; no later message is shared with earlier targets. Known credential patterns and generated instruction blocks are filtered before transmission. The combined worker input is capped at 32 KiB. Opting in sends this limited conversation context to TypeSafe.
|
|
22
|
+
|
|
23
|
+
A hidden Node worker sends isolated requests concurrently, with a shared 800 ms network deadline including response bodies and a 1,500 ms outer process timeout. There are no retries or configurable endpoint redirects. Any selected request failure or invalid response retains the complete original classification. Missing configuration, OFF, or a missing/invalid key makes no Jev request. If Haiku fails but Jev succeeds, positive correction-only signals can be emitted without inventing a translation.
|
|
24
|
+
|
|
25
|
+
## Local feedback
|
|
26
|
+
|
|
27
|
+
- `<project>/.kenaz/jev/decisions.jsonl`: best-effort, size-limited decision metadata, including hashed message IDs, probabilities, original correction flags, changed flags, latency, and fallback reasons. It excludes message text, keys, headers, and provider error bodies.
|
|
28
|
+
- `<project>/.kenaz/jev/feedback.jsonl`: up to 200 **Helpful / Incorrect judgment** feedback events, with timestamp, model, enabled preference, verdict, and feature scope. The unified control records `feature: all`; historical correction/Thalamus classifications are preserved, and missing historical feature means correction. These describe overall experience, not adjudicated labels for individual cases.
|
|
29
|
+
|
|
30
|
+
These new records are not automatically uploaded. They do not change the pre-existing user-profile storage behavior. General feedback and disagreement with the baseline cannot be interpreted as accuracy measurements without case-level review.
|
|
31
|
+
|
|
32
|
+
Settings requests validate the expected project. Writes preserve omitted fields, serialize concurrent updates, and replace files atomically. The frontend only displays acknowledged changes, discards old project responses, imposes a five-second request/body deadline, and does not automatically retry mutations whose acknowledgement may have been lost.
|
|
33
|
+
|
|
34
|
+
## Verification and remaining limits
|
|
35
|
+
|
|
36
|
+
Verified locally:
|
|
37
|
+
|
|
38
|
+
- 16 runtime tests, including real `extractUserSignals` to `updateUserModel` integration against an isolated temporary profile; positive, negative, OFF, timeout, privacy, and fallback cases.
|
|
39
|
+
- 54 existing research tests; combined Node run: **70/70**.
|
|
40
|
+
- 13 frontend controller/SSR tests; full Dashboard TypeScript check passed.
|
|
41
|
+
- 8 Rust settings tests, including actual temporary-file persistence, concurrent partial updates, explicit false, project/stale-write rejection, and feedback retention.
|
|
42
|
+
- ESLint completed with zero errors and 31 style warnings; diff whitespace checks passed.
|
|
43
|
+
- A real isolated Node worker without a key returned `missing_key` without network access.
|
|
44
|
+
- Local HTTP rejection probes: status GET with a wrong project returned `409 project_mismatch`; models POST with a wrong project returned 409; invalid feedback verdict returned 422. These probes did not enable the feature or save feedback.
|
|
45
|
+
|
|
46
|
+
No real historical messages were submitted to Jev during this implementation. The production adapter has not undergone a live model-quality evaluation, browser DOM interaction acceptance, or installed-build validation. Previous synthetic latency results are evidence for the earlier smoke requests, not a latency guarantee for this integrated workflow.
|
|
47
|
+
|
|
48
|
+
Changes remain local and uncommitted; no installer was built or released. The packaging script `dashboard/scripts/kenaz-core-staging.js` stages Git-tracked files only. New runtime files must be included in the tracked release changes before packaging, or the optional adapter will be absent and Muninn will retain the baseline.
|
|
49
|
+
|
|
50
|
+
Follow-up [packaging verification](PACKAGING_VERIFICATION.md) exercised the real staging function in an isolated candidate and passed staged-worker no-key and mocked decision checks. The real index still omits the untracked helper; this verification does not remove that release prerequisite or establish installed-build acceptance.
|
|
51
|
+
|
|
52
|
+
See [evaluation status](EVALUATION_STATUS.md) for the unchanged quality evidence and historical-corpus limitations.
|
|
@@ -1,31 +1,31 @@
|
|
|
1
|
-
# Historical context review
|
|
2
|
-
|
|
3
|
-
This stage restores observed conversational evidence for the local pending corpus. It does not assign correction labels or infer a task goal. No source text is sent to Jev or another external model.
|
|
4
|
-
|
|
5
|
-
## Review order
|
|
6
|
-
|
|
7
|
-
1. Confirm the record is an actual user interaction rather than a smoke test, generated prompt, pasted transcript or unrelated task. A session cwd proves workspace association, not authorship or topic.
|
|
8
|
-
2. Read the target message and only the observed preceding context. Check source lineage, missing/truncated/redacted context and compaction boundaries. When context is insufficient, keep the decision unresolved.
|
|
9
|
-
3. Identify the exact prior assertion or plan that the user changes. Quote the local evidence in the review, distinguishing an observed statement from an annotator inference.
|
|
10
|
-
4. A correction must change the previous understanding or direction; a question, status request, new requirement, agreement or quoted correction is not automatically positive. See ANNOTATION_GUIDE.md.
|
|
11
|
-
5. Keep decisions blank until reviewed. Model/agent proposals must remain agent_proposed. A context restoration status is never human_verified and does not make a case scoring-ready.
|
|
12
|
-
|
|
13
|
-
## Provenance limits
|
|
14
|
-
|
|
15
|
-
Original pending manifests contain source-relative paths and line numbers, but no historic source snapshot hash. Target identity/text checks and newly computed prefix hashes bind this review to the source as read now; they cannot prove all preceding text is unchanged since the earlier export. Context reconstruction is bounded and intentionally omits tool output, internal reasoning, secrets and unsupported branches. Session and family grouping need review before any held-out split.
|
|
16
|
-
|
|
17
|
-
Private review artifacts stay in the ignored `.kenaz/jev-corpus/` directory. Public reports contain aggregate counts only. The viewer is a read-only aid; opening it does not submit or save a review. Labels, reviewer identity and approval remain separate, explicit review records.
|
|
18
|
-
|
|
19
|
-
## Verified result ? 2026-09-20
|
|
20
|
-
|
|
21
|
-
All 458 candidate IDs retained: 419 partial context packets, 23 missing context, 16 rejected for ambiguous source identity. No case is marked fully restored or human-reviewed. The 16 rejections are in one Claude file with repeated attachment UUIDs whose parent links differ; duplicate display metadata alone is tolerated.
|
|
22
|
-
|
|
23
|
-
54/54 tests passed. Parent independently verified 442 source-prefix hashes and 2,352 exported context items against original visible text, source line order and filtering rules. All 458 HTML articles are escaped, without scripts, iframes, images or external stylesheets. Source pending file remains unchanged. Aggregate evidence: [context report](reports/context-review-2026-09-20.json).
|
|
24
|
-
|
|
25
|
-
Current private artifact: `.kenaz/jev-corpus/context-2026-09-20-v2/review.html`; matching `review-context.jsonl` holds blank decisions. First export `context-2026-09-20` is REJECTED: a JSON-escaped quoted credential escaped a precheck; raw-text, normalized-text and final-output checks now prevent that known pattern, with a regression test. This does not guarantee removal of every possible secret or personal detail.
|
|
26
|
-
|
|
27
|
-
```powershell
|
|
28
|
-
node dev/jev-eval/context.mjs --pending .kenaz/jev-corpus/history-2026-09-20-v2/pending.jsonl --manifest .kenaz/jev-corpus/history-2026-09-20-v2/manifest.json --out-dir .kenaz/jev-corpus/context-new
|
|
29
|
-
```
|
|
30
|
-
|
|
31
|
-
The next stage is evidence-based candidate triage and annotation. Cases lacking the relevant prior assumption remain unresolved. No accuracy claim, held-out split, real-data external upload or production activation has occurred.
|
|
1
|
+
# Historical context review
|
|
2
|
+
|
|
3
|
+
This stage restores observed conversational evidence for the local pending corpus. It does not assign correction labels or infer a task goal. No source text is sent to Jev or another external model.
|
|
4
|
+
|
|
5
|
+
## Review order
|
|
6
|
+
|
|
7
|
+
1. Confirm the record is an actual user interaction rather than a smoke test, generated prompt, pasted transcript or unrelated task. A session cwd proves workspace association, not authorship or topic.
|
|
8
|
+
2. Read the target message and only the observed preceding context. Check source lineage, missing/truncated/redacted context and compaction boundaries. When context is insufficient, keep the decision unresolved.
|
|
9
|
+
3. Identify the exact prior assertion or plan that the user changes. Quote the local evidence in the review, distinguishing an observed statement from an annotator inference.
|
|
10
|
+
4. A correction must change the previous understanding or direction; a question, status request, new requirement, agreement or quoted correction is not automatically positive. See ANNOTATION_GUIDE.md.
|
|
11
|
+
5. Keep decisions blank until reviewed. Model/agent proposals must remain agent_proposed. A context restoration status is never human_verified and does not make a case scoring-ready.
|
|
12
|
+
|
|
13
|
+
## Provenance limits
|
|
14
|
+
|
|
15
|
+
Original pending manifests contain source-relative paths and line numbers, but no historic source snapshot hash. Target identity/text checks and newly computed prefix hashes bind this review to the source as read now; they cannot prove all preceding text is unchanged since the earlier export. Context reconstruction is bounded and intentionally omits tool output, internal reasoning, secrets and unsupported branches. Session and family grouping need review before any held-out split.
|
|
16
|
+
|
|
17
|
+
Private review artifacts stay in the ignored `.kenaz/jev-corpus/` directory. Public reports contain aggregate counts only. The viewer is a read-only aid; opening it does not submit or save a review. Labels, reviewer identity and approval remain separate, explicit review records.
|
|
18
|
+
|
|
19
|
+
## Verified result ? 2026-09-20
|
|
20
|
+
|
|
21
|
+
All 458 candidate IDs retained: 419 partial context packets, 23 missing context, 16 rejected for ambiguous source identity. No case is marked fully restored or human-reviewed. The 16 rejections are in one Claude file with repeated attachment UUIDs whose parent links differ; duplicate display metadata alone is tolerated.
|
|
22
|
+
|
|
23
|
+
54/54 tests passed. Parent independently verified 442 source-prefix hashes and 2,352 exported context items against original visible text, source line order and filtering rules. All 458 HTML articles are escaped, without scripts, iframes, images or external stylesheets. Source pending file remains unchanged. Aggregate evidence: [context report](reports/context-review-2026-09-20.json).
|
|
24
|
+
|
|
25
|
+
Current private artifact: `.kenaz/jev-corpus/context-2026-09-20-v2/review.html`; matching `review-context.jsonl` holds blank decisions. First export `context-2026-09-20` is REJECTED: a JSON-escaped quoted credential escaped a precheck; raw-text, normalized-text and final-output checks now prevent that known pattern, with a regression test. This does not guarantee removal of every possible secret or personal detail.
|
|
26
|
+
|
|
27
|
+
```powershell
|
|
28
|
+
node dev/jev-eval/context.mjs --pending .kenaz/jev-corpus/history-2026-09-20-v2/pending.jsonl --manifest .kenaz/jev-corpus/history-2026-09-20-v2/manifest.json --out-dir .kenaz/jev-corpus/context-new
|
|
29
|
+
```
|
|
30
|
+
|
|
31
|
+
The next stage is evidence-based candidate triage and annotation. Cases lacking the relevant prior assumption remain unresolved. No accuracy claim, held-out split, real-data external upload or production activation has occurred.
|
|
@@ -1,39 +1,39 @@
|
|
|
1
|
-
# Jev evaluation status ? 2026-09-20
|
|
2
|
-
|
|
3
|
-
**Conclusion: promising synthetic latency and engineering feasibility; real-workflow quality and replacement value remain unproven.** A separate, default-off production correction adapter and Dashboard switch are now implemented locally at the user's request. The research harness itself is not imported into production. See [active correction experiment](ACTIVE_CORRECTION.md) for behavior, verification, and release limits.
|
|
4
|
-
|
|
5
|
-
| Evidence | Observed result | Interpretation limit |
|
|
6
|
-
| --- | --- | --- |
|
|
7
|
-
| Live Jev smoke | Two retained runs, 11/11 successful each; p50 272.66ms and 271.89ms | Same 11 synthetic cases; middle run excluded after output truncation; cache state unknown |
|
|
8
|
-
| Synthetic correction | 8/8 matching agent-proposed labels per run | Repeated cases, not 16 independent samples; not human-reviewed accuracy |
|
|
9
|
-
| Synthetic reranking | NDCG@5 = 1 on two evaluable queries | Third zero-relevance query excluded; far below 100-query target |
|
|
10
|
-
| Synthetic length probe | 9/9 succeeded at 800ms deadline; maximum observed 715ms | Repetitive filler, three calls per length; no robust p99 or concurrency evidence |
|
|
11
|
-
| Engineering checks | Latest 54/54 tests passed | Research harness and context tests, not installed-application acceptance |
|
|
12
|
-
| Historical corpus | 458 unique candidates, 68 session IDs; 42 short of 500 | User-role history is not proof of human authorship or representativeness |
|
|
13
|
-
| Restored context | 419 partial, 23 missing, 16 rejected | Context availability is not annotation or scoring readiness |
|
|
14
|
-
| Annotation pilot | 12 cases: 9 matching non-null agent proposals, 1 shared null, 2 disagreements | Same model family, no independent human review; not a Jev result |
|
|
15
|
-
| Haiku baseline | No successful batch in recorded run; CLI weekly quota blocked it | No fair quality or speedup comparison available |
|
|
16
|
-
|
|
17
|
-
## Thalamus local integration
|
|
18
|
-
|
|
19
|
-
A separate default-off Thalamus delivery-ranking experiment is implemented locally. It affects native Claude Active Sensing order and MCP calls that provide taskContext; it does not change neural dynamics or claim automatic Bifrost integration. See [Thalamus ranking](THALAMUS_RANKING.md) for tests and release limits. These engineering checks do not establish real Jev relevance improvement.
|
|
20
|
-
|
|
21
|
-
## New local annotation findings
|
|
22
|
-
|
|
23
|
-
Two separately produced agent reviews covered 12 short-context cases from distinct inferred groups. Selection was deterministic by ID, with four cases in each of three target-length bins and a maximum of 1,800 characters of target plus observed context. This convenience pilot is neither representative nor held out.
|
|
24
|
-
|
|
25
|
-
Nine non-null proposals matched: two correction and seven non-correction. One case was unresolved by both reviewers; two had null versus false disagreements. All three remain unresolved in the comparison artifact. Parent verified all 12 IDs in each review and all 49 exact evidence substrings. Both reviewers are agents from the same model family: agreement does not establish truth. Human-verified labels remain zero, and no Jev predictions were made on these real cases.
|
|
26
|
-
|
|
27
|
-
The difficult boundaries were: whether a session restart addresses a suggested MCP-server restart; whether a question about retiring an older task changes a prior plan absent from context; and whether a new downstream error contradicts a narrowly scoped successful repair. These require precise premise and scope definitions. A question mark or a negative word in an error log cannot decide the label by itself.
|
|
28
|
-
|
|
29
|
-
Private pilot artifacts remain under `.kenaz/jev-corpus/pilot-2026-09-20/`. Public aggregate: [annotation pilot](reports/annotation-pilot-2026-09-20.json). The nine agreements are still `agent_proposed`, and the comparison explicitly sets scoring eligibility false. No original corpus labels or splits were changed.
|
|
30
|
-
|
|
31
|
-
## Product implication
|
|
32
|
-
|
|
33
|
-
Continue bounded research on correction detection and recall ranking. The user explicitly authorized an optional active experiment: OFF keeps existing behavior; ON lets validated Jev responses add/remove correction and prioritize positive correction for existing primary-type consumers. This opt-in decision does not establish quality. The actual Muninn/Haiku routine still translates and emits seven signal labels; Jev has a smaller output contract. A fair comparison must account for those downstream consumers rather than comparing latency alone.
|
|
34
|
-
|
|
35
|
-
The separate active adapter now has 16 passing runtime tests, including downstream user-profile effects and OFF/failure behavior; the unchanged research suite passes 54 tests. Dashboard controller/SSR tests pass 13 cases and Rust settings tests pass 8 cases. These are engineering checks, not quality scores or installed-build acceptance. Users without their own Jev key keep the existing workflow. No shared developer key is stored or shipped.
|
|
36
|
-
|
|
37
|
-
Next evidence needed: recover the material premises for unresolved cases and freeze the annotation rule; expand context-grounded agent proposals without treating them as gold; independent human adjudication; fixed model/question/threshold and lineage split; equivalent Haiku/Synapse baselines; installed-build no-key/offline/failure acceptance. Current real cases have not been submitted to Jev or Haiku.
|
|
38
|
-
|
|
39
|
-
Sources: [live smoke](reports/live-smoke-2026-09-20.json), [validation progress](VALIDATION_PROGRESS.md), [baseline](BASELINE_COMPARISON.md), [historical corpus](HISTORICAL_CORPUS.md), [context review](CONTEXT_REVIEW.md).
|
|
1
|
+
# Jev evaluation status ? 2026-09-20
|
|
2
|
+
|
|
3
|
+
**Conclusion: promising synthetic latency and engineering feasibility; real-workflow quality and replacement value remain unproven.** A separate, default-off production correction adapter and Dashboard switch are now implemented locally at the user's request. The research harness itself is not imported into production. See [active correction experiment](ACTIVE_CORRECTION.md) for behavior, verification, and release limits.
|
|
4
|
+
|
|
5
|
+
| Evidence | Observed result | Interpretation limit |
|
|
6
|
+
| --- | --- | --- |
|
|
7
|
+
| Live Jev smoke | Two retained runs, 11/11 successful each; p50 272.66ms and 271.89ms | Same 11 synthetic cases; middle run excluded after output truncation; cache state unknown |
|
|
8
|
+
| Synthetic correction | 8/8 matching agent-proposed labels per run | Repeated cases, not 16 independent samples; not human-reviewed accuracy |
|
|
9
|
+
| Synthetic reranking | NDCG@5 = 1 on two evaluable queries | Third zero-relevance query excluded; far below 100-query target |
|
|
10
|
+
| Synthetic length probe | 9/9 succeeded at 800ms deadline; maximum observed 715ms | Repetitive filler, three calls per length; no robust p99 or concurrency evidence |
|
|
11
|
+
| Engineering checks | Latest 54/54 tests passed | Research harness and context tests, not installed-application acceptance |
|
|
12
|
+
| Historical corpus | 458 unique candidates, 68 session IDs; 42 short of 500 | User-role history is not proof of human authorship or representativeness |
|
|
13
|
+
| Restored context | 419 partial, 23 missing, 16 rejected | Context availability is not annotation or scoring readiness |
|
|
14
|
+
| Annotation pilot | 12 cases: 9 matching non-null agent proposals, 1 shared null, 2 disagreements | Same model family, no independent human review; not a Jev result |
|
|
15
|
+
| Haiku baseline | No successful batch in recorded run; CLI weekly quota blocked it | No fair quality or speedup comparison available |
|
|
16
|
+
|
|
17
|
+
## Thalamus local integration
|
|
18
|
+
|
|
19
|
+
A separate default-off Thalamus delivery-ranking experiment is implemented locally. It affects native Claude Active Sensing order and MCP calls that provide taskContext; it does not change neural dynamics or claim automatic Bifrost integration. See [Thalamus ranking](THALAMUS_RANKING.md) for tests and release limits. These engineering checks do not establish real Jev relevance improvement.
|
|
20
|
+
|
|
21
|
+
## New local annotation findings
|
|
22
|
+
|
|
23
|
+
Two separately produced agent reviews covered 12 short-context cases from distinct inferred groups. Selection was deterministic by ID, with four cases in each of three target-length bins and a maximum of 1,800 characters of target plus observed context. This convenience pilot is neither representative nor held out.
|
|
24
|
+
|
|
25
|
+
Nine non-null proposals matched: two correction and seven non-correction. One case was unresolved by both reviewers; two had null versus false disagreements. All three remain unresolved in the comparison artifact. Parent verified all 12 IDs in each review and all 49 exact evidence substrings. Both reviewers are agents from the same model family: agreement does not establish truth. Human-verified labels remain zero, and no Jev predictions were made on these real cases.
|
|
26
|
+
|
|
27
|
+
The difficult boundaries were: whether a session restart addresses a suggested MCP-server restart; whether a question about retiring an older task changes a prior plan absent from context; and whether a new downstream error contradicts a narrowly scoped successful repair. These require precise premise and scope definitions. A question mark or a negative word in an error log cannot decide the label by itself.
|
|
28
|
+
|
|
29
|
+
Private pilot artifacts remain under `.kenaz/jev-corpus/pilot-2026-09-20/`. Public aggregate: [annotation pilot](reports/annotation-pilot-2026-09-20.json). The nine agreements are still `agent_proposed`, and the comparison explicitly sets scoring eligibility false. No original corpus labels or splits were changed.
|
|
30
|
+
|
|
31
|
+
## Product implication
|
|
32
|
+
|
|
33
|
+
Continue bounded research on correction detection and recall ranking. The user explicitly authorized an optional active experiment: OFF keeps existing behavior; ON lets validated Jev responses add/remove correction and prioritize positive correction for existing primary-type consumers. This opt-in decision does not establish quality. The actual Muninn/Haiku routine still translates and emits seven signal labels; Jev has a smaller output contract. A fair comparison must account for those downstream consumers rather than comparing latency alone.
|
|
34
|
+
|
|
35
|
+
The separate active adapter now has 16 passing runtime tests, including downstream user-profile effects and OFF/failure behavior; the unchanged research suite passes 54 tests. Dashboard controller/SSR tests pass 13 cases and Rust settings tests pass 8 cases. These are engineering checks, not quality scores or installed-build acceptance. Users without their own Jev key keep the existing workflow. No shared developer key is stored or shipped.
|
|
36
|
+
|
|
37
|
+
Next evidence needed: recover the material premises for unresolved cases and freeze the annotation rule; expand context-grounded agent proposals without treating them as gold; independent human adjudication; fixed model/question/threshold and lineage split; equivalent Haiku/Synapse baselines; installed-build no-key/offline/failure acceptance. Current real cases have not been submitted to Jev or Haiku.
|
|
38
|
+
|
|
39
|
+
Sources: [live smoke](reports/live-smoke-2026-09-20.json), [validation progress](VALIDATION_PROGRESS.md), [baseline](BASELINE_COMPARISON.md), [historical corpus](HISTORICAL_CORPUS.md), [context review](CONTEXT_REVIEW.md).
|
|
@@ -1,33 +1,33 @@
|
|
|
1
|
-
# Local historical corpus preparation ? 2026-09-20
|
|
2
|
-
|
|
3
|
-
User authorized local Kenaz-related historical sessions on this machine. No transcript was sent to Jev, Claude, or another external model. Source history was read only.
|
|
4
|
-
|
|
5
|
-
## Result
|
|
6
|
-
|
|
7
|
-
638 files discovered in the four Kenaz-named Claude project directories and Codex session tree. Exact cwd checks restrict candidate records to the KenazAI tree and installed Kenaz core. This is workspace-based relevance, not proof that every message discusses Kenaz product code. Nested subagents are excluded. No archived Codex files were found by the parent inventory.
|
|
8
|
-
|
|
9
|
-
580 eligible candidate records became **458 unique pending messages across 68 session IDs / 68 inferred groups** after 122 duplicate removals. Sources: Claude 160, Codex 298. Target 500 has a **42-message shortfall**; no synthetic padding was added. These are source-history candidates, not independently verified human-authored gold. Shared lineage cannot be completely inferred from logs; session IDs are not proof of statistical independence.
|
|
10
|
-
|
|
11
|
-
Private artifacts: `.kenaz/jev-corpus/history-2026-09-20-v2/{pending.jsonl,manifest.json,aggregate.json}`. The private root is excluded using local `.git/info/exclude`. Do not commit or upload it. The manifest maps hashed IDs to local relative source paths and line numbers. Public aggregate: [report](reports/history-corpus-2026-09-20.json).
|
|
12
|
-
|
|
13
|
-
The first export (`history-2026-09-20`) is marked REJECTED: parent spot checks caught dispatch and summarization prompts incorrectly recorded as user-role events. Source confirmation in `.formation/tools/prompt.js:HEADLESS_PREFIX` and `.formation/tools/muninn.js:translateChunk` led to whole-session exclusion, including matching session IDs across files. Its 500-row count is invalid and superseded.
|
|
14
|
-
|
|
15
|
-
## Contract and remaining review
|
|
16
|
-
|
|
17
|
-
Every row is `pending_historical_review`, `unreviewed`, with no fabricated goal, assumptions or labels. All splits remain unassigned. Context contains at most the preceding eligible user message; 57 rows have none, and all rows explicitly flag incomplete correction context because assistant assumptions are absent. This file intentionally cannot be scored as a schema-v1 labelled evaluation dataset or applied directly with review.mjs.
|
|
18
|
-
|
|
19
|
-
Next review must verify human authorship and topical suitability, retrieve preceding assistant context where necessary from the local manifest, establish the actual goal/assumptions, and annotate independently. Ambiguous cases remain excluded from scoring. Assign splits by lineage only after this review. Local manual review is also needed for privacy: pattern filters are not a guarantee against every credential or personal detail.
|
|
20
|
-
|
|
21
|
-
44/44 offline tests passed, including 11 corpus-specific tests. Parent separately checked exact row/ID/normalized-message counts, manifest coverage, pending provenance and obvious credential/automation markers, then spot-checked message prefixes. Neither these checks nor 458 pending rows complete the planned human-reviewed 500-message dataset. Recall and handoff targets also remain separate work. Production Jev remains disabled; the Haiku baseline remains blocked by CLI weekly quota.
|
|
22
|
-
|
|
23
|
-
## Reproduce
|
|
24
|
-
|
|
25
|
-
```powershell
|
|
26
|
-
# Inventory only, aggregate stdout; no files written
|
|
27
|
-
node dev/jev-eval/corpus.mjs --target 500
|
|
28
|
-
# Explicit local export: requires ignored root and a NEW direct-child directory
|
|
29
|
-
node dev/jev-eval/corpus.mjs --target 500 --out-dir .kenaz/jev-corpus/history-new
|
|
30
|
-
node --test dev/jev-eval/test/*.test.mjs
|
|
31
|
-
```
|
|
32
|
-
|
|
33
|
-
The extractor streams with line/file/candidate budgets and reports truncations. Current scan discarded 10 oversized lines; it did not exhaust file or candidate budgets. Counts of exclusion events may overlap and are not distinct-person or distinct-message counts. Credentials trigger whole-message omission, followed by context reset; known automated session seeds reject the whole source/session. Generic repeated short messages are globally deduplicated without asserting common lineage; copied messages of at least 160 characters conservatively join groups. No split has been assigned.
|
|
1
|
+
# Local historical corpus preparation ? 2026-09-20
|
|
2
|
+
|
|
3
|
+
User authorized local Kenaz-related historical sessions on this machine. No transcript was sent to Jev, Claude, or another external model. Source history was read only.
|
|
4
|
+
|
|
5
|
+
## Result
|
|
6
|
+
|
|
7
|
+
638 files discovered in the four Kenaz-named Claude project directories and Codex session tree. Exact cwd checks restrict candidate records to the KenazAI tree and installed Kenaz core. This is workspace-based relevance, not proof that every message discusses Kenaz product code. Nested subagents are excluded. No archived Codex files were found by the parent inventory.
|
|
8
|
+
|
|
9
|
+
580 eligible candidate records became **458 unique pending messages across 68 session IDs / 68 inferred groups** after 122 duplicate removals. Sources: Claude 160, Codex 298. Target 500 has a **42-message shortfall**; no synthetic padding was added. These are source-history candidates, not independently verified human-authored gold. Shared lineage cannot be completely inferred from logs; session IDs are not proof of statistical independence.
|
|
10
|
+
|
|
11
|
+
Private artifacts: `.kenaz/jev-corpus/history-2026-09-20-v2/{pending.jsonl,manifest.json,aggregate.json}`. The private root is excluded using local `.git/info/exclude`. Do not commit or upload it. The manifest maps hashed IDs to local relative source paths and line numbers. Public aggregate: [report](reports/history-corpus-2026-09-20.json).
|
|
12
|
+
|
|
13
|
+
The first export (`history-2026-09-20`) is marked REJECTED: parent spot checks caught dispatch and summarization prompts incorrectly recorded as user-role events. Source confirmation in `.formation/tools/prompt.js:HEADLESS_PREFIX` and `.formation/tools/muninn.js:translateChunk` led to whole-session exclusion, including matching session IDs across files. Its 500-row count is invalid and superseded.
|
|
14
|
+
|
|
15
|
+
## Contract and remaining review
|
|
16
|
+
|
|
17
|
+
Every row is `pending_historical_review`, `unreviewed`, with no fabricated goal, assumptions or labels. All splits remain unassigned. Context contains at most the preceding eligible user message; 57 rows have none, and all rows explicitly flag incomplete correction context because assistant assumptions are absent. This file intentionally cannot be scored as a schema-v1 labelled evaluation dataset or applied directly with review.mjs.
|
|
18
|
+
|
|
19
|
+
Next review must verify human authorship and topical suitability, retrieve preceding assistant context where necessary from the local manifest, establish the actual goal/assumptions, and annotate independently. Ambiguous cases remain excluded from scoring. Assign splits by lineage only after this review. Local manual review is also needed for privacy: pattern filters are not a guarantee against every credential or personal detail.
|
|
20
|
+
|
|
21
|
+
44/44 offline tests passed, including 11 corpus-specific tests. Parent separately checked exact row/ID/normalized-message counts, manifest coverage, pending provenance and obvious credential/automation markers, then spot-checked message prefixes. Neither these checks nor 458 pending rows complete the planned human-reviewed 500-message dataset. Recall and handoff targets also remain separate work. Production Jev remains disabled; the Haiku baseline remains blocked by CLI weekly quota.
|
|
22
|
+
|
|
23
|
+
## Reproduce
|
|
24
|
+
|
|
25
|
+
```powershell
|
|
26
|
+
# Inventory only, aggregate stdout; no files written
|
|
27
|
+
node dev/jev-eval/corpus.mjs --target 500
|
|
28
|
+
# Explicit local export: requires ignored root and a NEW direct-child directory
|
|
29
|
+
node dev/jev-eval/corpus.mjs --target 500 --out-dir .kenaz/jev-corpus/history-new
|
|
30
|
+
node --test dev/jev-eval/test/*.test.mjs
|
|
31
|
+
```
|
|
32
|
+
|
|
33
|
+
The extractor streams with line/file/candidate budgets and reports truncations. Current scan discarded 10 oversized lines; it did not exhaust file or candidate budgets. Counts of exclusion events may overlap and are not distinct-person or distinct-message counts. Credentials trigger whole-message omission, followed by context reset; known automated session seeds reject the whole source/session. Generic repeated short messages are globally deduplicated without asserting common lineage; copied messages of at least 160 characters conservatively join groups. No split has been assigned.
|