virtual-context 0.2.9__tar.gz → 0.3.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {virtual_context-0.2.9 → virtual_context-0.3.0}/PKG-INFO +129 -29
- {virtual_context-0.2.9 → virtual_context-0.3.0}/README.md +128 -28
- {virtual_context-0.2.9 → virtual_context-0.3.0}/pyproject.toml +1 -1
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_cli_init.py +67 -11
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/__init__.py +1 -1
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/cli/main.py +21 -4
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/config.py +9 -1
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/proxy/formats.py +332 -63
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/proxy/message_filter.py +4 -4
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/proxy/server.py +30 -18
- {virtual_context-0.2.9 → virtual_context-0.3.0}/.gitignore +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/LICENSE +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/assets/dashboard.png +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/assets/hero.png +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/models.yaml +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/REGRESSION_MAP.md +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/conftest.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/docker-compose.test.yml +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/haiku/__init__.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/haiku/conftest.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/haiku/test_compaction.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/haiku/test_retrieval.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/haiku/test_tagging.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/ollama/__init__.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/ollama/conftest.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/ollama/test_compactor.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/ollama/test_pipeline.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/ollama/test_provider.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/ollama/test_tag_generator.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/proxy/__init__.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/proxy/test_dashboard_cors.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/proxy/test_metrics.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_assembler.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_backend_integration.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_budget_enforcement.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_compaction_commit_prune.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_compactor.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_compactor_concurrent.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_composite_store.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_config.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_context_bleed.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_conversation_identity.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_conversation_lifecycle.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_conversation_scoping.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_embedding_tag_generator.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_empty_turn_skip.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_engine_integration.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_engine_lookback.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_engine_state.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_fact_enrichment.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_fact_graph_integration.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_fact_link_checker.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_fact_link_query.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_fact_link_types.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_fact_links_sqlite.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_fact_redesign.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_fill_pass.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_find_quote.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_format_agnostic.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_headless.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_history_filter.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_idf_retrieval.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_ingest_index_integrity.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_longmemeval_auth.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_mcp_server.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_media.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_message_filter.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_metrics_persistence.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_model_catalog.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_model_limits.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_monitor.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_multi_instance.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_noop_fact_link_store.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_openrouter_provider.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_paging.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_passthrough_filter.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_presets.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_prev_context_leak.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_provider_adapters.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_proxy.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_proxy_dashboard.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_proxy_formats.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_proxy_message_filter.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_proxy_session.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_proxy_streaming.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_raw_content.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_recall_all.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_registry_lifecycle.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_request_captures_persistence.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_retriever.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_rrf_scoring.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_segmenter.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_semantic_search.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_sender_identity.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_session_cache.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_session_date.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_storage_protocols.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_store_recovery.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_store_sqlite.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_stub_turn_handling.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_supersession.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_supersession_migration.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_tag_canonicalizer.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_tag_consolidator.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_tag_generator.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_tag_splitter.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_telemetry.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_telemetry_integration.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_tool_loop.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_tool_output_interceptor.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_tool_result_filter.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_tool_tags.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_tui.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_turn_grouping.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_turn_tag_index.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_unified_budget.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_upstream_trim.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/tests/test_verb_expansion.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual-context.yaml +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual-context.yaml.example +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/cli/__init__.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/conversation_identity.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/core/__init__.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/core/assembler.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/core/compaction_pipeline.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/core/compactor.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/core/composite_store.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/core/conversation_store.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/core/embedding_provider.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/core/embedding_tag_generator.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/core/engine_utils.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/core/fact_query.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/core/fts_preprocessor.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/core/hint_builder.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/core/llm_utils.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/core/math_utils.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/core/model_catalog.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/core/monitor.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/core/paging_manager.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/core/protocols.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/core/provider_adapters.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/core/quote_search.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/core/retrieval_assembler.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/core/retrieval_scoring.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/core/retriever.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/core/search_engine.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/core/segmenter.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/core/semantic_search.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/core/store.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/core/tag_canonicalizer.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/core/tag_consolidator.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/core/tag_generator.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/core/tag_scoring.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/core/tag_splitter.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/core/tagging_pipeline.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/core/telemetry.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/core/temporal_resolver.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/core/tool_loop.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/core/tool_query.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/core/turn_tag_index.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/data/anthropic-tokenizer/tokenizer.json +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/engine.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/ingest/__init__.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/ingest/curator.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/ingest/date_resolver.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/ingest/parsers.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/ingest/supersession.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/mcp/__init__.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/mcp/server.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/model_limits.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/openclaw/virtual-context.mjs +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/patterns.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/presets/__init__.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/presets/agentic.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/presets/base.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/presets/coding.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/providers/__init__.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/providers/anthropic.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/providers/base.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/providers/generic_openai.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/providers/ollama_native.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/proxy/__init__.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/proxy/_envelope.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/proxy/dashboard.html +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/proxy/dashboard.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/proxy/handlers.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/proxy/helpers.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/proxy/media.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/proxy/metrics.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/proxy/multi.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/proxy/registry.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/proxy/session_cache.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/proxy/state.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/proxy/static/android-chrome-192x192.png +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/proxy/static/android-chrome-512x512.png +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/proxy/static/apple-touch-icon.png +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/proxy/static/favicon-16x16.png +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/proxy/static/favicon-32x32.png +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/proxy/static/favicon.ico +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/proxy/static/site.webmanifest +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/proxy/tool_output_interceptor.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/storage/__init__.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/storage/falkordb.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/storage/filesystem.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/storage/helpers.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/storage/neo4j.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/storage/noop_fact_link_store.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/storage/postgres.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/storage/sqlite.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/token_counter.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/tui/__init__.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/tui/app.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/tui/chat.tcss +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/tui/chat_provider.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/tui/headless.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/tui/modals/__init__.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/tui/modals/turn_inspector.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/tui/state.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/tui/widgets/__init__.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/tui/widgets/budget_bar.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/tui/widgets/chat_view.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/tui/widgets/input_box.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/tui/widgets/tag_panel.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/tui/widgets/turn_list.py +0 -0
- {virtual_context-0.2.9 → virtual_context-0.3.0}/virtual_context/types.py +0 -0
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.4
|
|
2
2
|
Name: virtual-context
|
|
3
|
-
Version: 0.
|
|
3
|
+
Version: 0.3.0
|
|
4
4
|
Summary: OS-style virtual memory for LLM session context management
|
|
5
5
|
Project-URL: Homepage, https://virtual-context.com
|
|
6
6
|
Project-URL: Repository, https://github.com/virtual-context/virtual-context
|
|
@@ -180,7 +180,7 @@ proxy:
|
|
|
180
180
|
**Daemon mode:** run as a background service:
|
|
181
181
|
|
|
182
182
|
```bash
|
|
183
|
-
# Creates config
|
|
183
|
+
# Creates ~/.virtualcontext/ with config + data, installs + starts daemon
|
|
184
184
|
virtual-context daemon install --upstream https://api.anthropic.com
|
|
185
185
|
|
|
186
186
|
# Or: guided interactive setup with daemon
|
|
@@ -282,6 +282,7 @@ Exposes virtual-context as an MCP server for integration with Claude Desktop, Cu
|
|
|
282
282
|
| Tool | `collapse_topic` | Collapse a topic back to summary or none |
|
|
283
283
|
| Tool | `find_quote` | Full-text search across all stored conversation text |
|
|
284
284
|
| Tool | `query_facts` | Structured fact lookup with subject/verb/object/status filters |
|
|
285
|
+
| Tool | `restore_tool` | Recover full content from compacted chain stubs or compressed media |
|
|
285
286
|
| Resource | `virtualcontext://domains` | List all tags |
|
|
286
287
|
| Resource | `virtualcontext://domains/{tag}` | Summaries for a specific tag |
|
|
287
288
|
| Prompt | `recall` | Suggest context retrieval for a topic |
|
|
@@ -293,11 +294,12 @@ Exposes virtual-context as an MCP server for integration with Claude Desktop, Cu
|
|
|
293
294
|
User message arrives
|
|
294
295
|
│
|
|
295
296
|
▼
|
|
296
|
-
|
|
297
|
-
│ ├─ Extract
|
|
298
|
-
│ ├─ Route to existing
|
|
299
|
-
│ ├─ No marker? →
|
|
300
|
-
│
|
|
297
|
+
Conversation routing (proxy mode)
|
|
298
|
+
│ ├─ Extract conversation ID from <!-- vc:conversation=UUID --> markers
|
|
299
|
+
│ ├─ Route to existing conversation or load persisted state from store + Redis
|
|
300
|
+
│ ├─ No marker? → derive stable ID from system prompt hash + format
|
|
301
|
+
│ ├─ Redis session cache: lossless restart, write-through history persistence
|
|
302
|
+
│ └─ Strip conversation markers before forwarding upstream
|
|
301
303
|
│
|
|
302
304
|
▼
|
|
303
305
|
Strip client envelope + extract metadata
|
|
@@ -307,13 +309,20 @@ Strip client envelope + extract metadata
|
|
|
307
309
|
│ └─ Metadata preserved on Message.metadata for downstream use
|
|
308
310
|
│
|
|
309
311
|
▼
|
|
312
|
+
Media compression (all paths — passthrough and active)
|
|
313
|
+
│ ├─ Detect base64 images across all 4 formats (Anthropic, OpenAI Chat, Responses, Gemini)
|
|
314
|
+
│ ├─ Compress to JPEG, store originals to disk for later recovery
|
|
315
|
+
│ ├─ Image token counting uses Anthropic formula: (width × height) / 750
|
|
316
|
+
│ └─ A 391KB screenshot → ~40KB compressed, saving ~88k tokens
|
|
317
|
+
│
|
|
318
|
+
▼
|
|
310
319
|
History ingestion (first request only)
|
|
311
320
|
│ ├─ Extract and tag all prior user+assistant pairs → bootstrap TurnTagIndex
|
|
312
321
|
│ ├─ Stub detection: media attachments/image placeholders get _stub tag (skip LLM tagger)
|
|
313
322
|
│ └─ Conversation-scoped: each conversation's index is independent
|
|
314
323
|
│
|
|
315
324
|
▼
|
|
316
|
-
Inbound tagging
|
|
325
|
+
Inbound tagging — identify what this message is about
|
|
317
326
|
│ ├─ Embedding tagger (recommended): cosine similarity against existing tag vocabulary
|
|
318
327
|
│ │ (closed-set, deterministic, can't hallucinate novel tags)
|
|
319
328
|
│ ├─ LLM / keyword tagger: alternative with vocabulary feedback
|
|
@@ -335,39 +344,66 @@ Assemble context within token budget
|
|
|
335
344
|
│ └─ Tag sections: retrieved summaries ordered by tag priority
|
|
336
345
|
│
|
|
337
346
|
▼
|
|
347
|
+
Chain collapse — compress tool-bearing history turns
|
|
348
|
+
│ ├─ Group raw messages into logical turns via group_into_turns()
|
|
349
|
+
│ ├─ Tool chains (assistant tool_use → user tool_result → assistant) → compact stubs
|
|
350
|
+
│ ├─ Stubs contain tool names, truncated previews, and restore refs
|
|
351
|
+
│ ├─ Handles all 4 formats: Anthropic, OpenAI Chat, OpenAI Responses, Gemini
|
|
352
|
+
│ ├─ Deep compaction: drops stubs entirely past configurable age threshold
|
|
353
|
+
│ └─ Non-tool turns outside protected window are dropped (summaries cover them)
|
|
354
|
+
│
|
|
355
|
+
▼
|
|
356
|
+
Store-backed recovery (when client truncates history)
|
|
357
|
+
│ ├─ Detect truncation: payload turns < 70% of stored turns
|
|
358
|
+
�� ├─ Recover chain snapshots from durable store (compact stubs with metadata)
|
|
359
|
+
│ ├─ Recover recent turns from stored turn_messages
|
|
360
|
+
│ └─ Sanitize restored turns: strip thinking blocks, replace media with placeholders
|
|
361
|
+
│
|
|
362
|
+
▼
|
|
338
363
|
Filter conversation history
|
|
339
364
|
│ ├─ Drop turns whose tags don't overlap with inbound tags
|
|
340
365
|
│ ├─ Preserve tool chains atomically (tool_use ↔ tool_result never separated)
|
|
366
|
+
│ ├─ Protected zone intrusion: stub tool results in protected turns 3+ when zone exceeds budget %
|
|
341
367
|
│ ├─ Protect recent turns (always kept regardless of tags)
|
|
342
368
|
│ └─ Temporal queries skip filtering entirely
|
|
343
369
|
│
|
|
344
370
|
▼
|
|
345
|
-
|
|
371
|
+
Budget enforcement — iterative payload reduction
|
|
372
|
+
│ ├─ Scan reducible items: conversation text, tool results, thinking blocks, images
|
|
373
|
+
│ ├─ Cut largest reducible item per iteration until under budget
|
|
374
|
+
│ ├─ Bloat fallback: if VC enrichment exceeds inbound size, fall back to pure passthrough
|
|
375
|
+
│ └─ Upstream trim: final trim to model's actual context window limit
|
|
376
|
+
│
|
|
377
|
+
▼
|
|
378
|
+
Fill pass — replenish context after compression
|
|
379
|
+
│ ├─ Phase 1a: overflow tag summaries (topics that didn't fit during assembly)
|
|
380
|
+
│ ├─ Phase 1b: breadth summaries (sample from remaining tags)
|
|
381
|
+
│ ├─ Phase 2: recent turns from store (newest first)
|
|
382
|
+
│ └─ Target: soft threshold between floor and budget ceiling
|
|
383
|
+
│
|
|
384
|
+
▼
|
|
385
|
+
Inject <virtual-context> block into last user message → forward to LLM
|
|
386
|
+
│ (injected into user messages, not system prompt, for Anthropic cache stability)
|
|
346
387
|
│
|
|
347
388
|
▼
|
|
348
389
|
LLM processes enriched context → produces response
|
|
349
390
|
│
|
|
350
391
|
▼
|
|
351
|
-
Inject
|
|
352
|
-
│ ├─ Streaming: emit final SSE delta with <!-- vc:
|
|
353
|
-
│
|
|
392
|
+
Inject conversation marker into response (proxy mode)
|
|
393
|
+
│ ├─ Streaming: emit final SSE delta with <!-- vc:conversation=UUID -->
|
|
394
|
+
│ ├─ Non-streaming: append marker to last text content block
|
|
395
|
+
│ └─ Skip injection when payload already contains a conversation marker
|
|
354
396
|
│
|
|
355
397
|
▼
|
|
356
|
-
Response tagging
|
|
398
|
+
Response tagging — LLM tags the full user+assistant pair (background thread)
|
|
357
399
|
│ ├─ Context lookback: feed N recent pairs as tagger context for short/ambiguous messages
|
|
358
400
|
│ ├─ Context bleed gate: embedding similarity blocks stale context on topic shifts
|
|
359
401
|
│ ├─ Retry on _general: if tagger returns only _general, retry with expanded context
|
|
360
402
|
│ ├─ Authoritative tags written to TurnTagIndex (vocabulary-building)
|
|
361
|
-
│ ├─ Fact signal extraction: lightweight subject/verb/object triples per turn
|
|
362
403
|
│ ├─ Related tags generated for cross-vocabulary retrieval
|
|
363
404
|
│ └─ Compactor generates related_tags at write time (vocabulary bridging)
|
|
364
405
|
│
|
|
365
406
|
▼
|
|
366
|
-
Fact curation (on inbound, before assembly)
|
|
367
|
-
│ └─ LLM scores retrieved facts for relevance to current query
|
|
368
|
-
│ Low-relevance facts dropped before assembly
|
|
369
|
-
│
|
|
370
|
-
▼
|
|
371
407
|
Check token thresholds (soft 70%, hard 85%)
|
|
372
408
|
│
|
|
373
409
|
▼ (if threshold exceeded)
|
|
@@ -377,14 +413,14 @@ Segment by tag → summarize each segment (concurrent, ThreadPoolExecutor)
|
|
|
377
413
|
│ ├─ Stub segments: media/attachment stubs get passthrough (no LLM), inherit neighbor's tags
|
|
378
414
|
│ ├─ XML-tagged prev_context: structural separation prevents context leak into summaries
|
|
379
415
|
│ ├─ Tags preserved: LLM can ADD refined/related tags but never REMOVE originals
|
|
380
|
-
│ ├─ Fact
|
|
416
|
+
│ ├─ Fact extraction: delete-and-replace per segment, code_mode filters investigatory noise
|
|
381
417
|
│ └─ Related tags written into stored segments for future cross-vocabulary retrieval
|
|
382
418
|
│
|
|
383
419
|
▼
|
|
384
420
|
Compute greedy set cover → build/update per-tag summaries (Layer 2)
|
|
385
421
|
│
|
|
386
422
|
▼
|
|
387
|
-
Persist engine state (TurnTagIndex + compaction watermark → store)
|
|
423
|
+
Persist engine state (TurnTagIndex + compaction watermark → store + Redis)
|
|
388
424
|
```
|
|
389
425
|
|
|
390
426
|
## Key Capabilities
|
|
@@ -451,7 +487,7 @@ For proxy/OpenClaw conversations, session dates come from envelope metadata time
|
|
|
451
487
|
|
|
452
488
|
### Context Awareness Hints
|
|
453
489
|
|
|
454
|
-
After compaction, the LLM loses visibility into what topics have been stored. virtual-context injects a lightweight `<context-topics>` block into the system prompt:
|
|
490
|
+
After compaction, the LLM loses visibility into what topics have been stored. virtual-context injects a lightweight `<context-topics>` block into the last user message (not the system prompt, so the system prompt remains stable and cacheable):
|
|
455
491
|
|
|
456
492
|
```xml
|
|
457
493
|
<context-topics>
|
|
@@ -469,9 +505,9 @@ This costs ~50-200 tokens and enables a natural drill-down loop: the user asks f
|
|
|
469
505
|
|
|
470
506
|
Summaries compress information but inevitably lose specific details. When the user says "I run 5K every morning" at turn 14, a summary might retain "runs regularly" but drop the exact distance and timing. Most memory systems extract facts in a single LLM pass and trust the output directly: raw text goes in, extracted facts come out, and those facts are stored as-is. virtual-context takes a fundamentally different approach with a two-phase pipeline where per-turn signals are treated as hints, not ground truth.
|
|
471
507
|
|
|
472
|
-
**
|
|
508
|
+
**Fact extraction (at compaction).** Facts are extracted from the full turn group when the compactor processes a segment, not per-turn during ingestion. This produces higher-quality facts because the LLM sees the complete conversation flow across multiple turns: what the user asked, how the assistant responded, what was clarified or corrected. The result is a structured `Fact` with full provenance: subject, verb, what (the core assertion), `fact_type` classification (`preference`, `biographical`, `decision`, `plan`, `opinion`, `routine`, `relationship`, `skill`, `medical`, `financial`, `general`), temporal status (active/completed/planned/abandoned/recurring), associated tags, session ID, and source turn numbers. Facts are stored in dedicated SQLite tables with indexes for efficient querying.
|
|
473
509
|
|
|
474
|
-
**
|
|
510
|
+
**Delete-and-replace on re-compaction.** When a segment is re-compacted (e.g., after new turns are added to an existing topic), all facts for that segment are atomically deleted and re-extracted from the full turn group. This prevents fact duplication across re-compaction cycles and ensures facts always reflect the latest understanding. A `code_mode` prompt modifier filters investigatory noise from coding conversations (e.g., "assistant examined file X") and focuses extraction on outcomes: what was built, fixed, changed, or decided.
|
|
475
511
|
|
|
476
512
|
**Why two phases matter.** A single-pass extractor processing "yes, let's go with PostgreSQL" in isolation has no idea what "yes" refers to. It might extract nothing, or hallucinate a fact. virtual-context's response tagger sees the surrounding turns ("Should we use PostgreSQL or MySQL for the user table?") and generates the correct signal. The consolidation pass then verifies it against the full segment before storing a permanent fact. Two chances to get it right, each with progressively more context.
|
|
477
513
|
|
|
@@ -494,6 +530,61 @@ vc_query_facts(fact_type="preference")
|
|
|
494
530
|
|
|
495
531
|
**Semantic fact search.** When structured filters return sparse results, a fallback embedding search matches the query intent against all stored facts' `what` fields by cosine similarity, surfacing relevant facts even when the subject/verb/object decomposition doesn't align.
|
|
496
532
|
|
|
533
|
+
### Chain Collapse and Tool Compression
|
|
534
|
+
|
|
535
|
+
Agent conversations are dominated by tool calls. A coding session with 50 tool rounds might have 900K tokens of tool output but only 60K of actual conversation. Raw tool output (file contents, search results, command output) is high-volume, low-reuse information that crushes the context window.
|
|
536
|
+
|
|
537
|
+
virtual-context collapses entire tool chains into compact stubs:
|
|
538
|
+
|
|
539
|
+
```
|
|
540
|
+
Before (3 messages, ~18K tokens):
|
|
541
|
+
assistant: [tool_use: Read file.py]
|
|
542
|
+
user: [tool_result: <full 500-line file contents>]
|
|
543
|
+
assistant: "The file has a bug on line 42..."
|
|
544
|
+
|
|
545
|
+
After (2 messages, ~200 tokens):
|
|
546
|
+
user: [compacted turn — tool activity: Read(file.py) — vc_restore_tool can recover full content]
|
|
547
|
+
assistant: "The file has a bug on line 42..."
|
|
548
|
+
```
|
|
549
|
+
|
|
550
|
+
Chain collapse handles all four provider formats (Anthropic `tool_use`/`tool_result`, OpenAI Chat `tool_calls`/`role:tool`, OpenAI Responses `function_call`/`function_call_output`, Gemini `functionCall`/`functionResponse`). Full raw tool output is stored durably with content-addressed refs and recoverable via `vc_restore_tool`.
|
|
551
|
+
|
|
552
|
+
**Deep compaction** drops stubs entirely past a configurable age threshold (`deep_compaction_ratio`). A stub from turn 5 in a 200-turn conversation adds no value; the segment summaries already cover that content.
|
|
553
|
+
|
|
554
|
+
### Media Compression
|
|
555
|
+
|
|
556
|
+
Base64 images in API payloads are enormous: a single screenshot is 300-500KB of base64, consuming ~100K tokens. Multimodal conversations with dozens of images can reach 10MB+.
|
|
557
|
+
|
|
558
|
+
virtual-context compresses images on first sight, before any pipeline processing:
|
|
559
|
+
|
|
560
|
+
- Detect base64 images across all 4 formats
|
|
561
|
+
- Compress to JPEG at configurable quality, store originals to disk
|
|
562
|
+
- Replace in-flight payload with compressed version
|
|
563
|
+
- A 391KB screenshot → ~40KB compressed, saving ~88K tokens per image
|
|
564
|
+
- Recovery via `vc_restore_tool` returns the original uncompressed content
|
|
565
|
+
|
|
566
|
+
Media compression runs on both passthrough and active paths, so even conversations that haven't triggered compaction benefit.
|
|
567
|
+
|
|
568
|
+
### Fill Pass
|
|
569
|
+
|
|
570
|
+
After chain collapse and budget enforcement compress the payload, the context window may have significant unused capacity. The fill pass replenishes it with high-value content from the store:
|
|
571
|
+
|
|
572
|
+
- **Phase 1a**: Overflow tag summaries — topics that didn't fit during initial assembly
|
|
573
|
+
- **Phase 1b**: Breadth summaries — sample from remaining tags for broader coverage
|
|
574
|
+
- **Phase 2**: Recent turns from store — newest first, filling toward a soft target threshold
|
|
575
|
+
|
|
576
|
+
The fill pass runs after all compression, so it never fights against budget enforcement. When clients truncate history (detected via `payload_turns < store_turns * 0.70`), the fill pass works with store-backed recovery to restore the most valuable context.
|
|
577
|
+
|
|
578
|
+
### Store-Backed Recovery
|
|
579
|
+
|
|
580
|
+
Clients (Claude Code, OpenClaw) sometimes truncate conversation history to manage their own context windows. When this happens, virtual-context detects the truncation and recovers from its durable store:
|
|
581
|
+
|
|
582
|
+
- **Chain snapshot recovery**: Restore compact tool chain stubs from stored `chain_snapshots`
|
|
583
|
+
- **Turn recovery**: Restore recent raw turns from stored `turn_messages`
|
|
584
|
+
- **Sanitization**: Strip thinking blocks, replace media with passive placeholders, remove orphaned tool scaffolding
|
|
585
|
+
|
|
586
|
+
Recovery is transparent to the client. The payload that reaches the LLM contains the recovered context as if it had never been truncated.
|
|
587
|
+
|
|
497
588
|
### Virtual Memory Paging
|
|
498
589
|
|
|
499
590
|
RAG retrieves content and appends it to the context window. It never frees space from what's already there. When a 100k document needs to enter a 120k window that already has 60k of conversation history, RAG has three options: truncate (lossy), error (useless), or chunk (every chunking approach either costs extra user turns, loses cross-chunk coherence, or both). Nobody touches the existing 60k. It sits there, potentially full of stale context from 30 turns ago that nobody needs anymore.
|
|
@@ -696,19 +787,23 @@ virtual-context chat --replay vc-session.json
|
|
|
696
787
|
|
|
697
788
|
### Proxy Deep Dive
|
|
698
789
|
|
|
699
|
-
**
|
|
790
|
+
**Conversation continuity.** The proxy injects an invisible `<!-- vc:conversation=UUID -->` marker into every assistant response. On subsequent requests, the proxy extracts the first marker in the conversation history, routes to the correct conversation, and strips markers before forwarding upstream. Stable conversation identity is derived from a format-specific hash of the system prompt and early messages, so the same client session always routes to the same conversation even across restarts. Valid UUID markers in the payload are accepted as-is without re-hashing.
|
|
791
|
+
|
|
792
|
+
**Redis session cache.** A write-through Redis cache persists conversation history and engine state across container restarts. On startup, conversations are restored from Redis with full history, eliminating cold-start re-ingestion. Degraded mode: if Redis is unavailable, the proxy falls back to store-only persistence with no interruption.
|
|
700
793
|
|
|
701
794
|
**Conversation-scoped retrieval.** All store retrieval methods are scoped by `conversation_id`. Multiple conversations sharing the same SQLite database are fully isolated; a new conversation never gets context from another conversation's segments.
|
|
702
795
|
|
|
703
|
-
**
|
|
796
|
+
**Pipeline suppression.** When a conversation has no compacted data, the pipeline is suppressed; requests pass through as-is. Once the first compaction runs, the pipeline activates automatically.
|
|
704
797
|
|
|
705
798
|
**History ingestion.** On the first request, the proxy extracts user+assistant pairs from the client's existing conversation history and tags each to bootstrap the TurnTagIndex. No cold-start period.
|
|
706
799
|
|
|
707
|
-
**
|
|
800
|
+
**Four-format support.** Auto-detects Anthropic, OpenAI Chat, OpenAI Responses, and Gemini request formats. Every pipeline stage (chain collapse, media compression, budget enforcement, context injection, stub generation, token counting) is format-aware through the `PayloadFormat` abstraction. A single proxy instance handles all formats on one port.
|
|
801
|
+
|
|
802
|
+
**Image-aware token counting.** Base64 images are counted using the Anthropic formula `(width × height) / 750` after scaling to 1568px max dimension, instead of tokenizing the raw base64 string. This prevents massive over-counting (a 10MB image payload is ~107K tokens, not 7M) and ensures budget enforcement and tier checks operate on accurate numbers.
|
|
708
803
|
|
|
709
804
|
**Streaming with zero added latency.** SSE streams are forwarded byte-for-byte. Text deltas are accumulated in the background for response tagging.
|
|
710
805
|
|
|
711
|
-
**Error-resilient.** If the engine fails, the request is forwarded to upstream unmodified. The proxy never blocks your LLM calls.
|
|
806
|
+
**Error-resilient.** If the engine fails, the request is forwarded to upstream unmodified. The proxy never blocks your LLM calls. If VC enrichment produces a larger payload than the original, the bloat fallback reverts to a pure passthrough with the original client body.
|
|
712
807
|
|
|
713
808
|
**Envelope stripping + metadata extraction.** Strips client metadata while extracting sender identity and timestamps from labeled JSON blocks. Group chat participants appear as "Sania" and "Yur" instead of generic "User". Original message timestamps give segments accurate chronological ordering.
|
|
714
809
|
|
|
@@ -766,8 +861,13 @@ Plugin for OpenClaw agents using lifecycle hooks for sync retrieval (`message.pr
|
|
|
766
861
|
| **ProxyServer** | `proxy/server.py` | HTTP proxy factory (`create_app`), delegates to state/registry/handlers |
|
|
767
862
|
| **ProxyState** | `proxy/state.py` | Session state machine: ingestion, tagging, compaction lifecycle |
|
|
768
863
|
| **SessionRegistry** | `proxy/registry.py` | Multi-session routing with fingerprint matching |
|
|
864
|
+
| **MessageFilter** | `proxy/message_filter.py` | Chain collapse, turn grouping, budget enforcement, fill pass, upstream trim |
|
|
865
|
+
| **TokenCounter** | `token_counter.py` | Image-aware token counting (Anthropic formula for images, tiktoken/estimate for text) |
|
|
866
|
+
| **MediaCompressor** | `proxy/media.py` | Base64 image compression, disk storage, intrusion detection |
|
|
867
|
+
| **ConversationIdentity** | `conversation_identity.py` | Stable per-format conversation ID hashing, marker extraction |
|
|
769
868
|
| **ProxyHandlers** | `proxy/handlers.py` | Streaming/non-streaming/passthrough HTTP request handlers |
|
|
770
869
|
| **MultiInstance** | `proxy/multi.py` | Multi-instance launcher: N uvicorn listeners, shared or per-port engine/store |
|
|
870
|
+
| **RedisSessionCache** | `proxy/redis_cache.py` | Write-through history cache for lossless restarts |
|
|
771
871
|
| **ProxyDashboard** | `proxy/dashboard.py` | Live SSE dashboard with request grid, turn inspector, session stats (auth-gated mutations) |
|
|
772
872
|
| **ProxyMetrics** | `proxy/metrics.py` | Thread-safe event collector with bounded deque + request capture ring buffer |
|
|
773
873
|
|
|
@@ -816,7 +916,7 @@ Both providers reuse a persistent `httpx.Client` across calls (connection poolin
|
|
|
816
916
|
|
|
817
917
|
**Tag preservation.** During compaction, the LLM can add refined tags but never remove original ones. A segment tagged `[ux, recipes, frontend]` stays tagged with all three even after summarization, ensuring cross-topic retrieval always works.
|
|
818
918
|
|
|
819
|
-
**Tool chain
|
|
919
|
+
**Tool chain collapse.** Historical tool chains (assistant `tool_use` → user `tool_result` → assistant response, and their OpenAI/Gemini equivalents) are collapsed into compact stubs containing tool names, truncated previews, and content-addressed restore refs. This is the single largest compression lever: a 937K-token payload with 52 tool chains collapses to ~65K. Full tool output is stored durably and recoverable via `vc_restore_tool`. Deep compaction drops stubs entirely past a configurable age threshold. The history filter preserves API-required message dependencies atomically; tool chains are never partially broken.
|
|
820
920
|
|
|
821
921
|
**The virtual memory analogy is literal, not metaphorical.** Every component in VC maps to a systems-level equivalent:
|
|
822
922
|
|
|
@@ -133,7 +133,7 @@ proxy:
|
|
|
133
133
|
**Daemon mode:** run as a background service:
|
|
134
134
|
|
|
135
135
|
```bash
|
|
136
|
-
# Creates config
|
|
136
|
+
# Creates ~/.virtualcontext/ with config + data, installs + starts daemon
|
|
137
137
|
virtual-context daemon install --upstream https://api.anthropic.com
|
|
138
138
|
|
|
139
139
|
# Or: guided interactive setup with daemon
|
|
@@ -235,6 +235,7 @@ Exposes virtual-context as an MCP server for integration with Claude Desktop, Cu
|
|
|
235
235
|
| Tool | `collapse_topic` | Collapse a topic back to summary or none |
|
|
236
236
|
| Tool | `find_quote` | Full-text search across all stored conversation text |
|
|
237
237
|
| Tool | `query_facts` | Structured fact lookup with subject/verb/object/status filters |
|
|
238
|
+
| Tool | `restore_tool` | Recover full content from compacted chain stubs or compressed media |
|
|
238
239
|
| Resource | `virtualcontext://domains` | List all tags |
|
|
239
240
|
| Resource | `virtualcontext://domains/{tag}` | Summaries for a specific tag |
|
|
240
241
|
| Prompt | `recall` | Suggest context retrieval for a topic |
|
|
@@ -246,11 +247,12 @@ Exposes virtual-context as an MCP server for integration with Claude Desktop, Cu
|
|
|
246
247
|
User message arrives
|
|
247
248
|
│
|
|
248
249
|
▼
|
|
249
|
-
|
|
250
|
-
│ ├─ Extract
|
|
251
|
-
│ ├─ Route to existing
|
|
252
|
-
│ ├─ No marker? →
|
|
253
|
-
│
|
|
250
|
+
Conversation routing (proxy mode)
|
|
251
|
+
│ ├─ Extract conversation ID from <!-- vc:conversation=UUID --> markers
|
|
252
|
+
│ ├─ Route to existing conversation or load persisted state from store + Redis
|
|
253
|
+
│ ├─ No marker? → derive stable ID from system prompt hash + format
|
|
254
|
+
│ ├─ Redis session cache: lossless restart, write-through history persistence
|
|
255
|
+
│ └─ Strip conversation markers before forwarding upstream
|
|
254
256
|
│
|
|
255
257
|
▼
|
|
256
258
|
Strip client envelope + extract metadata
|
|
@@ -260,13 +262,20 @@ Strip client envelope + extract metadata
|
|
|
260
262
|
│ └─ Metadata preserved on Message.metadata for downstream use
|
|
261
263
|
│
|
|
262
264
|
▼
|
|
265
|
+
Media compression (all paths — passthrough and active)
|
|
266
|
+
│ ├─ Detect base64 images across all 4 formats (Anthropic, OpenAI Chat, Responses, Gemini)
|
|
267
|
+
│ ├─ Compress to JPEG, store originals to disk for later recovery
|
|
268
|
+
│ ├─ Image token counting uses Anthropic formula: (width × height) / 750
|
|
269
|
+
│ └─ A 391KB screenshot → ~40KB compressed, saving ~88k tokens
|
|
270
|
+
│
|
|
271
|
+
▼
|
|
263
272
|
History ingestion (first request only)
|
|
264
273
|
│ ├─ Extract and tag all prior user+assistant pairs → bootstrap TurnTagIndex
|
|
265
274
|
│ ├─ Stub detection: media attachments/image placeholders get _stub tag (skip LLM tagger)
|
|
266
275
|
│ └─ Conversation-scoped: each conversation's index is independent
|
|
267
276
|
│
|
|
268
277
|
▼
|
|
269
|
-
Inbound tagging
|
|
278
|
+
Inbound tagging — identify what this message is about
|
|
270
279
|
│ ├─ Embedding tagger (recommended): cosine similarity against existing tag vocabulary
|
|
271
280
|
│ │ (closed-set, deterministic, can't hallucinate novel tags)
|
|
272
281
|
│ ├─ LLM / keyword tagger: alternative with vocabulary feedback
|
|
@@ -288,39 +297,66 @@ Assemble context within token budget
|
|
|
288
297
|
│ └─ Tag sections: retrieved summaries ordered by tag priority
|
|
289
298
|
│
|
|
290
299
|
▼
|
|
300
|
+
Chain collapse — compress tool-bearing history turns
|
|
301
|
+
│ ├─ Group raw messages into logical turns via group_into_turns()
|
|
302
|
+
│ ├─ Tool chains (assistant tool_use → user tool_result → assistant) → compact stubs
|
|
303
|
+
│ ├─ Stubs contain tool names, truncated previews, and restore refs
|
|
304
|
+
│ ├─ Handles all 4 formats: Anthropic, OpenAI Chat, OpenAI Responses, Gemini
|
|
305
|
+
│ ├─ Deep compaction: drops stubs entirely past configurable age threshold
|
|
306
|
+
│ └─ Non-tool turns outside protected window are dropped (summaries cover them)
|
|
307
|
+
│
|
|
308
|
+
▼
|
|
309
|
+
Store-backed recovery (when client truncates history)
|
|
310
|
+
│ ├─ Detect truncation: payload turns < 70% of stored turns
|
|
311
|
+
�� ├─ Recover chain snapshots from durable store (compact stubs with metadata)
|
|
312
|
+
│ ├─ Recover recent turns from stored turn_messages
|
|
313
|
+
│ └─ Sanitize restored turns: strip thinking blocks, replace media with placeholders
|
|
314
|
+
│
|
|
315
|
+
▼
|
|
291
316
|
Filter conversation history
|
|
292
317
|
│ ├─ Drop turns whose tags don't overlap with inbound tags
|
|
293
318
|
│ ├─ Preserve tool chains atomically (tool_use ↔ tool_result never separated)
|
|
319
|
+
│ ├─ Protected zone intrusion: stub tool results in protected turns 3+ when zone exceeds budget %
|
|
294
320
|
│ ├─ Protect recent turns (always kept regardless of tags)
|
|
295
321
|
│ └─ Temporal queries skip filtering entirely
|
|
296
322
|
│
|
|
297
323
|
▼
|
|
298
|
-
|
|
324
|
+
Budget enforcement — iterative payload reduction
|
|
325
|
+
│ ├─ Scan reducible items: conversation text, tool results, thinking blocks, images
|
|
326
|
+
│ ├─ Cut largest reducible item per iteration until under budget
|
|
327
|
+
│ ├─ Bloat fallback: if VC enrichment exceeds inbound size, fall back to pure passthrough
|
|
328
|
+
│ └─ Upstream trim: final trim to model's actual context window limit
|
|
329
|
+
│
|
|
330
|
+
▼
|
|
331
|
+
Fill pass — replenish context after compression
|
|
332
|
+
│ ├─ Phase 1a: overflow tag summaries (topics that didn't fit during assembly)
|
|
333
|
+
│ ├─ Phase 1b: breadth summaries (sample from remaining tags)
|
|
334
|
+
│ ├─ Phase 2: recent turns from store (newest first)
|
|
335
|
+
│ └─ Target: soft threshold between floor and budget ceiling
|
|
336
|
+
│
|
|
337
|
+
▼
|
|
338
|
+
Inject <virtual-context> block into last user message → forward to LLM
|
|
339
|
+
│ (injected into user messages, not system prompt, for Anthropic cache stability)
|
|
299
340
|
│
|
|
300
341
|
▼
|
|
301
342
|
LLM processes enriched context → produces response
|
|
302
343
|
│
|
|
303
344
|
▼
|
|
304
|
-
Inject
|
|
305
|
-
│ ├─ Streaming: emit final SSE delta with <!-- vc:
|
|
306
|
-
│
|
|
345
|
+
Inject conversation marker into response (proxy mode)
|
|
346
|
+
│ ├─ Streaming: emit final SSE delta with <!-- vc:conversation=UUID -->
|
|
347
|
+
│ ├─ Non-streaming: append marker to last text content block
|
|
348
|
+
│ └─ Skip injection when payload already contains a conversation marker
|
|
307
349
|
│
|
|
308
350
|
▼
|
|
309
|
-
Response tagging
|
|
351
|
+
Response tagging — LLM tags the full user+assistant pair (background thread)
|
|
310
352
|
│ ├─ Context lookback: feed N recent pairs as tagger context for short/ambiguous messages
|
|
311
353
|
│ ├─ Context bleed gate: embedding similarity blocks stale context on topic shifts
|
|
312
354
|
│ ├─ Retry on _general: if tagger returns only _general, retry with expanded context
|
|
313
355
|
│ ├─ Authoritative tags written to TurnTagIndex (vocabulary-building)
|
|
314
|
-
│ ├─ Fact signal extraction: lightweight subject/verb/object triples per turn
|
|
315
356
|
│ ├─ Related tags generated for cross-vocabulary retrieval
|
|
316
357
|
│ └─ Compactor generates related_tags at write time (vocabulary bridging)
|
|
317
358
|
│
|
|
318
359
|
▼
|
|
319
|
-
Fact curation (on inbound, before assembly)
|
|
320
|
-
│ └─ LLM scores retrieved facts for relevance to current query
|
|
321
|
-
│ Low-relevance facts dropped before assembly
|
|
322
|
-
│
|
|
323
|
-
▼
|
|
324
360
|
Check token thresholds (soft 70%, hard 85%)
|
|
325
361
|
│
|
|
326
362
|
▼ (if threshold exceeded)
|
|
@@ -330,14 +366,14 @@ Segment by tag → summarize each segment (concurrent, ThreadPoolExecutor)
|
|
|
330
366
|
│ ├─ Stub segments: media/attachment stubs get passthrough (no LLM), inherit neighbor's tags
|
|
331
367
|
│ ├─ XML-tagged prev_context: structural separation prevents context leak into summaries
|
|
332
368
|
│ ├─ Tags preserved: LLM can ADD refined/related tags but never REMOVE originals
|
|
333
|
-
│ ├─ Fact
|
|
369
|
+
│ ├─ Fact extraction: delete-and-replace per segment, code_mode filters investigatory noise
|
|
334
370
|
│ └─ Related tags written into stored segments for future cross-vocabulary retrieval
|
|
335
371
|
│
|
|
336
372
|
▼
|
|
337
373
|
Compute greedy set cover → build/update per-tag summaries (Layer 2)
|
|
338
374
|
│
|
|
339
375
|
▼
|
|
340
|
-
Persist engine state (TurnTagIndex + compaction watermark → store)
|
|
376
|
+
Persist engine state (TurnTagIndex + compaction watermark → store + Redis)
|
|
341
377
|
```
|
|
342
378
|
|
|
343
379
|
## Key Capabilities
|
|
@@ -404,7 +440,7 @@ For proxy/OpenClaw conversations, session dates come from envelope metadata time
|
|
|
404
440
|
|
|
405
441
|
### Context Awareness Hints
|
|
406
442
|
|
|
407
|
-
After compaction, the LLM loses visibility into what topics have been stored. virtual-context injects a lightweight `<context-topics>` block into the system prompt:
|
|
443
|
+
After compaction, the LLM loses visibility into what topics have been stored. virtual-context injects a lightweight `<context-topics>` block into the last user message (not the system prompt, so the system prompt remains stable and cacheable):
|
|
408
444
|
|
|
409
445
|
```xml
|
|
410
446
|
<context-topics>
|
|
@@ -422,9 +458,9 @@ This costs ~50-200 tokens and enables a natural drill-down loop: the user asks f
|
|
|
422
458
|
|
|
423
459
|
Summaries compress information but inevitably lose specific details. When the user says "I run 5K every morning" at turn 14, a summary might retain "runs regularly" but drop the exact distance and timing. Most memory systems extract facts in a single LLM pass and trust the output directly: raw text goes in, extracted facts come out, and those facts are stored as-is. virtual-context takes a fundamentally different approach with a two-phase pipeline where per-turn signals are treated as hints, not ground truth.
|
|
424
460
|
|
|
425
|
-
**
|
|
461
|
+
**Fact extraction (at compaction).** Facts are extracted from the full turn group when the compactor processes a segment, not per-turn during ingestion. This produces higher-quality facts because the LLM sees the complete conversation flow across multiple turns: what the user asked, how the assistant responded, what was clarified or corrected. The result is a structured `Fact` with full provenance: subject, verb, what (the core assertion), `fact_type` classification (`preference`, `biographical`, `decision`, `plan`, `opinion`, `routine`, `relationship`, `skill`, `medical`, `financial`, `general`), temporal status (active/completed/planned/abandoned/recurring), associated tags, session ID, and source turn numbers. Facts are stored in dedicated SQLite tables with indexes for efficient querying.
|
|
426
462
|
|
|
427
|
-
**
|
|
463
|
+
**Delete-and-replace on re-compaction.** When a segment is re-compacted (e.g., after new turns are added to an existing topic), all facts for that segment are atomically deleted and re-extracted from the full turn group. This prevents fact duplication across re-compaction cycles and ensures facts always reflect the latest understanding. A `code_mode` prompt modifier filters investigatory noise from coding conversations (e.g., "assistant examined file X") and focuses extraction on outcomes: what was built, fixed, changed, or decided.
|
|
428
464
|
|
|
429
465
|
**Why two phases matter.** A single-pass extractor processing "yes, let's go with PostgreSQL" in isolation has no idea what "yes" refers to. It might extract nothing, or hallucinate a fact. virtual-context's response tagger sees the surrounding turns ("Should we use PostgreSQL or MySQL for the user table?") and generates the correct signal. The consolidation pass then verifies it against the full segment before storing a permanent fact. Two chances to get it right, each with progressively more context.
|
|
430
466
|
|
|
@@ -447,6 +483,61 @@ vc_query_facts(fact_type="preference")
|
|
|
447
483
|
|
|
448
484
|
**Semantic fact search.** When structured filters return sparse results, a fallback embedding search matches the query intent against all stored facts' `what` fields by cosine similarity, surfacing relevant facts even when the subject/verb/object decomposition doesn't align.
|
|
449
485
|
|
|
486
|
+
### Chain Collapse and Tool Compression
|
|
487
|
+
|
|
488
|
+
Agent conversations are dominated by tool calls. A coding session with 50 tool rounds might have 900K tokens of tool output but only 60K of actual conversation. Raw tool output (file contents, search results, command output) is high-volume, low-reuse information that crushes the context window.
|
|
489
|
+
|
|
490
|
+
virtual-context collapses entire tool chains into compact stubs:
|
|
491
|
+
|
|
492
|
+
```
|
|
493
|
+
Before (3 messages, ~18K tokens):
|
|
494
|
+
assistant: [tool_use: Read file.py]
|
|
495
|
+
user: [tool_result: <full 500-line file contents>]
|
|
496
|
+
assistant: "The file has a bug on line 42..."
|
|
497
|
+
|
|
498
|
+
After (2 messages, ~200 tokens):
|
|
499
|
+
user: [compacted turn — tool activity: Read(file.py) — vc_restore_tool can recover full content]
|
|
500
|
+
assistant: "The file has a bug on line 42..."
|
|
501
|
+
```
|
|
502
|
+
|
|
503
|
+
Chain collapse handles all four provider formats (Anthropic `tool_use`/`tool_result`, OpenAI Chat `tool_calls`/`role:tool`, OpenAI Responses `function_call`/`function_call_output`, Gemini `functionCall`/`functionResponse`). Full raw tool output is stored durably with content-addressed refs and recoverable via `vc_restore_tool`.
|
|
504
|
+
|
|
505
|
+
**Deep compaction** drops stubs entirely past a configurable age threshold (`deep_compaction_ratio`). A stub from turn 5 in a 200-turn conversation adds no value; the segment summaries already cover that content.
|
|
506
|
+
|
|
507
|
+
### Media Compression
|
|
508
|
+
|
|
509
|
+
Base64 images in API payloads are enormous: a single screenshot is 300-500KB of base64, consuming ~100K tokens. Multimodal conversations with dozens of images can reach 10MB+.
|
|
510
|
+
|
|
511
|
+
virtual-context compresses images on first sight, before any pipeline processing:
|
|
512
|
+
|
|
513
|
+
- Detect base64 images across all 4 formats
|
|
514
|
+
- Compress to JPEG at configurable quality, store originals to disk
|
|
515
|
+
- Replace in-flight payload with compressed version
|
|
516
|
+
- A 391KB screenshot → ~40KB compressed, saving ~88K tokens per image
|
|
517
|
+
- Recovery via `vc_restore_tool` returns the original uncompressed content
|
|
518
|
+
|
|
519
|
+
Media compression runs on both passthrough and active paths, so even conversations that haven't triggered compaction benefit.
|
|
520
|
+
|
|
521
|
+
### Fill Pass
|
|
522
|
+
|
|
523
|
+
After chain collapse and budget enforcement compress the payload, the context window may have significant unused capacity. The fill pass replenishes it with high-value content from the store:
|
|
524
|
+
|
|
525
|
+
- **Phase 1a**: Overflow tag summaries — topics that didn't fit during initial assembly
|
|
526
|
+
- **Phase 1b**: Breadth summaries — sample from remaining tags for broader coverage
|
|
527
|
+
- **Phase 2**: Recent turns from store — newest first, filling toward a soft target threshold
|
|
528
|
+
|
|
529
|
+
The fill pass runs after all compression, so it never fights against budget enforcement. When clients truncate history (detected via `payload_turns < store_turns * 0.70`), the fill pass works with store-backed recovery to restore the most valuable context.
|
|
530
|
+
|
|
531
|
+
### Store-Backed Recovery
|
|
532
|
+
|
|
533
|
+
Clients (Claude Code, OpenClaw) sometimes truncate conversation history to manage their own context windows. When this happens, virtual-context detects the truncation and recovers from its durable store:
|
|
534
|
+
|
|
535
|
+
- **Chain snapshot recovery**: Restore compact tool chain stubs from stored `chain_snapshots`
|
|
536
|
+
- **Turn recovery**: Restore recent raw turns from stored `turn_messages`
|
|
537
|
+
- **Sanitization**: Strip thinking blocks, replace media with passive placeholders, remove orphaned tool scaffolding
|
|
538
|
+
|
|
539
|
+
Recovery is transparent to the client. The payload that reaches the LLM contains the recovered context as if it had never been truncated.
|
|
540
|
+
|
|
450
541
|
### Virtual Memory Paging
|
|
451
542
|
|
|
452
543
|
RAG retrieves content and appends it to the context window. It never frees space from what's already there. When a 100k document needs to enter a 120k window that already has 60k of conversation history, RAG has three options: truncate (lossy), error (useless), or chunk (every chunking approach either costs extra user turns, loses cross-chunk coherence, or both). Nobody touches the existing 60k. It sits there, potentially full of stale context from 30 turns ago that nobody needs anymore.
|
|
@@ -649,19 +740,23 @@ virtual-context chat --replay vc-session.json
|
|
|
649
740
|
|
|
650
741
|
### Proxy Deep Dive
|
|
651
742
|
|
|
652
|
-
**
|
|
743
|
+
**Conversation continuity.** The proxy injects an invisible `<!-- vc:conversation=UUID -->` marker into every assistant response. On subsequent requests, the proxy extracts the first marker in the conversation history, routes to the correct conversation, and strips markers before forwarding upstream. Stable conversation identity is derived from a format-specific hash of the system prompt and early messages, so the same client session always routes to the same conversation even across restarts. Valid UUID markers in the payload are accepted as-is without re-hashing.
|
|
744
|
+
|
|
745
|
+
**Redis session cache.** A write-through Redis cache persists conversation history and engine state across container restarts. On startup, conversations are restored from Redis with full history, eliminating cold-start re-ingestion. Degraded mode: if Redis is unavailable, the proxy falls back to store-only persistence with no interruption.
|
|
653
746
|
|
|
654
747
|
**Conversation-scoped retrieval.** All store retrieval methods are scoped by `conversation_id`. Multiple conversations sharing the same SQLite database are fully isolated; a new conversation never gets context from another conversation's segments.
|
|
655
748
|
|
|
656
|
-
**
|
|
749
|
+
**Pipeline suppression.** When a conversation has no compacted data, the pipeline is suppressed; requests pass through as-is. Once the first compaction runs, the pipeline activates automatically.
|
|
657
750
|
|
|
658
751
|
**History ingestion.** On the first request, the proxy extracts user+assistant pairs from the client's existing conversation history and tags each to bootstrap the TurnTagIndex. No cold-start period.
|
|
659
752
|
|
|
660
|
-
**
|
|
753
|
+
**Four-format support.** Auto-detects Anthropic, OpenAI Chat, OpenAI Responses, and Gemini request formats. Every pipeline stage (chain collapse, media compression, budget enforcement, context injection, stub generation, token counting) is format-aware through the `PayloadFormat` abstraction. A single proxy instance handles all formats on one port.
|
|
754
|
+
|
|
755
|
+
**Image-aware token counting.** Base64 images are counted using the Anthropic formula `(width × height) / 750` after scaling to 1568px max dimension, instead of tokenizing the raw base64 string. This prevents massive over-counting (a 10MB image payload is ~107K tokens, not 7M) and ensures budget enforcement and tier checks operate on accurate numbers.
|
|
661
756
|
|
|
662
757
|
**Streaming with zero added latency.** SSE streams are forwarded byte-for-byte. Text deltas are accumulated in the background for response tagging.
|
|
663
758
|
|
|
664
|
-
**Error-resilient.** If the engine fails, the request is forwarded to upstream unmodified. The proxy never blocks your LLM calls.
|
|
759
|
+
**Error-resilient.** If the engine fails, the request is forwarded to upstream unmodified. The proxy never blocks your LLM calls. If VC enrichment produces a larger payload than the original, the bloat fallback reverts to a pure passthrough with the original client body.
|
|
665
760
|
|
|
666
761
|
**Envelope stripping + metadata extraction.** Strips client metadata while extracting sender identity and timestamps from labeled JSON blocks. Group chat participants appear as "Sania" and "Yur" instead of generic "User". Original message timestamps give segments accurate chronological ordering.
|
|
667
762
|
|
|
@@ -719,8 +814,13 @@ Plugin for OpenClaw agents using lifecycle hooks for sync retrieval (`message.pr
|
|
|
719
814
|
| **ProxyServer** | `proxy/server.py` | HTTP proxy factory (`create_app`), delegates to state/registry/handlers |
|
|
720
815
|
| **ProxyState** | `proxy/state.py` | Session state machine: ingestion, tagging, compaction lifecycle |
|
|
721
816
|
| **SessionRegistry** | `proxy/registry.py` | Multi-session routing with fingerprint matching |
|
|
817
|
+
| **MessageFilter** | `proxy/message_filter.py` | Chain collapse, turn grouping, budget enforcement, fill pass, upstream trim |
|
|
818
|
+
| **TokenCounter** | `token_counter.py` | Image-aware token counting (Anthropic formula for images, tiktoken/estimate for text) |
|
|
819
|
+
| **MediaCompressor** | `proxy/media.py` | Base64 image compression, disk storage, intrusion detection |
|
|
820
|
+
| **ConversationIdentity** | `conversation_identity.py` | Stable per-format conversation ID hashing, marker extraction |
|
|
722
821
|
| **ProxyHandlers** | `proxy/handlers.py` | Streaming/non-streaming/passthrough HTTP request handlers |
|
|
723
822
|
| **MultiInstance** | `proxy/multi.py` | Multi-instance launcher: N uvicorn listeners, shared or per-port engine/store |
|
|
823
|
+
| **RedisSessionCache** | `proxy/redis_cache.py` | Write-through history cache for lossless restarts |
|
|
724
824
|
| **ProxyDashboard** | `proxy/dashboard.py` | Live SSE dashboard with request grid, turn inspector, session stats (auth-gated mutations) |
|
|
725
825
|
| **ProxyMetrics** | `proxy/metrics.py` | Thread-safe event collector with bounded deque + request capture ring buffer |
|
|
726
826
|
|
|
@@ -769,7 +869,7 @@ Both providers reuse a persistent `httpx.Client` across calls (connection poolin
|
|
|
769
869
|
|
|
770
870
|
**Tag preservation.** During compaction, the LLM can add refined tags but never remove original ones. A segment tagged `[ux, recipes, frontend]` stays tagged with all three even after summarization, ensuring cross-topic retrieval always works.
|
|
771
871
|
|
|
772
|
-
**Tool chain
|
|
872
|
+
**Tool chain collapse.** Historical tool chains (assistant `tool_use` → user `tool_result` → assistant response, and their OpenAI/Gemini equivalents) are collapsed into compact stubs containing tool names, truncated previews, and content-addressed restore refs. This is the single largest compression lever: a 937K-token payload with 52 tool chains collapses to ~65K. Full tool output is stored durably and recoverable via `vc_restore_tool`. Deep compaction drops stubs entirely past a configurable age threshold. The history filter preserves API-required message dependencies atomically; tool chains are never partially broken.
|
|
773
873
|
|
|
774
874
|
**The virtual memory analogy is literal, not metaphorical.** Every component in VC maps to a systems-level equivalent:
|
|
775
875
|
|