kalmanflow 0.1.0a1__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (36) hide show
  1. kalmanflow-0.1.0a1/Documentation/API_REFERENCE.md +107 -0
  2. kalmanflow-0.1.0a1/Documentation/ARCHITECTURE.md +83 -0
  3. kalmanflow-0.1.0a1/Documentation/BAYESIAN_NOISE_TUNING.md +217 -0
  4. kalmanflow-0.1.0a1/Documentation/CONFIGURATION.md +90 -0
  5. kalmanflow-0.1.0a1/Documentation/INFLOW_MODEL_BEHAVIOR.md +84 -0
  6. kalmanflow-0.1.0a1/Documentation/ONLINE_INFLOW_PIPELINE.md +63 -0
  7. kalmanflow-0.1.0a1/Documentation/VALIDATION.md +71 -0
  8. kalmanflow-0.1.0a1/LICENSE +202 -0
  9. kalmanflow-0.1.0a1/NOTICE +1 -0
  10. kalmanflow-0.1.0a1/PKG-INFO +151 -0
  11. kalmanflow-0.1.0a1/README.md +125 -0
  12. kalmanflow-0.1.0a1/pyproject.toml +79 -0
  13. kalmanflow-0.1.0a1/src/kalmanflow/__init__.py +91 -0
  14. kalmanflow-0.1.0a1/src/kalmanflow/_reservoir_checkpoint.py +356 -0
  15. kalmanflow-0.1.0a1/src/kalmanflow/_validation.py +79 -0
  16. kalmanflow-0.1.0a1/src/kalmanflow/bayesian_tuning/__init__.py +28 -0
  17. kalmanflow-0.1.0a1/src/kalmanflow/bayesian_tuning/_candidate.py +202 -0
  18. kalmanflow-0.1.0a1/src/kalmanflow/bayesian_tuning/_diagnostics.py +440 -0
  19. kalmanflow-0.1.0a1/src/kalmanflow/bayesian_tuning/_optimization.py +210 -0
  20. kalmanflow-0.1.0a1/src/kalmanflow/bayesian_tuning/_preparation.py +208 -0
  21. kalmanflow-0.1.0a1/src/kalmanflow/bayesian_tuning/_proxy.py +172 -0
  22. kalmanflow-0.1.0a1/src/kalmanflow/bayesian_tuning/_types.py +354 -0
  23. kalmanflow-0.1.0a1/src/kalmanflow/bayesian_tuning/_workflows.py +640 -0
  24. kalmanflow-0.1.0a1/src/kalmanflow/core.py +818 -0
  25. kalmanflow-0.1.0a1/src/kalmanflow/flags.py +27 -0
  26. kalmanflow-0.1.0a1/src/kalmanflow/kalman.py +750 -0
  27. kalmanflow-0.1.0a1/src/kalmanflow/models.py +154 -0
  28. kalmanflow-0.1.0a1/src/kalmanflow/observations.py +39 -0
  29. kalmanflow-0.1.0a1/src/kalmanflow/pandas_api.py +41 -0
  30. kalmanflow-0.1.0a1/src/kalmanflow/pipeline.py +581 -0
  31. kalmanflow-0.1.0a1/src/kalmanflow/py.typed +1 -0
  32. kalmanflow-0.1.0a1/src/kalmanflow/reservoir_backend.py +107 -0
  33. kalmanflow-0.1.0a1/src/kalmanflow/reservoir_config.py +120 -0
  34. kalmanflow-0.1.0a1/src/kalmanflow/rts.py +230 -0
  35. kalmanflow-0.1.0a1/src/kalmanflow/time_utils.py +48 -0
  36. kalmanflow-0.1.0a1/src/kalmanflow/units.py +83 -0
@@ -0,0 +1,107 @@
1
+ # API reference
2
+
3
+ This page covers the complete supported package surface in KalmanFlow 0.1.0.
4
+ Import application-facing names from `kalmanflow`. The advanced types below are
5
+ also package exports, but are intended for custom models and integrations.
6
+ Names in modules beginning with `_` are internal and may change without a
7
+ compatibility promise.
8
+
9
+ `kalmanflow.__version__` reports the installed package version.
10
+
11
+ ## Reservoir estimation
12
+
13
+ | Export | Purpose |
14
+ | --- | --- |
15
+ | `Observation(timestamp, storage, discharge)` | Immutable streaming input. Timestamp must be timezone-aware and no finer than microsecond precision; use `NaN` for a missing value after initialization. |
16
+ | `OnlineReservoirInflow` | Default acre-ft/cfs streaming estimator. Construct with the five scalar diagonal noise values; use `process`, `process_many`, `initialized`, and `pending_count`. |
17
+ | `OnlineReservoirInflow.from_config(config)` | Creates a stream from reviewed configuration, including its units and smoothing lag. |
18
+ | `OnlineReservoirInflow.from_checkpoint(checkpoint, config=...)` | Restores a configured stream. The reservoir ID and fingerprinted model, covariance, lag, and unit settings must match. |
19
+ | `OnlineReservoirInflow.checkpoint()` | Produces resumable state after a successful call. It requires a reservoir ID; streams made with `from_config` have one. |
20
+ | `ReservoirFlowEstimate` | One timestamped inflow value with `prediction_flag` and `smoothing_flag`. |
21
+ | `ReservoirFlowUpdate` | The streaming return value: `filtered_inflows` are causal and `revised_inflows` are absolute, finalized replacements. |
22
+ | `get_reservoir_inflow(storage, outflow, ...)` | Default acre-ft/cfs batch estimator. It returns the six documented inflow and provenance columns. |
23
+ | `get_reservoir_inflow_from_config(storage, outflow, config)` | Batch estimator using a `ReservoirConfig` and its unit system. |
24
+ | `run_inflow_model(observations, ...)` | DataFrame convenience form of `get_reservoir_inflow`; the frame must have `storage` and `outflow` columns. |
25
+ | `OutputFlag` | `NORMAL` or `PREDICTED` describes observation completeness; `SMOOTHED` or `NON_SMOOTHED` describes estimate provenance. |
26
+
27
+ Both batch series must share an exactly equal timezone-aware, strictly
28
+ increasing index with at most microsecond timestamp precision. The package
29
+ neither aligns nor cleans inputs. A revised
30
+ inflow always replaces the causal value at the same timestamp; it is never a
31
+ delta. See [model behavior](INFLOW_MODEL_BEHAVIOR.md) for initialization,
32
+ partial observations, and output columns.
33
+
34
+ The reservoir inflow state is an inferred net balance contribution from the
35
+ supplied storage and accounted outflow. Measured outflow should cover outlet
36
+ releases, spills, and outward diversions or withdrawals as applicable.
37
+ Precipitation, evaporation, seepage, and other water exchanges are not
38
+ separate API terms. The linear-Gaussian estimate is not constrained to be
39
+ nonnegative; sensor bias, storage-datum changes, and rating-curve changes can
40
+ also appear as model mismatch.
41
+
42
+ ## Configuration and units
43
+
44
+ | Export | Purpose |
45
+ | --- | --- |
46
+ | `ReservoirConfig` | Immutable reservoir identity, covariance, lag, units, version, and metadata. `q`, `r`, and `p0` are 3×3, 2×2, and 3×3 covariance matrices respectively. |
47
+ | `InitializationStrategy` | Initialization policy enum. The currently supported value is `FIRST_TWO_VALID_STORAGE`. |
48
+ | `InflowUnits` | Metadata enum: `CUBIC_FEET_PER_SECOND` or `SYSTEM_FLOW_RATE`. CFS requires a unit system whose flow label is `cfs`. |
49
+ | `UnitSystem` | Volume/flow labels and rate-to-volume-per-second conversion. Use `us_customary()`, `si()`, `flow_to_volume()`, and `volume_to_flow_rate()` as needed. |
50
+ | `CFS_TO_ACRE_FEET_PER_SECOND` | Conversion constant used by the default acre-ft/cfs model. |
51
+
52
+ `q` is continuous-time diffusion covariance, not a discrete fixed-interval Q
53
+ matrix. All covariance inputs must be finite, symmetric, and positive
54
+ semidefinite; `r` must have a strictly positive diagonal. See
55
+ [configuration](CONFIGURATION.md) for a complete example.
56
+
57
+ ## Bayesian noise evaluation and tuning
58
+
59
+ The Bayesian search API is **experimental** and requires
60
+ `pip install "kalmanflow[tuning]"`. Its API, selection rules, and result schema may
61
+ change between releases. Review and independently validate any proposed
62
+ configuration. `evaluate_configuration` is available with the base installation;
63
+ the SciPy/scikit-learn optimizer dependencies load only when search is invoked.
64
+
65
+ | Export | Purpose |
66
+ | --- | --- |
67
+ | `ValidationWindow(name, start, end, weight=None)` | Named, timezone-aware half-open interval `[start, end)` for causal scoring. Window boundaries may use nanosecond precision even though scored observation timestamps are limited to microseconds. |
68
+ | `BayesianEvaluationSettings` | Shared warmup, innovation, calibration-warning, and numerical-stability settings. |
69
+ | `BayesianTuningSettings` | Trial counts, covariance multiplier bounds, acquisition settings, and optional upstream-proxy shape/timing gate. |
70
+ | `tune_inflow_noise_bayesian(...)` | Proposes five diagonal `q`/`r` terms from causal innovation scores. It needs at least three windows and a new proposed configuration version. |
71
+ | `BayesianTuningResult` | Proposed `selected_config`, parameters, trial summary, selected-window diagnostics, optional proxy diagnostics, warnings, and timing. |
72
+ | `evaluate_configuration(...)` | Scores one already-frozen configuration without searching or modifying it. |
73
+ | `ConfigEvaluationResult` | Frozen configuration, compact candidate summary, window diagnostics, and warnings. |
74
+ | `BayesianTuningError` | Raised when no candidate meets the tuning workflow's hard requirements. |
75
+
76
+ Tuning requires diagonal `q`, `r`, and `p0`, clean matched pandas series, at
77
+ least three positive inflow-increment seeds, and at least three distinct,
78
+ non-overlapping windows. Every seed is evaluated, so `total_trials` must be at
79
+ least the number of seeds. An optional upstream series is a shape/timing
80
+ diagnostic only, not a total-inflow label or Kalman observation. See
81
+ [Bayesian noise tuning](BAYESIAN_NOISE_TUNING.md) for selection and report
82
+ details.
83
+
84
+ ## Advanced state-space interfaces
85
+
86
+ These exports make it possible to use KalmanFlow's filtering and smoothing with
87
+ a different state-space model. They operate on NumPy arrays and do not impose
88
+ reservoir units.
89
+
90
+ | Export | Purpose |
91
+ | --- | --- |
92
+ | `kalman_filter(...)` | Batch linear-Gaussian filter. Transition, process, offset, and control arrays may be shared or have exactly `n_times - 1` entries; observation arrays may be shared or have exactly `n_times` entries. Model arrays must be finite, covariance arrays symmetric positive semidefinite, and `NaN` is the only omitted observation value. |
93
+ | `predict_state(...)` | Predicts one state mean and covariance, optionally with a control offset. |
94
+ | `initial_filter_step(...)` / `kalman_step(...)` | Create immutable timestamped filter records for the initial or a later update. |
95
+ | `FilterStep` / `KalmanFilterResult` | Filter records and complete batch outputs, including innovations, innovation covariance, update mask, transitions, and log likelihood. |
96
+ | `smooth_filter_steps(steps)` | Runs full-record RTS smoothing and returns one `SmoothedStep` per input step. Empty input returns an empty tuple. |
97
+ | `OnlineFixedLagRTS` | Time-lagged RTS smoother with `add_step`, `provisional`, active-step inspection, and a bounded memory window. |
98
+ | `SmoothedStep` | Immutable smoothed timestamp, state mean/covariance, and provenance flags. |
99
+ | `StateSpaceModel` / `ReservoirStateSpaceModel` | Protocol and physical three-state reservoir implementation. The reservoir state order is `[storage, inflow_rate, true_outflow_rate]`. |
100
+ | `ReservoirBackend` | Connects `ReservoirStateSpaceModel` to the generic pipeline. |
101
+ | `OnlineInflowPipeline` / `PipelineUpdate` | Generic ordered-stream coordinator and its per-call results. Use when supplying a compatible custom backend and smoother. |
102
+
103
+ `OnlineFixedLagRTS` releases a state once the elapsed-time lag has passed and
104
+ raises `OverflowError` if its active window reaches `max_window_steps` before
105
+ release. Generic pipeline state and protocol types are available from
106
+ `kalmanflow.pipeline` for custom checkpoint/replay implementations, but the
107
+ reservoir checkpoint byte format remains internal.
@@ -0,0 +1,83 @@
1
+ # KalmanFlow architecture
2
+
3
+ `kalmanflow` is a Python package for estimating reservoir inflow from measured storage and discharge. It implements a three-state, linear-Gaussian water-balance model plus a causal Kalman filter and fixed-lag RTS smoothing. It is a library: callers provide clean, aligned data and select or supply the reservoir configuration.
4
+
5
+ ## Package layout
6
+
7
+ | Module | Responsibility |
8
+ | --- | --- |
9
+ | `kalmanflow.kalman` | Generic predict, update, and batch filtering primitives. |
10
+ | `kalmanflow.rts` | Full-record and bounded online fixed-lag RTS smoothing. |
11
+ | `kalmanflow.models` | The physical three-state reservoir model and model protocol. |
12
+ | `kalmanflow.reservoir_backend` | Adapts the physical model to the streaming pipeline. |
13
+ | `kalmanflow.pipeline` | Generic initialization, ordering, replay, and delayed-release coordinator. |
14
+ | `kalmanflow.core` | Public batch and reservoir-streaming adapters. |
15
+ | `kalmanflow.pandas_api` | DataFrame convenience wrapper for the default batch adapter. |
16
+ | `kalmanflow.bayesian_tuning` | Public façade for Bayesian diagonal-noise search and compact frozen-configuration evaluation; implementation is split across preparation, diagnostics, candidate, optimization, proxy, and workflow modules. |
17
+ | `kalmanflow.reservoir_config` | Immutable, validated per-reservoir configuration. |
18
+ | `kalmanflow.observations` / `flags` | Public input and output-provenance types. |
19
+ | `kalmanflow.units` | Volume/flow-rate conversion systems. |
20
+ | `kalmanflow.time_utils` | UTC normalization and elapsed-time validation helpers. |
21
+
22
+ The package-level imports are the supported starting point for applications.
23
+ The generic Kalman, RTS, model, backend, and pipeline types are also exported
24
+ for advanced integrations; their contracts are listed in the
25
+ [API reference](API_REFERENCE.md). Checkpoint encoding and low-level validation
26
+ modules are internal implementation details and are not public serialization
27
+ formats or extension points.
28
+
29
+ Repository-specific workflows live outside the package in `applications/`:
30
+ `applications/preparation.py` parses and aligns the checked-in Aquarius
31
+ exports, while `applications/run_bayesian_tuner.py` provides the offline
32
+ Bayesian calibration entry point.
33
+ `Notebooks/prepare_reservoir_data.py` remains a compatibility import for
34
+ existing notebooks.
35
+
36
+ ## Reservoir state-space model
37
+
38
+ The default state is `[storage, inflow_rate, true_outflow_rate]`. Storage and measured discharge are observations; discharge is not treated as a known control. For an interval `dt` seconds, the first state row is:
39
+
40
+ ```text
41
+ storage(t + dt) = storage(t)
42
+ + dt × flow_to_volume_per_second
43
+ × (inflow_rate(t) - true_outflow_rate(t))
44
+ ```
45
+
46
+ The rate states are random walks. `ReservoirStateSpaceModel` derives the transition and discrete process covariance from the actual elapsed time and the continuous-time 3×3 diffusion covariance `q_continuous`.
47
+
48
+ The `inflow_rate` state is the residual net balance contribution required by
49
+ the supplied storage and accounted outflow. Measured outflow should include
50
+ outlet releases, spills, and outward diversions or withdrawals as applicable.
51
+ The API has no separate terms for precipitation, evaporation, seepage, or
52
+ other water exchanges; applications must account for those separately with
53
+ justified inputs or interpret them as model mismatch. Sensor bias, storage
54
+ datum changes, and rating-curve changes are additional mismatch sources.
55
+ Estimates are unconstrained by sign, so a negative value is mathematically
56
+ valid even when a physical interpretation would prompt data or model review.
57
+
58
+ `UnitSystem.us_customary()` is the default: storage is acre-feet and flow rates are cfs. `UnitSystem.si()` supports cubic metres and cubic metres per second; custom unit systems supply labels and a rate-to-volume-per-second conversion.
59
+
60
+ ## Public layers
61
+
62
+ Use `get_reservoir_inflow` for a pandas batch result, or `OnlineReservoirInflow` for a reservoir-specific stream. The latter is built on the generic `OnlineInflowPipeline`, which can also be used with another backend and smoother implementation.
63
+
64
+ The public reservoir layer exposes causal filtered inflow after its
65
+ two-observation startup phase and later exposes an absolute fixed-lag RTS
66
+ revised inflow for the same timestamp. The first finite-storage observation
67
+ anchors the stream; it does not produce an inflow output by itself.
68
+ Storage, measured outflow, and latent true outflow remain model inputs or
69
+ internal state; no public outflow result is exposed. Generic filter and
70
+ smoothing types remain available for advanced integrations.
71
+
72
+ Offline calibration uses `tune_inflow_noise_bayesian` with clean, aligned
73
+ storage and discharge series. It proposes a new immutable configuration but
74
+ does not persist or activate it. `evaluate_configuration` scores one frozen
75
+ configuration on a separate period without performing another search.
76
+
77
+ ## Data ownership and boundaries
78
+
79
+ KalmanFlow does not read files, fetch or clean source data, align series, or
80
+ deduplicate records. Callers manage reviewed configuration artifacts; the
81
+ package validates the runtime contract—including timestamp indexes,
82
+ matrix shapes, covariance properties, and missing observations—at the library
83
+ boundary.
@@ -0,0 +1,217 @@
1
+ # Experimental Bayesian innovation noise tuning
2
+
3
+ Bayesian tuning is experimental. Its API, selection rules, and result schema
4
+ may change between releases. Treat selected settings as proposals requiring
5
+ independent validation before operational use.
6
+
7
+ Install the optional search dependencies with `pip install "kalmanflow[tuning]"`.
8
+ The optimizer loads only when `tune_inflow_noise_bayesian` is called. Importing
9
+ KalmanFlow, filtering, smoothing, and `evaluate_configuration` work with the base
10
+ installation. Existing package-level imports remain supported.
11
+
12
+ ## Scope and inputs
13
+
14
+ `applications/run_bayesian_tuner.py` is the offline tuning workflow. Its
15
+ `run_bayesian_tuner(...)` function accepts `project_root`, `reservoir`,
16
+ `data_start`, `data_end`, `output_path`, and the optional
17
+ `report_output_path`. It searches five diagonal covariance terms informed by
18
+ the observed storage and outflow innovations:
19
+
20
+ * `q_storage`, `q_inflow`, and `q_outflow`;
21
+ * `r_storage` and `r_outflow`.
22
+
23
+ The package API accepts already-cleaned `pandas.Series` inputs. Storage and
24
+ discharge indexes must match exactly and must be timezone-aware, strictly
25
+ increasing, unique, and nonmissing. The base configuration must use diagonal
26
+ `q`, `r`, and `p0` matrices with positive finite values for all five searched
27
+ diagonal `q`/`r` terms. At least three distinct positive
28
+ `inflow_increment_sd_seeds` and three uniquely named, non-overlapping,
29
+ half-open `ValidationWindow` intervals are required.
30
+ Every supplied seed is evaluated. Consequently, `total_trials` must be at
31
+ least the number of seeds; when `initial_trials` is smaller, the initial design
32
+ is expanded to include them all.
33
+
34
+ ## Objective and selection
35
+
36
+ The tuner scores the causal, pre-update predictive innovation distribution. It
37
+ does not use true inflow values or smoothed states. Only rows inside the
38
+ declared validation windows contribute to the objective, although earlier
39
+ observations outside a window may causally condition its starting state. The
40
+ supplied hourly inflow increment seeds are used as deterministic initial trials
41
+ and define the `q_inflow` bounds.
42
+ Remaining trials are selected by expected improvement from a Gaussian-process
43
+ surrogate in normalized log-parameter space. Proxy diagnostics use only the
44
+ same warmup-adjusted validation rows as the innovation objective.
45
+
46
+ Each trial's objective is the validation-window-weighted joint predictive
47
+ negative log density per observed component, plus the configured multiple of
48
+ its between-window standard error. Hard validity checks exclude failed or
49
+ insufficiently scored trials. Statistically competitive trials are identified
50
+ with paired window differences; final selection prefers lower calibration
51
+ violation, then proximity to the reviewed base configuration, then objective.
52
+
53
+ `eligible` and `calibrated` answer different questions in both search and frozen
54
+ evaluation. An eligible result has the required number of finite scored storage
55
+ observations in every window and finite joint NLPD values. A search result is
56
+ calibrated only when it is eligible and satisfies the configured NIS,
57
+ innovation-bias, and autocorrelation checks. Frozen evaluation reports
58
+ calibrated only when it is eligible and has no configured NIS or innovation-bias
59
+ warnings; the search-only autocorrelation threshold is not applied by
60
+ `evaluate_configuration`. Disabling calibration warning thresholds does not
61
+ make an empty or numerically invalid evaluation eligible or calibrated.
62
+
63
+ The elapsed-time autocorrelation diagnostic centers the finite innovations and
64
+ normalizes each tolerance bin with the energy of the actual paired values. A
65
+ bin with fewer than three finite pairs is reported as unavailable rather than
66
+ being used as a calibration pass.
67
+
68
+ ## Output files and held-out evaluation
69
+
70
+ The runner writes the proposed immutable `ReservoirConfig` to `output_path`.
71
+ This compact configuration JSON is the default artifact and contains the
72
+ complete configuration needed by a downstream application. Without an
73
+ explicit `output_path`, it is written beneath `Outputs/bayesian_tuning`.
74
+
75
+ Pass `report_output_path` to additionally write a detailed tuning audit. The
76
+ report contains the compact candidate table (trial, acquisition source, five
77
+ parameters, objective, eligibility, and selection flags), selected-trial
78
+ per-window efficacy, compact upstream-proxy diagnostics when supplied, and a
79
+ reference to the selected configuration artifact. It does not duplicate the
80
+ full configuration JSON. The two resolved output paths must be different.
81
+
82
+ A separate final evaluation interval should be reserved and evaluated with
83
+ `evaluate_configuration` and an explicit `ValidationWindow` after tuning. The
84
+ application runner divides its requested segment into three calibration
85
+ windows; it does not reserve a final test segment automatically.
86
+
87
+ Runtime diagnostics distinguish Gaussian-process `acquisition_seconds` from
88
+ `candidate_evaluation_seconds`, which covers each compact causal evaluation.
89
+
90
+ By default, the Bayesian search cannot increase `q_storage` or `r_storage`
91
+ above the reviewed base configuration (`0.5x` to `1.0x`). This prevents an
92
+ innovation-only objective from explaining storm-driven storage changes as
93
+ extra storage process or measurement noise. These bounds remain configurable
94
+ through `BayesianTuningSettings` when independent evidence supports a wider
95
+ range. The inflow and outflow terms retain wider base-relative bounds.
96
+
97
+ ## Optional upstream proxy
98
+
99
+ An upstream gauge may be supplied as an optional diagnostic proxy:
100
+
101
+ ```python
102
+ from datetime import timedelta
103
+
104
+ from kalmanflow import (
105
+ BayesianEvaluationSettings,
106
+ BayesianTuningSettings,
107
+ tune_inflow_noise_bayesian,
108
+ )
109
+
110
+ evaluation_settings = BayesianEvaluationSettings()
111
+
112
+ result = tune_inflow_noise_bayesian(
113
+ storage,
114
+ discharge,
115
+ base_config,
116
+ inflow_increment_sd_seeds,
117
+ validation_windows,
118
+ upstream_proxy=upstream_flow,
119
+ settings=evaluation_settings,
120
+ bayesian_settings=BayesianTuningSettings(
121
+ proxy_max_lag=timedelta(hours=12),
122
+ proxy_diagnostic_frequency="1h",
123
+ proxy_min_shape_correlation=0.20,
124
+ proxy_min_change_correlation=0.10,
125
+ ),
126
+ proposed_configuration_version="reviewed-bayesian-v1",
127
+ )
128
+ ```
129
+
130
+ Proxy checks first match the gauge to the model timestamps, then perform a
131
+ bounded lag search using standardized level and change shape. They use only
132
+ the warmup-adjusted validation rows and aggregate at an hourly cadence by
133
+ default. Only whole diagnostic-cadence shifts whose duration is within
134
+ `proxy_max_lag` are considered, while zero lag remains available for a
135
+ sub-cadence maximum. A positive reported lag means the proxy is shifted later
136
+ relative to the causal estimate. They report best lag, level/change correlation,
137
+ normalized shape RMSE, and a gate flag in the result and JSON report. Proxy magnitude is
138
+ deliberately not matched and the proxy is never added to the Kalman observation
139
+ vector or treated as true total inflow. If at least one innovation-valid trial
140
+ passes the configured gate, selection is restricted to those trials; if none
141
+ passes, the result is returned with an explicit diagnostic-only warning.
142
+
143
+ The runtime preparation helper is `applications.preparation`. It retains
144
+ `upstream_flow` and spillway audit columns in `PreparedReservoirData`. When the
145
+ requested window contains at least one finite, nonnegative spillway value,
146
+ finite spillway values are added to outlet discharge; missing or negative
147
+ spillway values make combined outflow missing at those timestamps instead of
148
+ being treated as zero. If the spillway file is absent or the requested window
149
+ contains no usable spillway value, outlet-only outflow is preserved and the
150
+ window audit records that limitation. The application tuner uses this helper;
151
+ `Notebooks.prepare_reservoir_data` is retained only as a compatibility import
152
+ for existing notebooks.
153
+
154
+ ## Application runner
155
+
156
+ Run the workflow from the project root:
157
+
158
+ ```bash
159
+ uv run --extra tuning python applications/run_bayesian_tuner.py
160
+ ```
161
+
162
+ Direct execution uses the reviewed constants near the top of the runner;
163
+ edit those constants first or call `run_bayesian_tuner(...)` from Python with
164
+ explicit arguments. Without an explicit `output_path`, the JSON configuration is
165
+ written beneath `Outputs/bayesian_tuning` as the compact selected
166
+ configuration. Pass `report_output_path` to additionally write the detailed
167
+ tuning audit report; it must resolve to a different file.
168
+
169
+ For example, to produce both artifacts explicitly:
170
+
171
+ ```python
172
+ from pathlib import Path
173
+
174
+ from applications.run_bayesian_tuner import run_bayesian_tuner
175
+
176
+ run_bayesian_tuner(
177
+ output_path=Path("Outputs/bayesian_tuning/chesbro-config.json"),
178
+ report_output_path=Path("Outputs/bayesian_tuning/chesbro-report.json"),
179
+ )
180
+ ```
181
+
182
+ The package API is:
183
+
184
+ ```python
185
+ from kalmanflow import (
186
+ BayesianEvaluationSettings,
187
+ BayesianTuningSettings,
188
+ ValidationWindow,
189
+ evaluate_configuration,
190
+ tune_inflow_noise_bayesian,
191
+ )
192
+
193
+ result = tune_inflow_noise_bayesian(
194
+ storage,
195
+ discharge,
196
+ base_config,
197
+ inflow_increment_sd_seeds,
198
+ validation_windows,
199
+ settings=BayesianEvaluationSettings(),
200
+ bayesian_settings=BayesianTuningSettings(total_trials=40, initial_trials=12),
201
+ proposed_configuration_version="reviewed-bayesian-v1",
202
+ )
203
+
204
+ test_result = evaluate_configuration(
205
+ test_storage,
206
+ test_discharge,
207
+ result.selected_config,
208
+ evaluation_window=ValidationWindow("final-test", test_start, test_end),
209
+ settings=BayesianEvaluationSettings(),
210
+ )
211
+ ```
212
+
213
+ `evaluate_configuration` performs no search and does not modify the supplied
214
+ configuration. It returns a compact `ConfigEvaluationResult` containing one
215
+ configuration summary, per-window causal diagnostics, and warnings. When
216
+ `evaluation_window` is omitted it evaluates the full supplied series, but an
217
+ explicit untouched window is recommended for final approval.
@@ -0,0 +1,90 @@
1
+ # Configuration
2
+
3
+ ## Complete configuration example
4
+
5
+ Use `ReservoirConfig` when a reservoir has reviewed noise settings, units, and metadata. The values below are structurally valid examples, not recommended production parameters.
6
+
7
+ ```python
8
+ from datetime import timedelta
9
+
10
+ import numpy as np
11
+
12
+ from kalmanflow import (
13
+ InflowUnits,
14
+ InitializationStrategy,
15
+ ReservoirConfig,
16
+ UnitSystem,
17
+ )
18
+
19
+ config = ReservoirConfig(
20
+ reservoir_id="lexington",
21
+ reservoir_name="Lexington Reservoir",
22
+ q=np.diag([1.0, 0.01, 0.01]),
23
+ r=np.diag([100.0, 25.0]),
24
+ p0=np.diag([1_000.0, 100.0, 100.0]),
25
+ smoothing_lag=timedelta(hours=12),
26
+ initialization_strategy=InitializationStrategy.FIRST_TWO_VALID_STORAGE,
27
+ inflow_units=InflowUnits.CUBIC_FEET_PER_SECOND,
28
+ model_version="physical-rate-v1",
29
+ configuration_version="2026-01",
30
+ metadata={"source": "initial calibration"},
31
+ unit_system=UnitSystem.us_customary(),
32
+ )
33
+ ```
34
+
35
+ Then pass it to either adapter:
36
+
37
+ ```python
38
+ from kalmanflow import OnlineReservoirInflow, get_reservoir_inflow_from_config
39
+
40
+ stream = OnlineReservoirInflow.from_config(config)
41
+ batch_result = get_reservoir_inflow_from_config(storage, discharge, config)
42
+ ```
43
+
44
+ Both adapters use discharge as an observation input, but only return causal
45
+ and revised inflow. A revised inflow is an absolute fixed-lag-smoothed value
46
+ that replaces the causal inflow at its timestamp; it is not an adjustment to
47
+ add. The latent true outflow state is never a public result.
48
+
49
+ ## Matrix and unit conventions
50
+
51
+ The state order is `[storage, inflow_rate, true_outflow_rate]`.
52
+
53
+ | Field | Shape | Meaning |
54
+ | --- | --- | --- |
55
+ | `q` | 3×3 | Continuous-time process-diffusion covariance; `q[i,j]` has units `state_i × state_j / second`. |
56
+ | `r` | 2×2 | Measurement covariance for `[storage, measured_outflow]`; `r[i,j]` has units `observation_i × observation_j`. |
57
+ | `p0` | 3×3 | Initial covariance; `p0[i,j]` has units `state_i × state_j`. |
58
+
59
+ All covariance matrices must be finite, symmetric, and positive semidefinite. The two diagonal elements of `r` must be strictly positive. KalmanFlow derives each discrete process covariance from `q` and the actual elapsed time between observations, so never reuse a discrete, fixed-interval Q matrix as `q`.
60
+
61
+ The balance model infers a net storage balance contribution. Measured outflow
62
+ should include outlet releases, spills, and outward diversions or withdrawals
63
+ as applicable. Precipitation, evaporation, seepage, and other water exchanges
64
+ are not separate model terms; account for them with separate justified inputs
65
+ or document them as model mismatch. Sensor bias, storage-datum changes, and
66
+ rating-curve changes are additional mismatch sources. Estimates are not
67
+ constrained to be nonnegative.
68
+
69
+ `UnitSystem.us_customary()` uses acre-feet and cfs; `UnitSystem.si()` uses m³ and m³/s. A custom `UnitSystem` must supply a positive `flow_to_volume_per_second` conversion that matches the storage unit.
70
+
71
+ ## Selecting parameters
72
+
73
+ Choose `q`, `r`, and `p0` through an engineering review of each reservoir's
74
+ sensor accuracy, operating conditions, and historical data. Record the
75
+ rationale in `metadata`, version the reviewed configuration, and validate it
76
+ against a separate period before operational use.
77
+
78
+ ## Bayesian inflow-noise tuning
79
+
80
+ For an offline proposal, use `tune_inflow_noise_bayesian` with a fixed
81
+ `ReservoirConfig` and predeclared `ValidationWindow` intervals. The supplied
82
+ `inflow_increment_sd_seeds` initialize the bounded Bayesian search for the five
83
+ diagonal covariance terms. Bayesian evaluation requires `q`, `r`, and `p0` to
84
+ be diagonal. The tuner runs a causal chronology per trial and scores the joint
85
+ predictive innovation for whichever storage and outflow components are
86
+ observed, before that row is assimilated. The returned configuration is
87
+ proposed only; it is never persisted automatically. Supply a new
88
+ `proposed_configuration_version`, review the compact report, and use
89
+ `evaluate_configuration` with an explicit untouched `ValidationWindow` before
90
+ operational approval.
@@ -0,0 +1,84 @@
1
+ # KalmanFlow model behavior
2
+
3
+ ## Input contract
4
+
5
+ `get_reservoir_inflow` and `get_reservoir_inflow_from_config` accept pre-cleaned storage and discharge `pandas.Series` with exactly equal, timezone-aware, strictly increasing indexes. Observation timestamps support microsecond precision; finer pandas timestamps are rejected before computation. The default adapter expects storage in acre-feet and discharge in cfs. The configured adapter uses the labels in `ReservoirConfig.unit_system`.
6
+
7
+ `OnlineReservoirInflow.process` accepts either an `Observation` or `timestamp`, `storage`, and `discharge` keywords. Timestamps must be timezone-aware, no finer than microseconds, and strictly later than the previously accepted observation. All elapsed-time comparisons are normalized to UTC. Validation-window boundaries may retain nanosecond precision because they are interval labels rather than observations.
8
+
9
+ The package intentionally does not sort, align, deduplicate, interpolate, or impute inputs.
10
+
11
+ ## Initialization and missing values
12
+
13
+ The stream ignores leading rows with missing storage. Its first finite storage
14
+ reading must also have finite discharge; that row becomes the initialization
15
+ anchor. The next finite storage reading completes initialization, even if its
16
+ discharge is missing. At the anchor timestamp, the model uses measured outflow
17
+ as a steady-state inflow prior (zero initial storage-change assumption). That
18
+ first filtered value therefore depends only on the anchor observation, not on
19
+ the later storage reading. The second storage reading then updates inflow
20
+ through the normal causal predict/update step. A finite storage row with
21
+ missing discharge before an anchor is invalid rather than silently skipped.
22
+
23
+ After initialization, either storage or discharge may be `NaN`. A finite component is still used as a partial Kalman observation. If both are `NaN`, the step is predict-only. Positive and negative infinity are invalid; `NaN` is the only missing-value marker. A public estimate carrying any missing observation component receives the `PREDICTED` flag; fully observed steps are `NORMAL`.
24
+
25
+ ## Batch outputs
26
+
27
+ The batch adapters return a `DataFrame`, indexed like the inputs, with:
28
+
29
+ | Column | Meaning |
30
+ | --- | --- |
31
+ | `estimated_inflow` | Causal filtered inflow rate. |
32
+ | `revised_inflow` | Absolute fixed-lag-smoothed inflow replacement, or `NaN` until final. |
33
+ | `estimated_inflow_flag` / `revised_inflow_flag` | `NORMAL` or `PREDICTED`. |
34
+ | `estimated_inflow_smoothing_flag` | Always `NON_SMOOTHED`: inflow is causal. |
35
+ | `revised_inflow_smoothing_flag` | `SMOOTHED` when released; otherwise `NON_SMOOTHED`. |
36
+
37
+ When a revised inflow becomes available, it replaces `estimated_inflow` at the
38
+ same timestamp; it is not a delta to add to that value. Storage and outflow
39
+ are internal states and are not public batch outputs.
40
+
41
+ ## Streaming outputs
42
+
43
+ Each call to `OnlineReservoirInflow.process` returns a `ReservoirFlowUpdate`.
44
+ `filtered_inflows` contains newly available causal inflow estimates;
45
+ `revised_inflows` contains only newly finalized, absolute smoothed
46
+ replacements at their original timestamps. The initial successful call that
47
+ completes initialization can emit two filtered inflow estimates. The active
48
+ smoothing window is retained internally and is bounded by `max_window_steps`.
49
+ Although both startup estimates are emitted together, the anchor estimate was
50
+ computed from the anchor observation alone.
51
+
52
+ The anchor state mean is seeded from the anchor storage and measured outflow,
53
+ then that same anchor observation is assimilated by the filter. Its innovation
54
+ is therefore zero by construction and the covariance can be reduced. This startup
55
+ covariance is an engineering approximation rather than a measurement-
56
+ conditioned uncertainty calibration; interpret early uncertainty and early
57
+ calibration diagnostics accordingly.
58
+
59
+ `process_many` is transactional: if any item is invalid, the stream returns to its entry state and produces no partial group result.
60
+
61
+ ## Configuration
62
+
63
+ `ReservoirConfig` is immutable. It holds a reservoir identifier, model and configuration versions, continuous-time `q` (3×3), measurement covariance `r` (2×2), initial covariance `p0` (3×3), a positive smoothing lag, units, and metadata. Covariances must be finite, symmetric, positive semidefinite; the diagonal of `r` must be strictly positive.
64
+
65
+ Select and review continuous-time process noise (`q`), observation noise (`r`),
66
+ and initial covariance (`p0`) for each reservoir before operational use.
67
+ Document the rationale and configuration version in the configuration metadata.
68
+
69
+ ## Bayesian causal evaluation
70
+
71
+ `tune_inflow_noise_bayesian` is an offline calibration aid, not an adaptive
72
+ streaming mode. It uses the forward Kalman filter and never calls the RTS
73
+ smoother. Initialization rows and the configured warm-up period are excluded
74
+ from scores. Only rows inside the declared validation windows contribute to
75
+ the objective. Earlier observations outside a window may still condition its
76
+ causal starting state, and validation observations can condition later
77
+ predictions, as they would in operation; no observation can affect its own or
78
+ an earlier score.
79
+
80
+ Use `evaluate_configuration` once a proposed configuration is frozen for a
81
+ separate, untouched test period. Test results must not be used to change the
82
+ Bayesian search, windows, thresholds, or selected parameters. Filtered storage
83
+ closure is a reconstruction diagnostic; it is not an independent predictive
84
+ score when the ending storage observation has already been assimilated.
@@ -0,0 +1,63 @@
1
+ # Online inflow pipeline
2
+
3
+ `OnlineInflowPipeline` coordinates an ordered observation stream with a model-specific backend and a fixed-lag smoother. `OnlineReservoirInflow` is the reservoir-facing wrapper and is the normal choice for application code.
4
+
5
+ ## Lifecycle
6
+
7
+ 1. The pipeline accepts timezone-aware, strictly increasing observations with at most microsecond timestamp precision. Nanosecond-level `ValidationWindow` boundaries remain valid for half-open scoring intervals.
8
+ 2. It ignores leading missing-storage rows, then stores the first finite-storage observation (which must have finite discharge) and waits for the next finite storage sample to initialize the backend.
9
+ 3. Initialization creates two forward filter steps. The first uses outflow as a steady-state inflow prior and does not inspect the second observation; the second is a normal causal update. Each later input creates one forward step, using the actual elapsed seconds since its predecessor.
10
+ 4. Each step enters the fixed-lag RTS smoother. A state is released only after the configured elapsed-time lag has passed.
11
+
12
+ `PipelineUpdate.filtered_state` is the latest forward step; `filtered_states` contains every step created by that call; and `smoothed_states` contains just the newly finalized states. `provisional_states` may inspect the active lag window without releasing it.
13
+
14
+ ## Reservoir API
15
+
16
+ Create `OnlineReservoirInflow` with scalar noise parameters for the default acre-ft/cfs model, or use `OnlineReservoirInflow.from_config(config)` for a validated configuration:
17
+
18
+ ```python
19
+ from kalmanflow import Observation, OnlineReservoirInflow
20
+
21
+ stream = OnlineReservoirInflow(
22
+ q_storage=1.0, q_inflow=1.0, q_outflow=1.0,
23
+ r_storage=1.0, r_outflow=1.0,
24
+ )
25
+ update = stream.process(Observation(timestamp, storage, discharge))
26
+ ```
27
+
28
+ `filtered_inflows` are causal and marked `NON_SMOOTHED`. `revised_inflows`
29
+ are released only after smoothing and marked `SMOOTHED`; each is an absolute
30
+ inflow value that replaces the causal estimate at the same timestamp.
31
+ Only `NaN` denotes a missing storage or discharge value; infinities are
32
+ rejected before pipeline state changes.
33
+
34
+ ## Checkpoints
35
+
36
+ Configured reservoir streams support compact checkpoints after a successful `process` or `process_many` call:
37
+
38
+ ```python
39
+ checkpoint = stream.checkpoint()
40
+ restored = OnlineReservoirInflow.from_checkpoint(checkpoint, config=config)
41
+ ```
42
+
43
+ A checkpoint is bound to one `reservoir_id` and a fingerprint of the model,
44
+ covariance, smoothing-lag, and unit settings. Restore rejects a configuration
45
+ that does not match. Restoration rebuilds the active smoothing window from its
46
+ retained observations without re-emitting already released records.
47
+
48
+ Checkpoints are version-bound. Checkpoints created before the package rename
49
+ are rejected during restore and must be recreated.
50
+
51
+ Checkpoint timestamps preserve microsecond precision. Inputs with a nonzero
52
+ nanosecond remainder are rejected before pipeline state changes, so restoring a
53
+ checkpoint cannot silently change an accepted observation label.
54
+
55
+ Checkpoints are unavailable during processing, before a successful `process` or
56
+ `process_many` result, after a failed processing call, or for streams created
57
+ without a reservoir ID. The serialized bytes are intentionally an internal
58
+ format; retain the compatible configuration and use the package's restore API
59
+ instead of decoding them yourself.
60
+
61
+ ## Resource bounds and failures
62
+
63
+ `max_window_steps` must be at least two and bounds the active smoother window. Exceeding that bound, invalid timestamps, incompatible checkpoint data, or a backend failure raises an exception. `process_many` rolls back its entire input group on failure; single `process` validates before forwarding values to the pipeline.