Expand description
Adaptive timeout management for LLM requests.
Tracks per-model latency samples and computes dynamic timeouts based on
observed percentile latencies and an exponential moving average. Timeouts
are bounded between 5 s and 120 s and can be scaled up for high queue
depths via AdaptiveTimeoutManager::adjust_for_load.
Structsยง
- Adaptive
Timeout Manager - Thread-safe, per-model adaptive timeout manager.
- Model
Timeout Stats - Per-model statistics maintained by
AdaptiveTimeoutManager. - Timeout
Sample - A single latency/outcome sample for one LLM request.
- Timeout
Summary - A snapshot of timeout statistics for a single model.