Expand description
§Request Deduplication
Coalesces identical in-flight requests so that only one backend call is made when multiple callers submit the same prompt concurrently.
§Design
Each call to RequestDeduplicator::submit computes a SHA-256 key over
model_id + prompt_text. If a request with that key is already in-flight,
the caller receives a DedupDecision::Waiting containing a
[tokio::sync::oneshot::Receiver] that resolves when the original request
completes. The first caller for a key receives
DedupDecision::Original and is responsible for calling
RequestDeduplicator::complete when the result is ready.
§TTL pruning
Entries older than RequestDeduplicator::TTL_SECS (30 s) are pruned on
every call to submit() to prevent unbounded memory growth if the original
caller crashes without completing.
Structs§
- Dedup
Stats - Snapshot of deduplicator statistics.
- Request
Deduplicator - Deduplicates identical in-flight LLM requests by their content hash.
- Request
Id - A unique identifier for an original (non-deduplicated) request.
Enums§
- Dedup
Decision - Result of submitting a request to the deduplicator.