Skip to main content

Module cache

Module cache 

Source
Expand description

§Prompt Cache

Content-addressed, in-process LRU cache for LLM inference responses.

Cache keys are the SHA-256 hash of (model_id + prompt_text). Entries carry a TTL; expired entries are evicted lazily on PromptCache::get and eagerly via PromptCache::evict_expired. When the cache reaches CacheConfig::max_entries the least-recently-used entry is evicted to make room (LRU via VecDeque order tracking + HashMap for O(1) lookup).

The cache is fully thread-safe: the public handle is Arc<Mutex<CacheInner>>.

§Example

use tokio_prompt_orchestrator::cache::{CacheConfig, PromptCache};
use std::time::Duration;

let cfg = CacheConfig {
    max_entries: 100,
    default_ttl: Duration::from_secs(300),
    max_prompt_len: 8192,
};
let cache = PromptCache::new(cfg);

// Cache a response.
cache.insert("gpt-4o", "Hello, world!", vec!["Hi there!".to_string()], None);

// Retrieve it.
let result = cache.get("gpt-4o", "Hello, world!");
assert!(result.is_some());

Structs§

CacheConfig
Configuration for PromptCache.
CacheEntry
A single cached inference response.
CacheStats
Aggregate statistics for a PromptCache instance.
PromptCache
Thread-safe, content-addressed LRU prompt cache.