Expand description
§Prompt Cache
Content-addressed, in-process LRU cache for LLM inference responses.
Cache keys are the SHA-256 hash of (model_id + prompt_text). Entries
carry a TTL; expired entries are evicted lazily on PromptCache::get and
eagerly via PromptCache::evict_expired. When the cache reaches
CacheConfig::max_entries the least-recently-used entry is evicted to
make room (LRU via VecDeque order tracking + HashMap for O(1) lookup).
The cache is fully thread-safe: the public handle is Arc<Mutex<CacheInner>>.
§Example
use tokio_prompt_orchestrator::cache::{CacheConfig, PromptCache};
use std::time::Duration;
let cfg = CacheConfig {
max_entries: 100,
default_ttl: Duration::from_secs(300),
max_prompt_len: 8192,
};
let cache = PromptCache::new(cfg);
// Cache a response.
cache.insert("gpt-4o", "Hello, world!", vec!["Hi there!".to_string()], None);
// Retrieve it.
let result = cache.get("gpt-4o", "Hello, world!");
assert!(result.is_some());Structs§
- Cache
Config - Configuration for
PromptCache. - Cache
Entry - A single cached inference response.
- Cache
Stats - Aggregate statistics for a
PromptCacheinstance. - Prompt
Cache - Thread-safe, content-addressed LRU prompt cache.