Expand description
§Provider Arbitrage Engine
Routes inference requests to the cheapest provider that historically meets a caller-specified latency SLA.
§Problem
Multiple providers (Anthropic, OpenAI, vLLM, llama.cpp) have different per-token prices and different latency characteristics. A naïve router always uses the cheapest provider, but that provider may be slower or less reliable. A naïve latency router always uses the fastest provider, but wastes money.
The arbitrage engine tracks a rolling P95 latency window for each registered provider and, given a latency budget, selects whichever eligible provider (i.e., whose P95 is below the budget) has the lowest per-token cost.
§Guarantees
- Thread-safe: all hot-path state uses atomics; the latency ring buffer uses
a
Mutexbut the lock is held for microseconds. - Non-blocking:
select_providernever performs I/O. - Graceful degradation: if no provider meets the SLA, the one with the lowest P95 latency is returned (best-effort).
§Example
use tokio_prompt_orchestrator::routing::arbitrage::{ArbitrageEngine, ProviderProfile};
use std::time::Duration;
let engine = ArbitrageEngine::new();
engine.register(ProviderProfile {
name: "anthropic".to_string(),
cost_per_1k_input_tokens: 0.003,
cost_per_1k_output_tokens: 0.015,
priority: 0,
});
engine.register(ProviderProfile {
name: "openai".to_string(),
cost_per_1k_input_tokens: 0.005,
cost_per_1k_output_tokens: 0.015,
priority: 0,
});
// Record observed latencies
engine.record_latency("anthropic", Duration::from_millis(320));
engine.record_latency("openai", Duration::from_millis(180));
// With a 500ms SLA budget, pick the cheapest provider that historically
// completes within 500ms.
let sla = Duration::from_millis(500);
let winner = engine.select_provider(Some(sla));
// anthropic is cheaper AND within the 500ms SLA → selected
assert_eq!(winner.as_ref().map(|p| p.name.as_str()), Some("anthropic"));Structs§
- Arbitrage
Engine - The provider arbitrage engine.
- Provider
Profile - Static profile for a single inference provider.
- Provider
Snapshot - Point-in-time snapshot of a single provider’s state.