Skip to main content

Module arbitrage

Module arbitrage 

Source
Expand description

§Provider Arbitrage Engine

Routes inference requests to the cheapest provider that historically meets a caller-specified latency SLA.

§Problem

Multiple providers (Anthropic, OpenAI, vLLM, llama.cpp) have different per-token prices and different latency characteristics. A naïve router always uses the cheapest provider, but that provider may be slower or less reliable. A naïve latency router always uses the fastest provider, but wastes money.

The arbitrage engine tracks a rolling P95 latency window for each registered provider and, given a latency budget, selects whichever eligible provider (i.e., whose P95 is below the budget) has the lowest per-token cost.

§Guarantees

  • Thread-safe: all hot-path state uses atomics; the latency ring buffer uses a Mutex but the lock is held for microseconds.
  • Non-blocking: select_provider never performs I/O.
  • Graceful degradation: if no provider meets the SLA, the one with the lowest P95 latency is returned (best-effort).

§Example

use tokio_prompt_orchestrator::routing::arbitrage::{ArbitrageEngine, ProviderProfile};
use std::time::Duration;

let engine = ArbitrageEngine::new();

engine.register(ProviderProfile {
    name: "anthropic".to_string(),
    cost_per_1k_input_tokens: 0.003,
    cost_per_1k_output_tokens: 0.015,
    priority: 0,
});
engine.register(ProviderProfile {
    name: "openai".to_string(),
    cost_per_1k_input_tokens: 0.005,
    cost_per_1k_output_tokens: 0.015,
    priority: 0,
});

// Record observed latencies
engine.record_latency("anthropic", Duration::from_millis(320));
engine.record_latency("openai", Duration::from_millis(180));

// With a 500ms SLA budget, pick the cheapest provider that historically
// completes within 500ms.
let sla = Duration::from_millis(500);
let winner = engine.select_provider(Some(sla));
// anthropic is cheaper AND within the 500ms SLA → selected
assert_eq!(winner.as_ref().map(|p| p.name.as_str()), Some("anthropic"));

Structs§

ArbitrageEngine
The provider arbitrage engine.
ProviderProfile
Static profile for a single inference provider.
ProviderSnapshot
Point-in-time snapshot of a single provider’s state.