Skip to main content

Module tournament

Module tournament 

Source
Expand description

§Provider Tournament Mode

Dispatches the same PromptRequest to multiple inference workers in parallel and returns the highest-quality response according to a pluggable scoring function.

§Why tournament mode?

Different LLM providers (or different models on the same provider) excel at different tasks. For high-value requests — customer-facing responses, code generation, document summarisation — it can be worth spending 2–4× the normal inference cost to guarantee a better answer.

Tournament mode lets you:

  • A/B test providers automatically and record which wins most often.
  • Hedge against provider outages: the first successful response wins.
  • Tune quality vs. cost by changing the scoring function.

§Built-in scorers

ScorerStrategy
LongestResponseScorerPrefer the longest non-empty response (proxy for detail).
FastestResponseScorerPrefer the response that arrived first (minimise latency).
KeywordDensityScorerPrefer the response with the highest density of caller-supplied keywords.

Implement ResponseScorer to add your own.

§Example

use std::sync::Arc;
use std::collections::HashMap;
use tokio_prompt_orchestrator::{SessionId, PromptRequest, EchoWorker};
use tokio_prompt_orchestrator::enhanced::tournament::{
    TournamentRunner, TournamentConfig, LongestResponseScorer,
};

#[tokio::main]
async fn main() {
    let workers: Vec<Arc<dyn tokio_prompt_orchestrator::ModelWorker>> = vec![
        Arc::new(EchoWorker::new()),
        Arc::new(EchoWorker::new()),
    ];

    let runner = TournamentRunner::new(
        workers,
        Arc::new(LongestResponseScorer),
        TournamentConfig::default(),
    );

    let req = PromptRequest {
        session: SessionId::new("demo"),
        request_id: "t1".into(),
        input: "Explain quantum entanglement".into(),
        meta: HashMap::new(),
        deadline: None,
    };

    if let Ok(result) = runner.run(req).await {
        println!("Winner (worker {}): {}", result.winner_index, result.response);
    }
}

Structs§

FastestResponseScorer
Scores by latency — the fastest response scores highest.
KeywordDensityScorer
Scores by the density of caller-supplied keywords in the response.
LongestResponseScorer
Scores by response length — longer responses score higher.
TournamentConfig
Configuration for TournamentRunner.
TournamentResult
The outcome of a tournament run.
TournamentRunner
Runs a prompt through multiple workers and picks the best response.
TournamentStats
Statistics for TournamentRunner.

Traits§

ResponseScorer
Scores a candidate inference response.