Skip to main content

Module smart_batch

Module smart_batch 

Source
Expand description

§Smart Adaptive Batcher

Groups incoming PromptRequests into micro-batches within a configurable time window, dispatching them together to batch-capable inference workers.

§Why batching matters

GPU-accelerated inference is most efficient when multiple requests are processed together: a batch of 8 prompts may take only 20% longer than a single prompt while delivering 5–6× higher throughput. Without batching, each request occupies the full per-request GPU round-trip overhead.

§How it works

  1. Requests arrive via SmartBatcher::submit.
  2. The batcher collects them in an internal staging buffer.
  3. A batch is flushed when either:
  4. The caller drives flushing by calling SmartBatcher::poll_ready in a loop (typically from a dedicated Tokio task).

§Similarity grouping (optional)

When BatchConfig::group_by_prefix_len is > 0 the batcher places requests with the same prompt prefix (first N bytes) into the same batch. This improves KV-cache utilisation on prefix-caching inference servers (e.g. vLLM, SGLang).

§Example

use std::collections::HashMap;
use tokio_prompt_orchestrator::{SessionId, PromptRequest};
use tokio_prompt_orchestrator::enhanced::smart_batch::{SmartBatcher, BatchConfig};

#[tokio::main]
async fn main() {
    let batcher = SmartBatcher::new(BatchConfig {
        max_batch_size: 8,
        max_wait_ms: 50,
        group_by_prefix_len: 64,
    });

    // Producer: submit requests.
    let make = |id: &str| PromptRequest {
        session: SessionId::new(id),
        request_id: id.to_string(),
        input: format!("Summarise document {id}"),
        meta: HashMap::new(),
        deadline: None,
    };
    batcher.submit(make("a")).await;
    batcher.submit(make("b")).await;

    // Consumer: poll until a batch is ready.
    if let Some(batch) = batcher.poll_ready().await {
        println!("Dispatching batch of {} requests", batch.len());
        // Pass `batch` to your batch-capable ModelWorker implementation.
    }
}

Structs§

BatchConfig
Configuration for SmartBatcher.
BatcherStats
Statistics produced by SmartBatcher::stats.
SmartBatcher
An adaptive batcher that groups PromptRequests into micro-batches.