Expand description
§Smart Adaptive Batcher
Groups incoming PromptRequests into micro-batches within a configurable
time window, dispatching them together to batch-capable inference workers.
§Why batching matters
GPU-accelerated inference is most efficient when multiple requests are processed together: a batch of 8 prompts may take only 20% longer than a single prompt while delivering 5–6× higher throughput. Without batching, each request occupies the full per-request GPU round-trip overhead.
§How it works
- Requests arrive via
SmartBatcher::submit. - The batcher collects them in an internal staging buffer.
- A batch is flushed when either:
- The batch reaches
BatchConfig::max_batch_size, or BatchConfig::max_wait_msmilliseconds elapse since the first item in the batch arrived.
- The batch reaches
- The caller drives flushing by calling
SmartBatcher::poll_readyin a loop (typically from a dedicated Tokio task).
§Similarity grouping (optional)
When BatchConfig::group_by_prefix_len is > 0 the batcher places
requests with the same prompt prefix (first N bytes) into the same batch.
This improves KV-cache utilisation on prefix-caching inference servers
(e.g. vLLM, SGLang).
§Example
use std::collections::HashMap;
use tokio_prompt_orchestrator::{SessionId, PromptRequest};
use tokio_prompt_orchestrator::enhanced::smart_batch::{SmartBatcher, BatchConfig};
#[tokio::main]
async fn main() {
let batcher = SmartBatcher::new(BatchConfig {
max_batch_size: 8,
max_wait_ms: 50,
group_by_prefix_len: 64,
});
// Producer: submit requests.
let make = |id: &str| PromptRequest {
session: SessionId::new(id),
request_id: id.to_string(),
input: format!("Summarise document {id}"),
meta: HashMap::new(),
deadline: None,
};
batcher.submit(make("a")).await;
batcher.submit(make("b")).await;
// Consumer: poll until a batch is ready.
if let Some(batch) = batcher.poll_ready().await {
println!("Dispatching batch of {} requests", batch.len());
// Pass `batch` to your batch-capable ModelWorker implementation.
}
}Structs§
- Batch
Config - Configuration for
SmartBatcher. - Batcher
Stats - Statistics produced by
SmartBatcher::stats. - Smart
Batcher - An adaptive batcher that groups
PromptRequests into micro-batches.