Expand description
Token Budget Middleware
Pre-estimates token counts before requests reach the LLM, enforcing spend limits at the orchestration layer rather than discovering them via a billing surprise.
§Why estimate before sending?
LLM APIs charge per token. Without pre-flight counting:
- A single runaway prompt (e.g. 200 KB user-uploaded doc) consumes the entire daily budget in one call.
- Cost anomalies are invisible until the invoice arrives.
- There is no way to shed low-priority requests before they hit the API.
TokenBudgetGuard adds a zero-network-round-trip gate that estimates token
count, checks it against configurable per-request and period limits, and
rejects requests that would overflow the budget before they reach the wire.
§Token estimation
The estimator uses the ⌈len / 4⌉ heuristic (one token ≈ 4 UTF-8 bytes
in English prose). This intentionally over-estimates by ~10–15 % on code
and ~5 % on English, providing a conservative safety margin without
requiring a tokenizer dependency. For accurate accounting of actual tokens
used, pair this with crate::metrics which records real token counts
from provider responses.
§Example
use tokio_prompt_orchestrator::token_budget::{TokenBudgetGuard, TokenBudgetConfig};
let guard = TokenBudgetGuard::new(TokenBudgetConfig {
max_tokens_per_request: 4_096,
max_tokens_per_period: 1_000_000,
period: std::time::Duration::from_secs(3600), // 1-hour rolling window
});
let prompt = "Summarise this document in three bullet points.";
match guard.check(prompt) {
Ok(estimated) => println!("Allowed — estimated {estimated} tokens"),
Err(e) => eprintln!("Budget exceeded: {e}"),
}Structs§
- Token
Budget Config - Configuration for
TokenBudgetGuard. - Token
Budget Guard - Guards LLM requests against token over-spend.
Enums§
- Token
Budget Error - Token budget rejection reasons.
Functions§
- estimate_
tokens - Estimate the number of tokens in
text.