Skip to main content

Module token_budget

Module token_budget 

Source
Expand description

Token Budget Middleware

Pre-estimates token counts before requests reach the LLM, enforcing spend limits at the orchestration layer rather than discovering them via a billing surprise.

§Why estimate before sending?

LLM APIs charge per token. Without pre-flight counting:

  • A single runaway prompt (e.g. 200 KB user-uploaded doc) consumes the entire daily budget in one call.
  • Cost anomalies are invisible until the invoice arrives.
  • There is no way to shed low-priority requests before they hit the API.

TokenBudgetGuard adds a zero-network-round-trip gate that estimates token count, checks it against configurable per-request and period limits, and rejects requests that would overflow the budget before they reach the wire.

§Token estimation

The estimator uses the ⌈len / 4⌉ heuristic (one token ≈ 4 UTF-8 bytes in English prose). This intentionally over-estimates by ~10–15 % on code and ~5 % on English, providing a conservative safety margin without requiring a tokenizer dependency. For accurate accounting of actual tokens used, pair this with crate::metrics which records real token counts from provider responses.

§Example

use tokio_prompt_orchestrator::token_budget::{TokenBudgetGuard, TokenBudgetConfig};

let guard = TokenBudgetGuard::new(TokenBudgetConfig {
    max_tokens_per_request: 4_096,
    max_tokens_per_period: 1_000_000,
    period: std::time::Duration::from_secs(3600), // 1-hour rolling window
});

let prompt = "Summarise this document in three bullet points.";
match guard.check(prompt) {
    Ok(estimated) => println!("Allowed — estimated {estimated} tokens"),
    Err(e) => eprintln!("Budget exceeded: {e}"),
}

Structs§

TokenBudgetConfig
Configuration for TokenBudgetGuard.
TokenBudgetGuard
Guards LLM requests against token over-spend.

Enums§

TokenBudgetError
Token budget rejection reasons.

Functions§

estimate_tokens
Estimate the number of tokens in text.