Skip to main content

estimate_tokens

Function estimate_tokens 

Source
pub fn estimate_tokens(text: &str) -> u64
Expand description

Estimate the number of tokens in text.

Uses the ⌈len / 4βŒ‰ heuristic: one token β‰ˆ 4 UTF-8 bytes in English prose. Over-estimates by ~10–15 % on code; ~5 % on English. Never returns 0 for non-empty input (minimum 1 token).