Expand description
Prompt compression — reduce token count before sending to the model.
Compressing prompts before inference reduces cost and latency, especially for long conversation histories or large context retrievals. This module provides several complementary strategies that can be chained.
§Strategies
| Strategy | What it does | Tokens saved (typical) |
|---|---|---|
WhitespaceCompressor | Collapses redundant whitespace/newlines | 2–5% |
RepetitionRemover | Removes duplicate paragraphs/sentences | 5–15% |
StopWordFilter | Removes low-information filler words | 5–20% |
SentenceRanker | Keeps only the top-K most relevant sentences | 20–60% |
TruncationStrategy | Hard truncation with smart boundary detection | variable |
CompressionPipeline | Chains multiple strategies | additive |
§Usage
use tokio_prompt_orchestrator::compression::{CompressionPipeline, SentenceRanker, WhitespaceCompressor};
let pipeline = CompressionPipeline::new()
.with(Box::new(WhitespaceCompressor))
.with(Box::new(SentenceRanker::new(0.5))); // keep top 50% sentences
let original = "Long prompt with lots of redundancy...";
let (compressed, ratio) = pipeline.compress(original);
println!("Reduced by {:.1}%", (1.0 - ratio) * 100.0);Structs§
- Compression
Pipeline - Chains multiple compression strategies sequentially.
- Compression
Result - The output of a compression pass: compressed text + compression ratio.
- Repetition
Remover - Removes duplicate sentences or paragraphs.
- Sentence
Ranker - Keeps only the top-K most relevant sentences using TF-IDF scoring.
- Stop
Word Filter - Removes common English stop words from the text.
- Truncation
Strategy - Hard truncation at a character limit, respecting sentence boundaries.
- Whitespace
Compressor - Collapses redundant whitespace: multiple spaces → one, 3+ newlines → two.
Traits§
- Compressor
- Trait for a single compression strategy.