Skip to main content

Module compression

Module compression 

Source
Expand description

Prompt compression — reduce token count before sending to the model.

Compressing prompts before inference reduces cost and latency, especially for long conversation histories or large context retrievals. This module provides several complementary strategies that can be chained.

§Strategies

StrategyWhat it doesTokens saved (typical)
WhitespaceCompressorCollapses redundant whitespace/newlines2–5%
RepetitionRemoverRemoves duplicate paragraphs/sentences5–15%
StopWordFilterRemoves low-information filler words5–20%
SentenceRankerKeeps only the top-K most relevant sentences20–60%
TruncationStrategyHard truncation with smart boundary detectionvariable
CompressionPipelineChains multiple strategiesadditive

§Usage

use tokio_prompt_orchestrator::compression::{CompressionPipeline, SentenceRanker, WhitespaceCompressor};

let pipeline = CompressionPipeline::new()
    .with(Box::new(WhitespaceCompressor))
    .with(Box::new(SentenceRanker::new(0.5))); // keep top 50% sentences

let original = "Long prompt with lots of redundancy...";
let (compressed, ratio) = pipeline.compress(original);
println!("Reduced by {:.1}%", (1.0 - ratio) * 100.0);

Structs§

CompressionPipeline
Chains multiple compression strategies sequentially.
CompressionResult
The output of a compression pass: compressed text + compression ratio.
RepetitionRemover
Removes duplicate sentences or paragraphs.
SentenceRanker
Keeps only the top-K most relevant sentences using TF-IDF scoring.
StopWordFilter
Removes common English stop words from the text.
TruncationStrategy
Hard truncation at a character limit, respecting sentence boundaries.
WhitespaceCompressor
Collapses redundant whitespace: multiple spaces → one, 3+ newlines → two.

Traits§

Compressor
Trait for a single compression strategy.