Expand description
Custom plugin stage system for the LLM inference pipeline.
This module provides a first-class extension point so user code can inject arbitrary async processing logic into the pipeline without forking the library. Plugins are composable, ordered, and fully type-safe.
§Architecture
PromptRequest
|
v
[BeforeRag plugins] -> RAG stage -> [AfterRag plugins]
|
v
[BeforeAssemble plugins] -> Assemble stage -> [AfterAssemble plugins]
|
v
[BeforeInference plugins] -> Inference stage -> [AfterInference plugins]
|
v
[BeforePost plugins] -> Post stage -> [AfterPost plugins]
|
v
[BeforeStream plugins] -> Stream stage -> [AfterStream plugins]§Quick Start
use tokio_prompt_orchestrator::plugin::{
StagePlugin, PluginInput, PluginOutput, PluginPosition, PluginRegistry,
};
use tokio_prompt_orchestrator::PipelineStage;
use async_trait::async_trait;
use std::sync::Arc;
/// A simple logging plugin that records when the inference stage is entered.
struct InferenceLogPlugin;
#[async_trait]
impl StagePlugin for InferenceLogPlugin {
fn name(&self) -> &'static str { "inference-logger" }
async fn process(&self, input: PluginInput) -> PluginOutput {
tracing::info!(request_id = %input.request_id, "entering inference stage");
PluginOutput::passthrough(input)
}
}
#[tokio::main]
async fn main() {
let mut registry = PluginRegistry::new();
registry.register(
PluginPosition::Before(PipelineStage::Inference),
Arc::new(InferenceLogPlugin),
);
let chain = registry.chain_for(PluginPosition::Before(PipelineStage::Inference));
let input = PluginInput {
request_id: "req-1".to_string(),
session_id: "session-1".to_string(),
payload: serde_json::json!({"prompt": "hello"}),
metadata: std::collections::HashMap::new(),
};
let output = chain.run(input).await;
println!("Plugin chain produced: {:?}", output.status);
}Structs§
- Hook
Metrics - Metrics snapshot passed to pre- and post-inference hooks.
- Latency
Logger Plugin - A plugin that records per-request start/end timestamps (milliseconds since Unix epoch) into a shared latency log.
- Plugin
Chain - An ordered sequence of
StagePlugins that run at a singlePluginPosition. - Plugin
Context - Rich context passed to every pre- and post-inference hook.
- Plugin
Info - Snapshot of per-plugin call statistics.
- Plugin
Input - Input passed to every plugin in a
PluginChain. - Plugin
Output - Output returned by a
StagePlugin::processcall. - Plugin
Registry - Global runtime registry for
StagePlugininstances. - Plugin
V2Chain - Ordered list of
Plugins executed sequentially on each request/response. - Plugin
V2Registry - Runtime registry of
Plugininstances with enable/disable support and per-plugin call statistics. - Profanity
Filter Plugin - A plugin that blocks requests containing any word from a configurable list.
- Response
Length CapPlugin - A plugin that truncates response token lists to at most
max_tokensitems.
Enums§
- Plugin
Error - Error returned by
Pluginhooks. - Plugin
Position - Specifies where in the pipeline a plugin should be inserted.
- Plugin
Status - The outcome status of a plugin’s
processcall.
Traits§
- Inference
Hook - A focused hook trait for code that only needs to intercept inference
before or after it happens, without needing the full
StagePluginposition system. - Plugin
- Core trait for request/response plugins.
- Stage
Plugin - Core trait for custom pipeline stage plugins.