Expand description
Model worker abstraction and implementations
Provides the ModelWorker trait and production-ready implementations:
- EchoWorker: Testing/demo worker
- OpenAiWorker: OpenAI API (GPT-4, GPT-3.5, etc.)
- AnthropicWorker: Anthropic Claude API
- LlamaCppWorker: Local llama.cpp server
- VllmWorker: vLLM inference server
§Environment Variables
OPENAI_API_KEY: Required for OpenAiWorkerANTHROPIC_API_KEY: Required for AnthropicWorkerLLAMA_CPP_URL: llama.cpp server URL (default: http://localhost:8080)VLLM_URL: vLLM server URL (default: http://localhost:8000)
§Design note: workers are pipeline components
Workers are designed to be used through the pipeline orchestrated in
stages.rs, not called directly in production code. The pipeline provides
circuit breaking, backpressure, dead-letter queuing, and timeout handling
around every worker call.
If you call a worker directly (e.g. in tests or custom integrations), you must supply your own retry, timeout, and backoff strategy.
§Note (debug builds only)
In debug builds (cfg(debug_assertions)) consider adding assertions that
verify a pipeline context is present when calling workers directly, to
catch accidental direct use in integration tests.
Structs§
- Anthropic
Worker - Anthropic Claude API worker
- Echo
Worker - Dummy echo worker for testing
- Llama
CppWorker - llama.cpp HTTP server worker
- Load
Balanced Worker - A worker pool that distributes inference requests across multiple backends.
- Open
AiWorker - OpenAI API worker (GPT-4, GPT-3.5-turbo-instruct, etc.)
- Vllm
Worker - vLLM inference server worker
Enums§
- Load
Balance Strategy - Strategy for distributing requests across a pool of workers.
Traits§
- Model
Worker - Trait for model inference workers
Functions§
- stream_
worker - Stream tokens from any
ModelWorker.
Type Aliases§
- Token
Stream - Boxed streaming token iterator returned by
infer_stream.