Skip to main content

Module worker

Module worker 

Source
Expand description

Model worker abstraction and implementations

Provides the ModelWorker trait and production-ready implementations:

  • EchoWorker: Testing/demo worker
  • OpenAiWorker: OpenAI API (GPT-4, GPT-3.5, etc.)
  • AnthropicWorker: Anthropic Claude API
  • LlamaCppWorker: Local llama.cpp server
  • VllmWorker: vLLM inference server

§Environment Variables

  • OPENAI_API_KEY: Required for OpenAiWorker
  • ANTHROPIC_API_KEY: Required for AnthropicWorker
  • LLAMA_CPP_URL: llama.cpp server URL (default: http://localhost:8080)
  • VLLM_URL: vLLM server URL (default: http://localhost:8000)

§Design note: workers are pipeline components

Workers are designed to be used through the pipeline orchestrated in stages.rs, not called directly in production code. The pipeline provides circuit breaking, backpressure, dead-letter queuing, and timeout handling around every worker call.

If you call a worker directly (e.g. in tests or custom integrations), you must supply your own retry, timeout, and backoff strategy.

§Note (debug builds only)

In debug builds (cfg(debug_assertions)) consider adding assertions that verify a pipeline context is present when calling workers directly, to catch accidental direct use in integration tests.

Structs§

AnthropicWorker
Anthropic Claude API worker
EchoWorker
Dummy echo worker for testing
LlamaCppWorker
llama.cpp HTTP server worker
LoadBalancedWorker
A worker pool that distributes inference requests across multiple backends.
OpenAiWorker
OpenAI API worker (GPT-4, GPT-3.5-turbo-instruct, etc.)
VllmWorker
vLLM inference server worker

Enums§

LoadBalanceStrategy
Strategy for distributing requests across a pool of workers.

Traits§

ModelWorker
Trait for model inference workers

Functions§

stream_worker
Stream tokens from any ModelWorker.

Type Aliases§

TokenStream
Boxed streaming token iterator returned by infer_stream.