Skip to main content

Module bulkhead

Module bulkhead 

Source
Expand description

Bulkhead pattern — isolates concurrent execution pools to prevent cascade failures.

§What is the bulkhead pattern?

The bulkhead pattern (named after watertight compartments in a ship) isolates concurrent execution into separate, size-limited pools. If one pool fills up — because a downstream service is slow or unavailable — only callers targeting that pool are rejected. All other pools continue to operate normally, preventing a partial outage from cascading into a total system failure.

Each Bulkhead wraps a tokio semaphore. Callers acquire a permit before starting work; the permit is automatically released when it is dropped.

§When to use it

Use a bulkhead when you need to:

  • Protect a critical resource — e.g. limit the number of concurrent calls to an expensive GPU inference backend so it is never overwhelmed.
  • Prevent cascade failures — if the inference pool fills up, callers receive an immediate error rather than queuing indefinitely; other subsystems (RAG, post-processing) are unaffected.
  • Enforce fairness — give each tenant or request class its own pool so a burst in one class cannot monopolise shared resources.

§How to size the semaphore

A good starting point for max_concurrent:

SubsystemSuggested starting valueRationale
Local model inference (GPU)num_gpus × 2Leave headroom for batching.
Cloud API (e.g. OpenAI)10–20Typical per-key rate-limit allows ~20 concurrent requests.
RAG / retrieval50–100I/O-bound, can be higher.
Post-processing100–200CPU-light, tolerate high concurrency.

Start conservatively and increase based on observed queue depth and error rate. A max_concurrent that is too large degrades the backend; too small wastes throughput.

§Behaviour

  • acquire is non-blocking: it uses try_acquire under the hood and immediately returns Err if no permits are available.
  • The returned BulkheadPermit releases the semaphore slot on drop.
  • Multiple independent bulkheads can coexist for different subsystems.

§Example

use tokio_prompt_orchestrator::enhanced::Bulkhead;

// Allow at most 10 concurrent inference calls.
let inference_bh = Bulkhead::new("inference", 10);

match inference_bh.acquire() {
    Ok(permit) => {
        // Call the inference backend here.
        // The slot is released automatically when `permit` is dropped.
        drop(permit);
    }
    Err(e) => {
        // All 10 slots are occupied; shed this request or return an error.
        eprintln!("Bulkhead full: {e}");
    }
}

Structs§

Bulkhead
Limits concurrent executions to max_concurrent using a semaphore.
BulkheadPermit
An RAII permit that releases one bulkhead slot on drop.