Expand description
Bulkhead pattern — isolates concurrent execution pools to prevent cascade failures.
§What is the bulkhead pattern?
The bulkhead pattern (named after watertight compartments in a ship) isolates concurrent execution into separate, size-limited pools. If one pool fills up — because a downstream service is slow or unavailable — only callers targeting that pool are rejected. All other pools continue to operate normally, preventing a partial outage from cascading into a total system failure.
Each Bulkhead wraps a tokio semaphore. Callers acquire a permit before
starting work; the permit is automatically released when it is dropped.
§When to use it
Use a bulkhead when you need to:
- Protect a critical resource — e.g. limit the number of concurrent calls to an expensive GPU inference backend so it is never overwhelmed.
- Prevent cascade failures — if the inference pool fills up, callers receive an immediate error rather than queuing indefinitely; other subsystems (RAG, post-processing) are unaffected.
- Enforce fairness — give each tenant or request class its own pool so a burst in one class cannot monopolise shared resources.
§How to size the semaphore
A good starting point for max_concurrent:
| Subsystem | Suggested starting value | Rationale |
|---|---|---|
| Local model inference (GPU) | num_gpus × 2 | Leave headroom for batching. |
| Cloud API (e.g. OpenAI) | 10–20 | Typical per-key rate-limit allows ~20 concurrent requests. |
| RAG / retrieval | 50–100 | I/O-bound, can be higher. |
| Post-processing | 100–200 | CPU-light, tolerate high concurrency. |
Start conservatively and increase based on observed queue depth and error rate.
A max_concurrent that is too large degrades the backend; too small wastes
throughput.
§Behaviour
acquireis non-blocking: it usestry_acquireunder the hood and immediately returnsErrif no permits are available.- The returned
BulkheadPermitreleases the semaphore slot on drop. - Multiple independent bulkheads can coexist for different subsystems.
§Example
use tokio_prompt_orchestrator::enhanced::Bulkhead;
// Allow at most 10 concurrent inference calls.
let inference_bh = Bulkhead::new("inference", 10);
match inference_bh.acquire() {
Ok(permit) => {
// Call the inference backend here.
// The slot is released automatically when `permit` is dropped.
drop(permit);
}
Err(e) => {
// All 10 slots are occupied; shed this request or return an error.
eprintln!("Bulkhead full: {e}");
}
}Structs§
- Bulkhead
- Limits concurrent executions to
max_concurrentusing a semaphore. - Bulkhead
Permit - An RAII permit that releases one bulkhead slot on drop.