Expand description
Prometheus metrics for the orchestrator pipeline.
§Usage
Call init_metrics once at process startup before spawning any pipeline
stages. The helper functions (record_stage_latency, inc_request, …) are
no-ops if init_metrics was never called, so the pipeline is always safe to
run - observability simply degrades gracefully.
§Metrics Exposed
| Name | Type | Labels |
|---|---|---|
orchestrator_requests_total | Counter | stage |
orchestrator_requests_shed_total | Counter | stage |
orchestrator_requests_dropped_total | Counter | stage |
orchestrator_errors_total | Counter | stage, err_type |
orchestrator_stage_duration_seconds | Histogram | stage |
orchestrator_queue_depth | Gauge | stage |
inference_time_to_first_token_seconds | Histogram | worker, model |
orchestrator_requests_expired_total | Counter | (none) |
Structs§
- Metrics
- All Prometheus metrics for the orchestrator, bundled together so they can
be stored in a single
OnceLockand initialised atomically. - Metrics
Summary - A structured snapshot of key metric counters, used by the health endpoint.
Functions§
- gather
- Gather all registered metrics as a raw list of metric families.
- gather_
metrics - Gather and encode all metrics in the Prometheus text exposition format.
- get_
metrics_ summary - Return a structured summary of current metric counter values.
- inc_
cache_ hit - Increment the cache hit counter.
- inc_
cache_ miss - Increment the cache miss counter.
- inc_
cb_ rejected - Increment the circuit-breaker rejected-requests counter.
- inc_
cb_ transition - Increment the circuit-breaker state-transition counter for
state. - inc_
config_ reload_ error - Increment the config reload error counter.
- inc_
dedup_ hash_ collision - Increment the dedup hash collision counter.
- inc_
dedup_ hit - Increment the dedup hit counter.
- inc_
dedup_ waiter_ unblocked - Increment the dedup waiters-unblocked counter.
- inc_
dlq_ lock_ poisoned - Increment the DLQ lock-poisoned counter.
- inc_
dropped - Increment the dropped-request counter for a pipeline stage.
- inc_
error - Increment the error counter for a pipeline stage and error type.
- inc_
expired - Increment the deadline-expired counter.
- inc_
inference_ timeout - Increment the inference timeout counter.
- inc_
rag_ expired - Increment the RAG-stage deadline-expired counter.
- inc_
request - Increment the request counter for a pipeline stage.
- inc_
session_ affinity_ hit - Increment the session affinity hit counter.
- inc_
session_ affinity_ miss - Increment the session affinity miss counter.
- inc_
shed - Increment the shed-request counter for a pipeline stage.
- inc_
worker_ error - Increment the per-worker error counter.
- init_
metrics - Initialise all Prometheus metrics and register them with a private registry.
- record_
config_ reload_ duration - Record the duration of a config hot-reload operation.
- record_
inference_ cost - Record USD cost for an inference call.
- record_
stage_ latency - Record the processing latency for a pipeline stage.
- record_
ttft - Record the time-to-first-token (TTFT) for a streaming inference call.
- set_
queue_ depth - Set the queue depth gauge for a pipeline stage.
- set_
rate_ limiter_ tokens - Set the rate-limiter tokens-remaining gauge for a session.
- total_
inference_ cost_ usd - Return the total USD spent on inference calls this session.