Skip to main content

Module metrics

Module metrics 

Source
Expand description

Prometheus metrics for the orchestrator pipeline.

§Usage

Call init_metrics once at process startup before spawning any pipeline stages. The helper functions (record_stage_latency, inc_request, …) are no-ops if init_metrics was never called, so the pipeline is always safe to run - observability simply degrades gracefully.

§Metrics Exposed

NameTypeLabels
orchestrator_requests_totalCounterstage
orchestrator_requests_shed_totalCounterstage
orchestrator_requests_dropped_totalCounterstage
orchestrator_errors_totalCounterstage, err_type
orchestrator_stage_duration_secondsHistogramstage
orchestrator_queue_depthGaugestage
inference_time_to_first_token_secondsHistogramworker, model
orchestrator_requests_expired_totalCounter(none)

Structs§

Metrics
All Prometheus metrics for the orchestrator, bundled together so they can be stored in a single OnceLock and initialised atomically.
MetricsSummary
A structured snapshot of key metric counters, used by the health endpoint.

Functions§

gather
Gather all registered metrics as a raw list of metric families.
gather_metrics
Gather and encode all metrics in the Prometheus text exposition format.
get_metrics_summary
Return a structured summary of current metric counter values.
inc_cache_hit
Increment the cache hit counter.
inc_cache_miss
Increment the cache miss counter.
inc_cb_rejected
Increment the circuit-breaker rejected-requests counter.
inc_cb_transition
Increment the circuit-breaker state-transition counter for state.
inc_config_reload_error
Increment the config reload error counter.
inc_dedup_hash_collision
Increment the dedup hash collision counter.
inc_dedup_hit
Increment the dedup hit counter.
inc_dedup_waiter_unblocked
Increment the dedup waiters-unblocked counter.
inc_dlq_lock_poisoned
Increment the DLQ lock-poisoned counter.
inc_dropped
Increment the dropped-request counter for a pipeline stage.
inc_error
Increment the error counter for a pipeline stage and error type.
inc_expired
Increment the deadline-expired counter.
inc_inference_timeout
Increment the inference timeout counter.
inc_rag_expired
Increment the RAG-stage deadline-expired counter.
inc_request
Increment the request counter for a pipeline stage.
inc_session_affinity_hit
Increment the session affinity hit counter.
inc_session_affinity_miss
Increment the session affinity miss counter.
inc_shed
Increment the shed-request counter for a pipeline stage.
inc_worker_error
Increment the per-worker error counter.
init_metrics
Initialise all Prometheus metrics and register them with a private registry.
record_config_reload_duration
Record the duration of a config hot-reload operation.
record_inference_cost
Record USD cost for an inference call.
record_stage_latency
Record the processing latency for a pipeline stage.
record_ttft
Record the time-to-first-token (TTFT) for a streaming inference call.
set_queue_depth
Set the queue depth gauge for a pipeline stage.
set_rate_limiter_tokens
Set the rate-limiter tokens-remaining gauge for a session.
total_inference_cost_usd
Return the total USD spent on inference calls this session.