Skip to main content

Module eval_harness

Module eval_harness 

Source
Expand description

§Evaluation Harness

Compare prompt strategies across a shared set of EvalCase items. Supports multiple scoring metrics, composite ranking, tag filtering, and per-difficulty pass-rate breakdowns.

§Example

use tokio_prompt_orchestrator::eval_harness::{
    EvalCase, EvalHarness, EvalMetric,
};

let mut harness = EvalHarness::new();
harness.add_case(EvalCase {
    id: "q1".into(),
    prompt: "What is 2+2?".into(),
    reference_answer: Some("4".into()),
    tags: vec!["math".into()],
    difficulty: 1,
});
let report = harness.run_eval(
    "baseline",
    vec![("4".to_string(), 120, 0.001)],
);
assert!(report.pass_rate > 0.99);

Structs§

EvalCase
A single evaluation case.
EvalHarness
Evaluation harness for comparing prompt strategies.
EvalReport
Aggregated results for one strategy across all cases.
EvalResult
Result for a single case in a strategy evaluation.

Enums§

EvalMetric
Scoring metric applied to a strategy response.