Expand description
§Evaluation Harness
Compare prompt strategies across a shared set of EvalCase items.
Supports multiple scoring metrics, composite ranking, tag filtering,
and per-difficulty pass-rate breakdowns.
§Example
use tokio_prompt_orchestrator::eval_harness::{
EvalCase, EvalHarness, EvalMetric,
};
let mut harness = EvalHarness::new();
harness.add_case(EvalCase {
id: "q1".into(),
prompt: "What is 2+2?".into(),
reference_answer: Some("4".into()),
tags: vec!["math".into()],
difficulty: 1,
});
let report = harness.run_eval(
"baseline",
vec![("4".to_string(), 120, 0.001)],
);
assert!(report.pass_rate > 0.99);Structs§
- Eval
Case - A single evaluation case.
- Eval
Harness - Evaluation harness for comparing prompt strategies.
- Eval
Report - Aggregated results for one strategy across all cases.
- Eval
Result - Result for a single case in a strategy evaluation.
Enums§
- Eval
Metric - Scoring metric applied to a strategy response.