AI Model Regression Dashboard

Real-time evaluation metrics, statistical baseline comparison, and regression anomaly alerts.

All Evaluations
Quality / Exact Match
94.8%
-1.4%
Target: ≥92.0%
AI Judge ScoreAI JUDGE
96.2%
+0.8%
Semantic Correctness
Mean Latency
240ms
+6.2%
p95: 340ms
Avg Cost / Case
$0.00340
-2.1%
USD per call
Quality Pass Rate Trend
94.8 %
Run #101Run #105
Latency Trend (p50 / Mean)
240 ms
Run #101Run #105
Estimated Cost Trend
0.003 USD
Run #101Run #105

Active Production Baseline Comparison

Baseline: Production v2.4 Gold Baseline (gpt-4o)

Manage Baselines
No active regressions detected against baseline. All evaluated metrics are within safe operational thresholds.

Recent Evaluation Runs

View all runs →
Run IDDatasetModelStatusDecisionCasesTriggerStarted
No evaluation runs recorded yet. Click "Run Quick Evaluation" above to start.