AI Model Regression Dashboard
Real-time evaluation metrics, statistical baseline comparison, and regression anomaly alerts.
Quality / Exact Match
94.8%
-1.4%
Target: ≥92.0%AI Judge ScoreAI JUDGE
96.2%
+0.8%
Semantic CorrectnessMean Latency
240ms
+6.2%
p95: 340msAvg Cost / Case
$0.00340
-2.1%
USD per callQuality Pass Rate Trend
94.8 %
Run #101Run #105
Latency Trend (p50 / Mean)
240 ms
Run #101Run #105
Estimated Cost Trend
0.003 USD
Run #101Run #105
Active Production Baseline Comparison
Baseline: Production v2.4 Gold Baseline (gpt-4o)
No active regressions detected against baseline. All evaluated metrics are within safe operational thresholds.
Recent Evaluation Runs
View all runs →| Run ID | Dataset | Model | Status | Decision | Cases | Trigger | Started |
|---|---|---|---|---|---|---|---|
| No evaluation runs recorded yet. Click "Run Quick Evaluation" above to start. | |||||||