Campaign overview
Evaluation split results
Evaluation summary ↗
Grouped utility outcomes ↗
Pareto selection ↗
Intra-campaign experiments
Inter-campaign experiments
Evaluation
Select a campaign to inspect its aggregate tuning metrics.
Three-tier tuning results
Pooled tuning-set scores per configuration, ordered from the final delivered result to upstream guide identification.
Mapbox · GPS
Mapbox · Telemetry
Google Maps · GPS
Google Maps · Telemetry
Tier-specific FP rate
Tier 3
Final end-to-end Utility Score per maneuver zone
Tier 2
Expected symbolic advice or silence observed within the maneuver zone
Tier 1
Maneuver-zone hit rate
Tier 1
Distinct guide F1
Pooled distinct-guide identification balance per configuration.
Tier 1
Guide false-positive rate
Pooled guide-selection false-positive rate per configuration.
Campaign Tier 3 sensitivity analysis
Which configuration knobs influence useful outcome delivery, what delivered-advice FP cost they carry, and which ones can probably be ignored.
Utility effects use pooled Useful outcome delivered where counts are available. FP cost uses the delivered advice false-positive rate.
Select a campaign to inspect parameter sensitivity.
Pairwise mean per-run differences
Each cell is the mean paired run-level difference: row candidate minus column baseline.
Rows are candidates; columns are baselines. Values are mean per-run differences. Green means the row is better for the selected metric. No statistical significance is implied.
Tuning configurations
Click a column header to sort ascending or descending. One sort at a time, because chaos already has enough features.