Experiment setup
One Tier 3 tuning winner per navigation assembly, followed by an aggregate Utility comparison on held-out, tuning or pooled outcomes.
The four configurations are always selected from Tier 3 tuning results. Ties prefer lower delivered-advice FP, then lower feature complexity, then variant ID. The comparison dataset changes only the aggregate Utility outcomes; it never reselects the winners.
Experiment result
Tier 3 configurations are selected on tuning and compared using aggregate Utility and delivered-advice false-positive outcomes from the chosen dataset.
Select a campaign to run an intra-campaign experiment.