Evaluation Protocolο
For Research (Current)ο
All evaluation is offline using historical data:
Static diagnostic: NDCG@3, Hit@3 on validation queries
Rollout simulation: Single-agent and multi-agent simulators
Historical replay: Compare against actual trip outcomes
Shadow evaluation: Record recommendations without execution
A/B testing: Statistical comparison (simulation only)
All results are clearly labeled with their source:
SIMULATIONβ Multi-agent simulatorHISTORICAL_REPLAYβ Historical data replaySHADOWβ Shadow evaluationREAL A/Bβ Not yet available
For Pilot Deployment (Future)ο
Phase 1: Shadow Evaluationο
Deploy recommendation API in shadow mode
Record vehicle positions and current zone from telemetry
Generate recommendations WITHOUT showing them to drivers
Compare AI recommendations against actual driver decisions
Evaluate after 2-4 weeks of data collection
Phase 2: Controlled A/B Testο
Randomly assign drivers to control (Hot Zone) and treatment (Two-Step) groups
Show recommendations only to treatment group
Track: revenue, utilization, empty distance, trip completion rate
Run for minimum 4 weeks (to capture day-of-week and weather effects)
Use bootstrap paired comparison for statistical analysis
Monitor for spillover effects between groups
Phase 3: Full Rolloutο
Deploy to all consenting drivers
Continuous shadow evaluation against non-participating drivers
Monitor for policy degradation as adoption increases
A/B test policy variants (different models, different constraints)
Metricsο
Primaryο
Metric |
Definition |
|---|---|
Revenue per vehicle-hour |
Total fare / total vehicle-hours |
Utilization |
Time with passenger / total time |
Empty distance |
Distance driven without passenger |
Secondaryο
Metric |
Definition |
|---|---|
Trip completion rate |
Completed trips / trip opportunities |
Recommendation acceptance |
Times driver followed AI recommendation |
Zone exposure |
Distribution of recommended zones |
Market saturation |
% of time a zone has more drivers than demand |
Statistical Rigorο
Bootstrap confidence intervals (2000 resamples, 95% CI)
Paired comparison (same seeds, same demand realizations)
Effect size (Cohenβs d)
Multiple comparison correction (Bonferroni for >2 metrics)
Important Caveatsο
No real-world validation exists β All current metrics are simulation-based
Simulator omits key dynamics β No congestion, airport queues, driver adaptation
Recommendation β execution β Drivers may ignore AI recommendations
Market equilibrium unknown β Mass adoption may degrade policy performance