Urban Mobility Decision Intelligence

📖 Documentation

  • Problem Formulation
    • Spatial-Temporal Taxi Zone Recommendation
      • Definition
      • State Space
      • Action Space
      • Objective
      • Reward Model
      • Transition Dynamics
      • Relocation Cost
      • Evaluation Metrics
      • Constraints
  • Methodology
    • Problem definition
    • Temporal protocol
    • Data cleaning
    • Baseline 1: hot zones
    • Baseline 2: single-step utility
    • Finite-horizon continuation model
    • Corrected model-based MDP
    • Static evaluation
    • Rollout evaluation
    • Statistical analysis
    • Counterfactual evaluation
    • Exposure analysis
  • Supervised Forecasting
    • Scope
    • Data and temporal boundary
    • Features
    • Outputs
    • Results and ablation
    • Reproduction
    • Limitations
  • Graph Learning
    • Leakage boundary
    • Features and models
    • Results
    • Reproduction
  • Multi-Agent Simulator
    • Purpose
    • Event semantics
    • Demand/supply ratio
    • Metrics
    • Benchmark
    • Reproduction
    • Remaining limitations
  • DQN and Double-DQN Baselines
    • Scope
    • Temporal isolation
    • Environment
    • Algorithms
    • Benchmark
    • Reproduction
  • Combined Benchmark
    • Endpoint matrix
    • Main conclusions
    • Reproduction
  • Reproducible Sensitivity and Ablation Results
    • Planning horizon
    • Executable parameter grid
    • Stress tests
    • Exposure concentration
    • Unsupported ablations removed

🧭 Platform

  • Decision Engine
    • Overview
    • Architecture
    • Schema
      • Recommendation
      • RankedZone
    • Policies
    • Constraints
    • Usage
    • Important Notes
  • API Documentation
    • Decision Intelligence Platform API
      • Source Label
      • Endpoints
      • Error Handling
      • Model Names
      • Interactive Docs
  • Leaderboard
    • Scope & Honesty Statement
    • Policy Leaderboard
    • Forecast Leaderboard (held-out)
    • External Submissions
    • How to Regenerate
  • Benchmark Protocol v2.0
    • Dataset
    • Task Definition
      • Forecasting Sub-task
      • Decision Making Sub-task
      • Offline RL Sub-task
    • Metrics
      • Forecasting
      • Decision Making
      • Offline RL
    • Baselines
      • Forecasting
      • Decision Making
      • RL
    • Evaluation Procedure
    • Reproducibility Requirements
  • Shadow Evaluation
    • Purpose
    • Principle
    • Current Status
    • Usage
    • Output Schema
    • Limitations
    • Production Path
  • Historical Replay Evaluation
    • Goal
    • Methodology
    • Results
    • Comparison with Simulator
    • Limitations
  • LLM Mobility Agent
    • What it is
    • Design
    • Honesty contract
    • Providers
    • Wire a provider
    • Run
  • Cross-City Extension Framework
    • Goal
    • Current Status
    • City Configuration
    • Required Data for New Cities
    • Extension Steps
    • Limitations
    • Future Work

🔬 Research Notes

  • Decision-Aware Forecasting
    • Research Question
    • Key Finding
    • Experiment Design
      • Models Compared
      • Metrics
    • Hypotheses
      • H1: Forecast MAE correlates with decision quality
      • H2: Decision-aware training outperforms metric-oriented training
      • H3: Oracle forecasting provides an upper bound
    • Experiment Script
    • Research Directions
      • 1. Ranking-Optimized Forecasting
      • 2. Direct Policy Evaluation
      • 3. Calibration-Aware Training
      • 4. Contextual Decision Quality
    • References
  • Multi-Agent Market Effect of AI Policy Adoption
    • Experiment
    • Results
    • Interpretation
    • Limitations (read before citing)
    • Follow-up directions

🔧 API Reference

  • Unified Data Loader
    • DataLoader
      • DataLoader.project_root
      • DataLoader.zone_count
      • DataLoader.slot_count
      • DataLoader.__init__()
      • DataLoader.datetime_to_state()
      • DataLoader.load_train_data()
      • DataLoader.load_travel_time_matrix()
      • DataLoader.load_zone_statistics()
      • DataLoader.next_half_hour()
  • Configuration System
    • get_config()
    • load_config()
    • reload_config()
  • Finite-Horizon Strategy
  • MDP Solver
    • MDPValueIteration
      • MDPValueIteration.q_values()
      • MDPValueIteration.recommend()
    • bellman_backup()
    • recommend()
  • Demo Gallery
    • Scenario: Rainy Friday Evening in Manhattan
      • The problem
      • AI decision process
      • The result
    • Benchmark Dashboard
      • Static diagnostic (3,360 queries)
      • Rollout performance
      • Multi-agent competition
    • NYC Map Visualization
    • Reproducibility
    • More scenarios
Urban Mobility Decision Intelligence
  • Demo Gallery
  • View page source

Demo Gallery

Visual walkthrough of the NYC Taxi Zone Recommendation platform.


Scenario: Rainy Friday Evening in Manhattan

The problem

A taxi driver finishes a drop-off in Midtown at 7:30 PM on a rainy Friday. Without guidance, the driver cruises randomly, wasting fuel and time.

Before AI guidance:

  • Driver circles Midtown for 20 minutes

  • Finds passenger heading to Brooklyn ($18 fare)

  • 35% utilization across the shift

  • ~$350 daily revenue

AI decision process

Step 1: Demand forecast → JFK airport demand spikes at 8 PM (rain + Friday)
Step 2: Travel time matrix → 35 min from Midtown to JFK via highway
Step 3: Two-step horizon planner → Go to JFK now, pick up airport fare, reposition to Manhattan
Step 4: Recommendation → [JFK Zone 132, Upper East Side Zone 140, Midtown Zone 161]

The result

After AI guidance:

  • Driver heads to JFK, picks up $62 airport fare within 10 minutes

  • Then repositions based on next forecast window

  • 52% utilization across the shift

  • ~\(570 daily revenue (+\)220/day vs cruising)


Benchmark Dashboard

Static diagnostic (3,360 queries)

NDCG comparison

The Two-Step Horizon strategy achieves 0.9565 NDCG@3 on public validation queries — a 21.9% improvement over the naive Hot Zone baseline.

Rollout performance

Pickup comparison

100-seed paired rollout shows consistent improvement: Two-Step delivers +$139/day vs Hot Zone, with the improvement concentrated in the 7-10 AM and 6-9 PM peak windows.

Multi-agent competition

The multi-agent simulator reveals that as fleet size grows, simpler strategies degrade faster than horizon-aware policies. At 50 drivers, Two-Step maintains a +$25/day advantage over Single-Step.


NYC Map Visualization

The project uses NYC’s 263-taxi-zone geography. All-pairs travel times are precomputed via Dijkstra on a directed OD graph built from 1.8M+ training trips.

        Manhattan (Zones 100-199)
        ┌──────────────────────────┐
        │  Upper East   Upper West │
        │  ┌────┬────┐ ┌────┬────┐ │
        │  │140 │141 │ │142 │143 │ │
        │  └────┴────┘ └────┴────┘ │
        │  Midtown                 │
        │  ┌────┬────┐             │
        │  │161 │162 │   → Queens  │
        │  └────┴────┘      (JFK)  │
        │  Downtown          ↓     │
        │  ┌────┬────┐     ┌───┐   │
        │  │113 │114 │     │132│   │
        │  └────┴────┘     └───┘   │
        └──────────────────────────┘

Reproducibility

All results in this gallery are reproducible. Run:

make all          # Full pipeline
make static       # Static diagnostic only
make combined-benchmark  # Combined report

All metrics are checked in as reference snapshots in outputs/ with timestamped validation.


More scenarios

Scenario

Time

Weather

AI Decision

Outcome

Monday morning rush

8:15 AM

Clear

Upper East → Midtown

Commuter demand

Saturday night

11:00 PM

Clear

Greenwich Village → Meatpacking

Nightlife flow

Airport surge

4:00 PM

Thunderstorm

JFK → Manhattan

Airport backlog

Holiday eve

6:00 PM

Snow

Penn Station → residential

Transit hub exit

Previous

© Copyright 2026, Zefan Cai.

Built with Sphinx using a theme provided by Read the Docs.