The team built a front-end classifier that directs 95% of traffic to a lightweight model, adds a Redis cache for repeat answers (including “I don’t know”), compresses prompts and batches background jobs, trimming monthly costs from $420 to $28 without hurting latency.
By routing 85% of queries to cheaper models, caching repeat prompts, and compressing inputs, the team slashed monthly spend from $420 to $28 while keeping latency low and user experience intact.
The new architecture treats rover actions as language predictions, letting the robot choose the most valuable rock to sample while staying within safety limits, all logged in an immutable audit trail for engineers on Earth.
Ramp’s Router lets developers set “strategies” that route queries by benchmark, complexity or provider usage tier, while a built-in dashboard logs token spend, latency and fallback attempts, turning AI inference into a line-item expense.
The paper introduces two variants—Pandora’s Router, a centralized cost-aware selector, and Pandora’s Bidder, a self-bidding marketplace—that automatically adjust inspection intensity based on the price of evaluating each model.