When inference demand spikes, auctioning GPU priority to the highest bidder feels like textbook economics. But unconstrained bidding thrashes the KV cache and blows up latency by 12x. New work from Berkeley shows how to run priority auctions inside radix trees.
deanlee.info
The Inference Auction
Why allocating GPU priority through naive financial bidding breaks KV cache locality, and how mechanism design has to adapt to radix trees.