LLM cloud inference dominates usage, but should it? Local models and accelerators have improved massively over recent years.
Perfect routing to best local model "reduce energy consumption by 80.4%, compute by 77.3%, and cost by 73.8% versus cloud-only deployment"
arxiv.org/pdf/2511.07885