Will you explore client-side offloading? Some people have PCs which can actually run certain LLM models. There's a bunch of open source solutions that could help you implement this.
Hosting LLMs is expensive. Offloading could be a win-win solution for both Proton and the users. I have solar power.