Great open-source communities push enterprise hardware further. 🤝
New work from IBM Research & Red Hat on llm-d:
• 753B open model on 544 NVIDIA H100 GPUs
• 5–10x lower cost per token vs commercial APIs
• Serves thousands of concurrent agents
Blog: research.ibm.com/blog/running...
research.ibm.com
How llm-d makes the most of the hardware you already have
IBM Research and Red Hat deployed a 753B open model on H100 GPUs, serving thousands of concurrent coding agents at 5-10x lower cost than commercial APIs.