Sign in

Modular: Inference from Kernel to Cloud [Unofficial]

@modular.com.web.brid.gy
10 followers 0 following 236 posts

The unified AI inference stack - from custom GPU kernels to production cloud serving on NVIDIA and AMD. 2x performance. Top open models. Open source stack. 🌉 bridged from 🌐 modular.com: fed.brid.gy/web/modular.com

PostsRepliesMedia
Modular: Inference from Kernel to Cloud [Unofficial] @modular.com.web.brid.gy · 17/09/2026
modular.com
Modular: Modular 26.6: Open compiler contributions, audio generation, and expanded model support
Modular 26.6: Open compiler contributions, audio generation, and expanded model support
000
Modular: Inference from Kernel to Cloud [Unofficial] @modular.com.web.brid.gy · 18/08/2026
modular.com
Modular: Mojo🔥 is now open source!
Mojo🔥 is now open source!
000
Modular: Inference from Kernel to Cloud [Unofficial] @modular.com.web.brid.gy · 18/08/2026
modular.com
Modular: ModCon 2026: Open source, open cloud, open silicon
ModCon 2026: Open source, open cloud, open silicon
000
Modular: Inference from Kernel to Cloud [Unofficial] @modular.com.web.brid.gy · 29/07/2026
modular.com
Modular: Qualcomm Completes Acquisition of Modular
Qualcomm Completes Acquisition of Modular
000
Modular: Inference from Kernel to Cloud [Unofficial] @modular.com.web.brid.gy · 24/06/2026
modular.com
Modular: Qualcomm to Acquire Modular
Qualcomm to Acquire Modular
000
Modular: Inference from Kernel to Cloud [Unofficial] @modular.com.web.brid.gy · 18/06/2026
modular.com
Modular: Modular 26.4: SOTA MoE Serving, Model Bringup via Agent Skills, Mojo 1.0 Beta 2 and More
Modular 26.4: SOTA MoE Serving, Model Bringup via Agent Skills, Mojo 1.0 Beta 2 and More
000
Modular: Inference from Kernel to Cloud [Unofficial] @modular.com.web.brid.gy · 17/06/2026
modular.com
Modular: ModCon 2026: Modular’s Developer Conference
ModCon 2026: Modular’s Developer Conference
000
Modular: Inference from Kernel to Cloud [Unofficial] @modular.com.web.brid.gy · 11/06/2026
modular.com
Modular: Day Zero: MiniMax M3 Open Weights on Modular Cloud
Day Zero: MiniMax M3 Open Weights on Modular Cloud
000
Modular: Inference from Kernel to Cloud [Unofficial] @modular.com.web.brid.gy · 10/06/2026
modular.com
Modular: Modverse #55: Mojo 1.0 Beta, Community Mojo Libraries, and Real-Time Patient Conversations Powered by MAX
Modverse #55: Mojo 1.0 Beta, Community Mojo Libraries, and Real-Time Patient Conversations Powered by MAX
000
Modular: Inference from Kernel to Cloud [Unofficial] @modular.com.web.brid.gy · 05/06/2026
modular.com
Modular: Why LLM Inference Needs a New Kind of Router - Part 3
Why LLM Inference Needs a New Kind of Router - Part 3
000
Modular: Inference from Kernel to Cloud [Unofficial] @modular.com.web.brid.gy · 29/05/2026
modular.com
Modular: Three trends from MLSys 2026
Three trends from MLSys 2026
000
Modular: Inference from Kernel to Cloud [Unofficial] @modular.com.web.brid.gy · 21/05/2026
modular.com
Modular: Why LLM Inference Needs a New Kind of Router - Part 2
Why LLM Inference Needs a New Kind of Router - Part 2
000
Modular: Inference from Kernel to Cloud [Unofficial] @modular.com.web.brid.gy · 19/05/2026
modular.com
Modular: How I built a pure Mojo app (and 10 libraries) with AI agents
How I built a pure Mojo app (and 10 libraries) with AI agents
000
Modular: Inference from Kernel to Cloud [Unofficial] @modular.com.web.brid.gy · 18/05/2026
modular.com
Modular: Hippocratic AI partners with Modular to power flexible, high-quality inference for real-time patient conversations
Hippocratic AI partners with Modular to power flexible, high-quality inference for real-time patient conversations
000
Modular: Inference from Kernel to Cloud [Unofficial] @modular.com.web.brid.gy · 13/05/2026
modular.com
Modular: Translating to Mojo via AI Agents
Translating to Mojo via AI Agents
000
Modular: Inference from Kernel to Cloud [Unofficial] @modular.com.web.brid.gy · 12/05/2026
modular.com
Modular: Inkwell: Why Your Inference Platform Matters As Much As Your Model
Inkwell: Why Your Inference Platform Matters As Much As Your Model
000
Modular: Inference from Kernel to Cloud [Unofficial] @modular.com.web.brid.gy · 08/05/2026
modular.com
Modular: Why LLM Inference Needs a New Kind of Router - Part 1
Why LLM Inference Needs a New Kind of Router - Part 1
000
Modular: Inference from Kernel to Cloud [Unofficial] @modular.com.web.brid.gy · 07/05/2026
modular.com
Modular: Modular 26.3: Mojo 1.0 Beta, MAX Video Gen, and more
Modular 26.3: Mojo 1.0 Beta, MAX Video Gen, and more
000
Modular: Inference from Kernel to Cloud [Unofficial] @modular.com.web.brid.gy · 04/05/2026
modular.com
Modular: Modverse #54: AMD AI DevDay, New Modular Offices, and a Community That Keeps Shipping
Modverse #54: AMD AI DevDay, New Modular Offices, and a Community That Keeps Shipping
010
Modular: Inference from Kernel to Cloud [Unofficial] @modular.com.web.brid.gy · 16/04/2026
modular.com
Modular: How Frontier Coding Agents Built a Video Diffusion Pipeline on MAX
How Frontier Coding Agents Built a Video Diffusion Pipeline on MAX
000
Modular: Inference from Kernel to Cloud [Unofficial] @modular.com.web.brid.gy · 13/04/2026
modular.com
Modular: TileTensor Part 1 - Safer, More Efficient GPU Kernels
TileTensor Part 1 - Safer, More Efficient GPU Kernels
000
Modular: Inference from Kernel to Cloud [Unofficial] @modular.com.web.brid.gy · 10/04/2026
modular.com
Modular: Modular Opens Edinburgh & San Francisco Offices
Modular Opens Edinburgh & San Francisco Offices
000
Modular: Inference from Kernel to Cloud [Unofficial] @modular.com.web.brid.gy · 02/04/2026
modular.com
Modular: Day Zero Launch: Fastest Performance for Gemma 4 on NVIDIA and AMD
Day Zero Launch: Fastest Performance for Gemma 4 on NVIDIA and AMD
000
Modular: Inference from Kernel to Cloud [Unofficial] @modular.com.web.brid.gy · 31/03/2026
modular.com
Modular: Modverse #54: From GTC to Edinburgh, a Community Building Momentum
Modverse #54: From GTC to Edinburgh, a Community Building Momentum
010
Modular: Inference from Kernel to Cloud [Unofficial] @modular.com.web.brid.gy · 30/03/2026
modular.com
Modular: Software Pipelining for GPU Kernels: Part 1 - The Pipeline Problem
Software Pipelining for GPU Kernels: Part 1 - The Pipeline Problem
000
Modular: Inference from Kernel to Cloud [Unofficial] @modular.com.web.brid.gy · 26/03/2026
modular.com
Modular: Structured Mojo Kernels Part 3 - Composition in Practice
Structured Mojo Kernels Part 3 - Composition in Practice
010
Modular: Inference from Kernel to Cloud [Unofficial] @modular.com.web.brid.gy · 19/03/2026
modular.com
Modular: Modular 26.2: State-of-the-Art Image Generation and Upgraded AI Coding with Mojo
Modular 26.2: State-of-the-Art Image Generation and Upgraded AI Coding with Mojo
010
Modular: Inference from Kernel to Cloud [Unofficial] @modular.com.web.brid.gy · 16/03/2026
modular.com
Modular: Modular at NVIDIA GTC 2026: MAX on Blackwell, Mojo Kernel Porting, and DeepSeek V3 on B200
Modular at NVIDIA GTC 2026: MAX on Blackwell, Mojo Kernel Porting, and DeepSeek V3 on B200
010
Modular: Inference from Kernel to Cloud [Unofficial] @modular.com.web.brid.gy · 11/03/2026
modular.com
Modular: Structured Mojo Kernels Part 2 - The Three Pillars
Structured Mojo Kernels Part 2 - The Three Pillars
010
Modular: Inference from Kernel to Cloud [Unofficial] @modular.com.web.brid.gy · 06/03/2026
modular.com
Modular: Modverse #53: Community Builds, Research Milestones, and a Growing Ecosystem
Modverse #53: Community Builds, Research Milestones, and a Growing Ecosystem
000
Modular: Inference from Kernel to Cloud [Unofficial] @modular.com.web.brid.gy · 04/03/2026
modular.com
Modular: Structured Mojo Kernels Part 1 - Peak Performance, Half the Code
Structured Mojo Kernels Part 1 - Peak Performance, Half the Code
010
Modular: Inference from Kernel to Cloud [Unofficial] @modular.com.web.brid.gy · 27/02/2026
modular.com
Modular: Structured Mojo Kernels Part 1 - Why Structured Kernels?
Structured Mojo Kernels Part 1 - Why Structured Kernels?
000
Modular: Inference from Kernel to Cloud [Unofficial] @modular.com.web.brid.gy · 18/02/2026
modular.com
Modular: The Claude C Compiler: What It Reveals About the Future of Software
The Claude C Compiler: What It Reveals About the Future of Software
000
Modular: Inference from Kernel to Cloud [Unofficial] @modular.com.web.brid.gy · 10/02/2026
modular.com
Modular: BentoML Joins Modular
BentoML Joins Modular
010
Modular: Inference from Kernel to Cloud [Unofficial] @modular.com.web.brid.gy · 05/02/2026
modular.com
Modular: The Five Eras of KVCache
The Five Eras of KVCache
010
Modular: Inference from Kernel to Cloud [Unofficial] @modular.com.web.brid.gy · 29/01/2026
modular.com
Modular: Modular 26.1: A Big Step Towards More Programmable and Portable AI Infrastructure
Modular 26.1: A Big Step Towards More Programmable and Portable AI Infrastructure
010
Modular: Inference from Kernel to Cloud [Unofficial] @modular.com.web.brid.gy · 14/01/2026
modular.com
Modular: How to Beat Unsloth's CUDA Kernel Using Mojo—With Zero GPU Experience
How to Beat Unsloth's CUDA Kernel Using Mojo—With Zero GPU Experience
010
Modular: Inference from Kernel to Cloud [Unofficial] @modular.com.web.brid.gy · 12/01/2026
modular.com
Modular: How I Beat Unsloth's CUDA Kernel Using Mojo—With Zero GPU Experience
How I Beat Unsloth's CUDA Kernel Using Mojo—With Zero GPU Experience
010
Modular: Inference from Kernel to Cloud [Unofficial] @modular.com.web.brid.gy · 19/12/2025
modular.com
Modular: 🔥 Modular 2025 Year in Review
🔥 Modular 2025 Year in Review
010
Modular: Inference from Kernel to Cloud [Unofficial] @modular.com.web.brid.gy · 05/12/2025
modular.com
Modular: The path to Mojo 1.0
The path to Mojo 1.0
010
Modular: Inference from Kernel to Cloud [Unofficial] @modular.com.web.brid.gy · 03/12/2025
modular.com
Modular: Modverse #52: Advancing AI Together — Community Projects & Platform Milestones
Modverse #52: Advancing AI Together — Community Projects & Platform Milestones
000
Modular: Inference from Kernel to Cloud [Unofficial] @modular.com.web.brid.gy · 20/11/2025
modular.com
Modular: Modular 25.7: Faster Inference, Safer GPU Programming, and a More Unified Developer Experience
Modular 25.7: Faster Inference, Safer GPU Programming, and a More Unified Developer Experience
010
Modular: Inference from Kernel to Cloud [Unofficial] @modular.com.web.brid.gy · 07/11/2025
modular.com
Modular: "TTS 1 Max" (powered by Modular Platform) Ranked #1 Speech Model on Artificial Analysis
"TTS 1 Max" (powered by Modular Platform) Ranked #1 Speech Model on Artificial Analysis
010
Modular: Inference from Kernel to Cloud [Unofficial] @modular.com.web.brid.gy · 06/11/2025
modular.com
Modular: PyTorch and LLVM in 2025 — Keeping up With AI Innovation
PyTorch and LLVM in 2025 — Keeping up With AI Innovation
010
Modular: Inference from Kernel to Cloud [Unofficial] @modular.com.web.brid.gy · 17/10/2025
modular.com
Modular: Achieving State-of-the-Art Performance on AMD MI355 — in Just 14 Days
Achieving State-of-the-Art Performance on AMD MI355 — in Just 14 Days
010
Modular: Inference from Kernel to Cloud [Unofficial] @modular.com.web.brid.gy · 24/09/2025
modular.com
Modular: Modular Raises $250M to scale AI's Unified Compute Layer
Modular Raises $250M to scale AI's Unified Compute Layer
010
Modular: Inference from Kernel to Cloud [Unofficial] @modular.com.web.brid.gy · 22/09/2025
modular.com
Modular: Modular 25.6: Unifying the latest GPUs from NVIDIA, AMD, and Apple
Modular 25.6: Unifying the latest GPUs from NVIDIA, AMD, and Apple
010
Modular: Inference from Kernel to Cloud [Unofficial] @modular.com.web.brid.gy · 19/09/2025
modular.com
Modular: Matrix Multiplication on Blackwell: Part 4 - Breaking SOTA
Matrix Multiplication on Blackwell: Part 4 - Breaking SOTA
010
Modular: Inference from Kernel to Cloud [Unofficial] @modular.com.web.brid.gy · 12/09/2025
modular.com
Modular: Matrix Multiplication on Blackwell: Part 3 - The Optimizations Behind 85% of SOTA Performance
Matrix Multiplication on Blackwell: Part 3 - The Optimizations Behind 85% of SOTA Performance
010
Modular: Inference from Kernel to Cloud [Unofficial] @modular.com.web.brid.gy · 05/09/2025
modular.com
Modular: Matrix Multiplication on Blackwell: Part 2 - Using Hardware Features to Optimize Matmul
Matrix Multiplication on Blackwell: Part 2 - Using Hardware Features to Optimize Matmul
010