Sign in

Jonathan Ross

@jonathan-ross.bsky.social
998 followers 16 following 42 posts

CEO + Founder @ Groq, the Most Popular API for Fast Inference | Creator of the TPU and LPU, Two of the World’s Most Important AI Chips | On a Mission to Double the World's AI Compute by 2027

PostsRepliesMedia
Reposted by Jonathan Ross
Peter Bordes @peterbordes.bsky.social · 25/09/2025
Fantastic insight on the massive demand for AI inference infrastructure “The demand for AI compute is insatiable” @groq.com CEO @jonathan-ross.bsky.social, “Our mission is to provide over half of the world’s inference compute” - @cnbc.com cnb.cx/4nG7Pcm #AI
cnb.cx
Groq CEO: Our mission is to provide over half of the world’s inference compute
Jonathan Ross, CEO and founder of Groq, joins CNBC’s 'Squawk on the Street' to discuss the AI chip startup’s $750 million funding round, its push to deliver faster, lower-cost inference chips, and why...
054
Jonathan Ross @jonathan-ross.bsky.social · 06/09/2025
Founder Tip #2: You have to spend time to make time. Hiring, re-organizing, calendar clean up (across the team), preparation for meetings (internal and external), etc. Half my day is available for whatever I find important - because the other half is spent freeing up time.
021
Jonathan Ross @jonathan-ross.bsky.social · 19/08/2025
Clearly China doesn't have enough compute for scaled AI today: - GPT-OSS, Llama [US]: optimized for cheaper inference - R1, Kimi K2, Qwen [China]: optimized for cheaper training With China's population reducing inference costs is more important, and that means more training.
021
Reposted by Jonathan Ross
AI SDK @handle.invalid · 16/04/2025
Transcribe audio with @groq.com.
152
Jonathan Ross @jonathan-ross.bsky.social · 24/03/2025
I spent the weekend hanging out with a group of friends. A question we asked was what dreams did we have that we gave up on? When I was 18, I had two dreams: 1) Be an astronaut 2) Build AI chips I didn’t give up on one of them. 😀
020
Reposted by Jonathan Ross
Groq Inc. @groq.com · 27/02/2025
Big news! Mistral AI Saba 24B is on GroqCloud! The specialized regional language model is perfect for Middle East and South Asia-based devs and enterprises building AI solutions that need fast inference. Learn more: groq.com/mistral-saba...
hubs.la
Mistral Saba Added to GroqCloud™ Model Suite - Groq is Fast AI Inference
GroqCloud™ has added another openly-available model to our suite – Mistral Saba. Mistral Saba is Mistral AI’s first specialized regional language model,
171
Jonathan Ross @jonathan-ross.bsky.social · 17/02/2025
YouTube: www.youtube.com/watch?v=xBMR... Spotify: open.spotify.com/episode/30np... Try Groq: console.groq.com
youtube.com
Jonathan Ross, Founder & CEO @ Groq: NVIDIA vs Groq - The Future of Training vs Inference | E1260
YouTube video by 20VC with Harry Stebbings
161
Jonathan Ross @jonathan-ross.bsky.social · 17/02/2025
It was a pleasure being back on 20VC with Harry Stebbings. His craft of interviewing is second to none and we went deep. This is the interview after we just launched 19,000 LPUs in Saudi Arabia. We built the largest inference cluster in the region. Link to the interview in the comments below!
4657
Jonathan Ross @jonathan-ross.bsky.social · 09/02/2025
We built the region’s largest inference cluster in Saudi Arabia in 51 days and we just announced a $1.5B agreement for Groq to expand our advanced LPU-based AI inference infrastructure. Build fast.
081
Jonathan Ross @jonathan-ross.bsky.social · 01/02/2025
media.tenor.com
a close up of a man 's face with the words inconceivable written on it .
ALT: a close up of a man 's face with the words inconceivable written on it .
020
Jonathan Ross @jonathan-ross.bsky.social · 29/01/2025
My emergency episode with @harrystebbings.bsky.social at 20VC just launched on the impact of #DeepSeek on the AI world
715431
Reposted by Jonathan Ross
Yoshua Bengio @yoshuabengio.bsky.social · 23/01/2025
Yesterday at the World Economic Forum in Davos, I joined a constructive discussion on AGI alongside @andrewyng.bsky.social, @yejinchoinka.bsky.social, @jonathan-ross.bsky.social , @thomwolf.bsky.social and moderator @nxthompson.bsky.social. Full discussion here: www.weforum.org/meetings/wor...
1476
Jonathan Ross @jonathan-ross.bsky.social · 13/01/2025
0142
Jonathan Ross @jonathan-ross.bsky.social · 09/01/2025
Thank you! 🙏
030
Jonathan Ross @jonathan-ross.bsky.social · 08/01/2025
Over the next decade, we want to drive the cost down for generative AI 1,000x making a lot more activities profitable. And we think that that will cause a 100x spend increase. 🧵(5/5)
080
Jonathan Ross @jonathan-ross.bsky.social · 08/01/2025
Over the last 60 years, almost like clockwork, every decade compute gets about 1000x cheaper, people buy 100,000x as much of it, spending 100x times more overall.  Our mission at Groq is to drive the cost of compute towards zero;The cheaper we make compute the more people spend. 🧵(4/5)
150
Jonathan Ross @jonathan-ross.bsky.social · 08/01/2025
- The answer is when you make a steam engine more efficient, it reduces the OpEx; - When you reduce the OpEx, it increases the number of activities that are profitable; - Therefore, people will do more things using steam engines and coal demand rises. The same paradox applies to compute. 🧵(3/5)
120
Jonathan Ross @jonathan-ross.bsky.social · 08/01/2025
It’s a paradox because if they're more efficient, why are they buying more coal? 🧵(2/5)
110
Jonathan Ross @jonathan-ross.bsky.social · 08/01/2025
When you make compute cheaper do people buy more? Yes. It's called Jevons Paradox and it's a big part of our business thesis. In the 1860s, an Englishman wrote a treatise on coal where he noted that every time steam engines got more efficient people bought more coal. 🧵(1/5)
191
Jonathan Ross @jonathan-ross.bsky.social · 07/01/2025
This is insane, Groq is the #4 API on this list! 😮 OpenAI, Anthropic, and Azure are the top 3 LLM API providers on LangChain Groq is #4, and close behind Azure Google, Amazon, Mistral, and Hugging Face are the next 4. Ollama is for local development. Now add three more 747's worth of LPUs 😁
1162
Jonathan Ross @jonathan-ross.bsky.social · 05/01/2025
www.youtube.com/watch?v=HxNU...
youtube.com
2025 Predictions with bestie Gavin Baker
YouTube video by All-In Podcast
010
Jonathan Ross @jonathan-ross.bsky.social · 05/01/2025
Groq just got a shout out on the All-In pod as one of the big winners for 2025 alongside Nvidia. It’s the year of the AI chip and ours is the fastest 😃
150
Jonathan Ross @jonathan-ross.bsky.social · 24/12/2024
Welcome to Shipmas - Groq Style. Groq's second B747 this week. How many LPUs and GroqRacks can we load into a jumbo jet? Take a look. Have you been naughty or nice?
1121
Jonathan Ross @jonathan-ross.bsky.social · 23/12/2024
Santa rented two full 747s this week to make his holiday deliveries of GroqRacks. Ho ho ho! 🎅
2182
Jonathan Ross @jonathan-ross.bsky.social · 10/12/2024
(5/5) Learning: product-led growth works; even when your product is too large and expensive to let people have it for free, you just have to be more creative about it.
020
Jonathan Ross @jonathan-ross.bsky.social · 10/12/2024
(4/5) We're not shipping anyone millions of dollars of hardware. It’s not a big ask for them to try it. And when they try it, they love it.
120
Jonathan Ross @jonathan-ross.bsky.social · 10/12/2024
(3/5) By making Groq easy and low cost to try, we got 60,000 developers on our developer console in 30 days. Less than a year after that, and we're at 645,000 developers and growing.
210
Jonathan Ross @jonathan-ross.bsky.social · 10/12/2024
(2/5) That makes it almost impossible to do counterintuitive things. Like, “Try this new chip called an LPU", when everything in the zeitgeist is talking about GPUs. And if you're a startup? Forget it. That realization is why we made the strategic decision to put up our own cloud.
100
Jonathan Ross @jonathan-ross.bsky.social · 10/12/2024
(1/5) One of the reasons why chips are so hard to innovate in is because if you're asking someone to put up a 10 million, 100 million, or a billion dollar check they need to know that what they're buying is going to work.
271
Jonathan Ross @jonathan-ross.bsky.social · 06/12/2024
(5/5) The new Llama-3.3-70B model launched and is now available to all 645,000 GroqCloud™ developers as of this morning. Go cook, and don't forget to share what you build here. Thank you for making GroqCloud™ the #1 API for fast inference! This is only just the beginning.
041
Jonathan Ross @jonathan-ross.bsky.social · 06/12/2024
(4/5) It's also significantly less expensive and faster than the larger model. Meta continues to push the lead in open weight innovations, and is keeping the pressure high for proprietary model providers to attempt to keep ahead of the giant wave of open.
131
Jonathan Ross @jonathan-ross.bsky.social · 06/12/2024
(3/5) Though ~1/5th the size of Llama 3.1-405B, benchmarking shows Llama-3.3-70B performing neck and neck, and in many crucial cases substantially out performing the larger model (Instruction Following, Coding, Math, etc.), a suitable replacement for a majority of workloads.
111
Jonathan Ross @jonathan-ross.bsky.social · 06/12/2024
(2/5) Today, our partner Meta released its latest version of Llama-3.3-70B-Instruct. And to all those who speculated that the industry had hit the wall - maybe some have, but Meta hasn’t yet. 😉 This is a big deal.
131
Jonathan Ross @jonathan-ross.bsky.social · 06/12/2024
(1/5) "The reports of the LLM scaling laws' demise have been greatly exaggerated." techcrunch.com/2024/12/06/m...
143
Jonathan Ross @jonathan-ross.bsky.social · 28/11/2024
(5/5) In time we will look at intelligence the same way we look at the universe, we will be inspired by how vast and deep it is, and we won't be afraid.
030
Jonathan Ross @jonathan-ross.bsky.social · 28/11/2024
(4/5) When you look at something that's rapidly becoming that grand, that vast, and that capable it makes you feel small and afraid. Over time we've gotten used to this idea that the universe is large and that we're small. If anything we now see the beauty in how vast the universe is.
130
Jonathan Ross @jonathan-ross.bsky.social · 28/11/2024
(3/5) They did that because they were afraid. If people weren't the center of the universe that made us feel small. Large language models are the telescope of the mind. They give us this view of an intelligence that is quickly becoming larger and more vast than any single human being.
100
Jonathan Ross @jonathan-ross.bsky.social · 28/11/2024
(2/5) He got in trouble because he made a telescope, pointed it at the stars, and using data from other astronomers he put together a pretty convincing case that human beings weren't the center of the universe. They locked him in a tower.
100
Jonathan Ross @jonathan-ross.bsky.social · 28/11/2024
(1/5) The question I get asked a lot is, “Should I be afraid of AI?” There was this guy who got in a lot of trouble once, his name was Galileo.
293
Jonathan Ross @jonathan-ross.bsky.social · 26/11/2024
(5/5) That's why we expect that by the start of next year, we will be one of the most significant players in our space. And by the end of next year, we will provide more than half the world's generative AI inference compute.
040
Jonathan Ross @jonathan-ross.bsky.social · 26/11/2024
(4/5) The reason for that number? 25 million? If we do 25 million tokens per second then we're going to have as much compute capacity as a hyperscaler had at the start of this year. And from that point forward, we'll just ramp.
120
Jonathan Ross @jonathan-ross.bsky.social · 26/11/2024
(3/5) It may sound silly to distill our north star down to such a simple message but that is the hardest part about leadership. Getting alignment is finding that singular thing, that if you do it well, you're going to be successful.
100
Jonathan Ross @jonathan-ross.bsky.social · 26/11/2024
(2/5) Okay fine, cute prop. But how does a challenge coin create alignment? Well, if someone makes a proposal and it doesn't help with these priorities, anyone can just put their coin on the table and ask “how does it help with this?”
100
Jonathan Ross @jonathan-ross.bsky.social · 26/11/2024
(1/5) Everyone at Groq has one of these challenge coins on them. It’s how we create alignment. One side says its 25 million, because we're going to get to 25 million tokens per second by the end of the year On the other side, it says, “Make it real. Make it now. Make it wow.”
160
Reposted by Jonathan Ross
Rick Lamers @ricklamers.bsky.social · 23/11/2024
Image generation is just TOO MUCH FUN! Fast prompt generation with Groq ✅ Fast image generation with Fal.ai ✅ Open Source (MIT) ✅ ⚙️ pip install pyimagen
153
Jonathan Ross @jonathan-ross.bsky.social · 23/11/2024
Wow. How did this happen, and how do we keep it happening?
120
Jonathan Ross @jonathan-ross.bsky.social · 23/11/2024
What can you do with Llama quality and Groq speed? Instant. That's what. 3 months back: Llama 8B running at 750 Tokens/sec Now: Llama 70B model running at 3,200 Tokens/sec We're still going to get a liiiiiiitle bit faster, but this is our V1 14nm LPU - how fast will V2 be? 😉
1191