Sign in

Daron Yondem

@daron.me
215 followers 111 following 749 posts

Tech Lead at Microsoft | Applied AI Expert | Ex-CTO of SaaS Startups

PostsRepliesMedia
Daron Yondem @daron.me · 06/08/2025
- 128K context, advanced reasoning, top-tier price-performance (3× cheaper than comparable Gemini, 5× than DeepSeek-R1. 👉 Official announcement www.aboutamazon.com/news/aws/ope... ! Here’s to many more launches (accidental or otherwise). Views are my own. #AWS #OpenAI #GenerativeAI #Bedrock
aboutamazon.com
OpenAI open weight models available today on AWS
Customers will have access to advanced open weight models in Amazon Bedrock and Amazon SageMaker AI, bringing together OpenAI and the world’s most comprehensive and broadly adopted cloud.
000
Daron Yondem @daron.me · 06/08/2025
but I’m happy to take (mostly fictional) credit for “bringing OpenAI with me.” 😅 Why this drop is a big deal? - Builder freedom: open weights = fine-tune, inspect, self-host as needed. - Enterprise muscle: Bedrock + SageMaker guardrails, security, and scale.
100
Daron Yondem @daron.me · 06/08/2025
🚀 Apparently TSA missed the “3-oz limit on AI models” in my carry-on. #dadjoke #staywithme Thirty days after joining AWS, OpenAI’s brand-new open-weight models, gpt-oss-120b and gpt-oss-20b, just landed in Amazon Bedrock and SageMaker. Total coincidence…
111
Daron Yondem @daron.me · 28/07/2025
Keep chasing what matters, friends. See you on the course, or wherever your “next try” takes you. 💪 #SFMarathon #PRDay #ComebackStory #KeepGoing #RunningCommunity
100
Daron Yondem @daron.me · 28/07/2025
Attaching my race map and official time for anyone who loves the nerdy details 🙂
100
Daron Yondem @daron.me · 28/07/2025
If you’re staring at a goal that got derailed, injury, rejection, life-curveball, remember: one “no” doesn’t erase the “yes” that’s still possible.
100
Daron Yondem @daron.me · 28/07/2025
Celebrate the comeback. Today’s PR isn’t just a faster time; it’s proof that persistence + patience pay off.
100
Daron Yondem @daron.me · 28/07/2025
Progress isn’t linear. The road back was full of frustrating rehab days, slow runs, and doubts, but every small step stacked up.
100
Daron Yondem @daron.me · 28/07/2025
Re-evaluate, don’t quit. After the injury I asked myself if the goal still mattered. It did, so the plan changed, not the purpose.
100
Daron Yondem @daron.me · 28/07/2025
Setbacks aren’t stop signs. Missing the 2023 start line didn’t mean the story was over, it just meant there was another chapter to write.
100
Daron Yondem @daron.me · 28/07/2025
Fast-forward to this morning: same city, same Golden Gate views, but a very different ending. Crossing that finish line wasn’t just about a medal or a stopwatch; it was a reminder.
100
Daron Yondem @daron.me · 28/07/2025
Finished the San Francisco Marathon, and set a new PR:4 h 20 m! 🎉 Two years ago I stood on these same streets with a race bib and a busted leg, watching everyone else start while I sat on the sidelines. It hurt 😞 literally and figuratively. I’d trained, I’d planned, and life still said “not today.”
120
Daron Yondem @daron.me · 22/07/2025
AI isn’t a single “best” model, it’s a toolbox. Pick the wrench that fits the bolt, and watch your efficiency (and your API bill) improve. Are you experimenting with model size in your workflow? I’d love to hear what’s working for you. 👇 #AI #LLM #Productivity #Sustainability #TechInsights
001
Daron Yondem @daron.me · 22/07/2025
- Right tool, right job. Reserve the juggernauts (O3 Pro et al.) for truly complex reasoning. For routine writing, data cleanup, or quick ideation, a lighter model is the productivity hack no one talks about.
100
Daron Yondem @daron.me · 22/07/2025
- Iterative > “perfect.” Fast back-and-forth lets me steer, refine, and co-create. Smaller models aren’t worse, they just leave more room for human intuition to tie everything together.
100
Daron Yondem @daron.me · 22/07/2025
- Cost & carbon matter. Bigger models draw more compute, more electricity, and more dollars, often to draft a quick email or brainstorm titles. That resource mismatch adds up.
100
Daron Yondem @daron.me · 22/07/2025
- Speed fuels flow. Waiting 15–30 seconds for a response breaks the creative feedback loop. Snappier, lightweight models keep the conversation, and my momentum alive.
100
Daron Yondem @daron.me · 22/07/2025
But after a few months, I noticed something curious: for 80 % of my daily tasks, the largest model was actually slowing me down. Here’s what I’ve learned:
100
Daron Yondem @daron.me · 22/07/2025
💡 Bigger isn’t always better: why I’m reaching for smaller models in my everyday work When I first got access to O3, its raw power blew me away, I honestly thought I’d never settle for anything less.
100
Daron Yondem @daron.me · 10/07/2025
🔥 Deal alert: The first 10 people to snag a standard ticket can take 50 % off with code DARON50. If you’re building with open-source LLMs, or want a smarter, cheaper, faster stack, this is the summit to bookmark. Direct link to the site drn.fyi/pyk 😉 #DeepSeek #LLM #OpenSourceAI #GenerativeAI
drn.fyi
DeepSeek Demystified Summit
One day Summit. All access. Go beyond the buzz with expert-led sessions and a hands-on workshop built for real-world AI builders.
000
Daron Yondem @daron.me · 10/07/2025
• Real-world deployment playbooks: caching, quantization, rate-limit tricks—the stuff we all Google at 2 AM. Line-up: Paul Iusztin, Duarte O.Carmo, Karl Zhao, PhD, Miguel Otero Pedrido, Alex S. and more. When: 6 AM – 1:30 PM PDT / 9 AM – 4:30 PM ET, Aug 16 Where: 100 % online (live + replay)
110
Daron Yondem @daron.me · 10/07/2025
• Expert sessions that get past the buzzwords—think data security, model selection (DeepSeek-R1 vs V3), and agentic AI workflows. • A hands-on fine-tuning lab (LoRA + Unsloth) you can run on a consumer GPU.
100
Daron Yondem @daron.me · 10/07/2025
On Saturday, August 16 I’m trading weekend plans for something way more fun: the DeepSeek Demystified Summit, a one-day, all-access dive into the open-source LLM that’s making waves across cost, performance, and reasoning. Why I’m excited:
100
Daron Yondem @daron.me · 08/07/2025
Playground: chat.inceptionlabs.ai
000
Daron Yondem @daron.me · 08/07/2025
Diffusion isn’t just for images anymore. With Mercury shipping and giants circling, the next generation of language models may be noisy under the hood—but the output is crystal clear to users. 🧑🚀✨ #LLM #DiffusionModels #AI #NLP #DataScience
110
Daron Yondem @daron.me · 08/07/2025
Want to see it? I’m dropping the playground link + a fun speed-up video (video effect exaggerated for Twitter attention 😉).
100
Daron Yondem @daron.me · 08/07/2025
Fine print: • Training is still costlier per token than AR cousins • Stream UX isn’t char-by-char, Mercury shows a draft, then refines (UI tweaks needed) • Denoising blurs “first-token” vs “final-answer” latency
100
Daron Yondem @daron.me · 08/07/2025
Why you should care: • Voice agents & RAG get sub-100 ms replies • Fewer GPU-seconds → lower cloud bills • Bidirectional context → native infill & doc-wide edits • AR-centric tricks (KV cache, speculative decoding) need a rethink
100
Daron Yondem @daron.me · 08/07/2025
Big Tech’s onboard too: Google DeepMind’s Gemini Diffusion demo at I/O clocked 1479 tok/sec. 2025 is starting to look like diffusion’s “Transformers 2017” moment. 🌌
100
Daron Yondem @daron.me · 08/07/2025
Academic teasers like LLaDA (Feb ‘25) hinted at this. Mercury is the first public chat model to deliver the goods. Speed ≠ sloppy. On MMLU-Pro, Mercury matches GPT-4.1 Nano & Claude 3.5 Haiku—while running >7× faster. The “diffusion is quick but messy” era is over. ✅
100
Daron Yondem @daron.me · 08/07/2025
How? • Autoregressive (AR) LLMs: write left→right, token-by-token. • Diffusion LLMs: draft the whole sentence as noisy “static”, then iteratively denoise all tokens together. Parallel passes → giant throughput gains. 🚀
100
Daron Yondem @daron.me · 08/07/2025
Its code-specialized sibling goes even harder: >1 000 tok/sec. Numbers we used to associate with batch inference or exotic ASICs now happen live, in chat. 🔥
100
Daron Yondem @daron.me · 08/07/2025
Text-diffusion LLMs just graduated from lab demo to production reality. Let me show you why that matters… 👇 Inception Labs’ brand-new Mercury chat model streams ≈708 tokens /sec on ONE H100, \~7× faster than GPT-4-class “speed” variants on identical hardware. ⚡️
110
Daron Yondem @daron.me · 07/07/2025
Where could triangle attention shine first, code, math, or multimodal? 👇 Full paper link here arxiv.org/abs/2507.027... #DeepLearning #ScalingLaws #AIResearch #Transformers
arxiv.org
Fast and Simplex: 2-Simplicial Attention in Triton
Recent work has shown that training loss scales as a power law with both model size and the number of tokens, and that achieving compute-optimal models requires scaling model size and token count…
010
Daron Yondem @daron.me · 07/07/2025
⚠️ Caveat: The Triton kernel is still a prototype, and 2-simplicial layers add O(n × w₁ × w₂) FLOPs. Production-ready kernels and hardware co-design are the next hurdles.
100
Daron Yondem @daron.me · 07/07/2025
It also hints at richer relational reasoning, echoing AlphaFold’s triangle self-attention. Better structure understanding, Better reasoning, New horizons.
100
Daron Yondem @daron.me · 07/07/2025
Why it matters ➡️ Triangle attention rewrites Chinchilla-style scaling: you can grow parameters faster than tokens without hitting diminishing returns, crucial as we exhaust internet-scale data.
110
Daron Yondem @daron.me · 07/07/2025
⚡️ Compute check: Sliding-window (512 × 32) trilinear kernels keep complexity near standard attention at 48 k context ~55 ms per layer, peaking at 520 TFLOPS on one GPU.
100
Daron Yondem @daron.me · 07/07/2025
📈 Scaling laws bend: the params-to-loss exponent α jumps ≈ 20 % (0.142 → 0.168 on GSM8K, 0.090 → 0.108 on MMLU-pro). Each extra parameter now buys more accuracy when tokens are capped.
100
Daron Yondem @daron.me · 07/07/2025
📊 Same data, better model A 3.5 B-param model with 2-simplicial layers cut NLL by −2.3 % on GSM8K, −2.1 % on MMLU-pro, and −0.4 % on the coding benchmark MBPP vs. a vanilla Transformer.
100
Daron Yondem @daron.me · 07/07/2025
Modern LLMs are hitting a new wall: high-quality tokens are now scarcer than GPUs. The paper “Fast & Simplex” suggests a fix, swap dot-product attention (pairs) for 2-simplicial attention (triples).
100
Daron Yondem @daron.me · 07/07/2025
🤯 What if your Transformer paid attention to triangles instead of lines? Those extra corners buy you token efficiency and a steeper scaling curve.
220
Daron Yondem @daron.me · 14/06/2025
And if you do try it, let me know what you think. I’m always up for nerding out over health data. 👍
010
Daron Yondem @daron.me · 14/06/2025
This isn’t a sponsored post 🙂 (I know it might sound like one) just sharing a tool that quietly replaced a folder full of apps for me. If you’re juggling multiple trackers and want a single source of truth, it’s worth a look.
100
Daron Yondem @daron.me · 14/06/2025
The best is, It surfaces little insights I’d never spot on my own. Yesterday it linked a dip in glucose to a tougher-than-usual interval ride and suggested a quick carb refuel, simple, but it saved the rest of my day. Little moments like that make coaching clients (and myself!) so much easier.
100
Daron Yondem @daron.me · 14/06/2025
Seeing my heart-rate recovery, sleep trends, and real-time glucose on the same screen feels like someone finally gave my health data a common language.
100
Daron Yondem @daron.me · 14/06/2025
For years I’ve bounced between half-a-dozen trackers, but Bevel pulls everything, workouts from Strava, Apple Watch data, and even my Dexcom G7 glucose readings, into one clean, easy-to-read dashboard.
100
Daron Yondem @daron.me · 14/06/2025
Just wrapped up my first week with the Bevel Health app, and I’m honestly blown away. 📱✨
200
Daron Yondem @daron.me · 23/05/2025
I’ll have open slots on 27 May. If you’ll be on-site, DM me or drop a comment below—let’s grab a coffee and chat about all things Azure, AI, or whatever’s on your mind. See you in Düsseldorf! #ECS2025 #CloudSummitEU #CloudComputing #Azure #AI #Networking #TechCommunity
010
Daron Yondem @daron.me · 23/05/2025
🚀 Counting down to European AI and Cloud Summit next week! I’m excited to head to Düsseldorf for the European AI & Cloud Summit (26-28 May) where I’ll be taking the stage to share some fresh insights on Multi-Agent AI Workflows.
220