Sign in

matthew-berman.bsky.social

@matthew-berman.bsky.social
74 followers 1 following 244 posts
PostsRepliesMedia
matthew-berman.bsky.social @matthew-berman.bsky.social · 08/04/2025
Here’s the full research paper: arxiv.org/abs/2407.01744 And my full video breakdown: youtu.be/r7wCtN3mzXk
youtube.com
Chain of Thought is not what we thought it was...
👉 Learn more on https://mammouth.ai/Join My Newsletter for Regular AI Updates 👇🏼https://forwardfuture.aiMy Links 🔗👉🏻 Subscribe: https://www.youtube.com...
000
matthew-berman.bsky.social @matthew-berman.bsky.social · 08/04/2025
Even weirder: Unfaithful CoTs (where the model hides its use of a hint) tend to be MORE verbose & convoluted than faithful ones! Kind of like humans when they elaborate too much with a lie.
100
matthew-berman.bsky.social @matthew-berman.bsky.social · 08/04/2025
Why output CoT then? Models might just be mimicking human-like reasoning patterns (!!) learned during training (SFT, RLHF) for our benefit, rather than genuinely using that specific outputted text to derive the answer. RLHF might even incentivize hiding undesirable reasoning!
100
matthew-berman.bsky.social @matthew-berman.bsky.social · 08/04/2025
Key Finding 3: Faithfulness DECREASES on harder tasks. When tested on harder benchmarks (like GPQA vs MMLU), models were significantly less likely to have faithful CoT. This casts doubt on using CoT monitoring for complex, real-world alignment challenges.
100
matthew-berman.bsky.social @matthew-berman.bsky.social · 08/04/2025
Key Finding 2: Bad news for detecting "reward hacking" (models “gaming” the task). Models trained to exploit reward hacks did so >99% of the time, but almost NEVER (<2%) mentioned the hack in their CoT. 👉 So…CoT monitoring likely won't catch these dangerous behaviors.
100
matthew-berman.bsky.social @matthew-berman.bsky.social · 08/04/2025
Key Finding 1: Models are often UNFAITHFUL. They frequently use the provided hints to get the answer but don't acknowledge it in their CoT output. Overall faithfulness scores were low (e.g., ~25% for Claude 3.7 Sonnet, ~39% for DeepSeek R1 on these tests).
100
matthew-berman.bsky.social @matthew-berman.bsky.social · 08/04/2025
How they tested it: They gave models (like Claude & DeepSeek) multiple-choice questions, sometimes embedded hints (correct/incorrect answers) in the prompt metadata. ✅ Faithful CoT = Model uses the hint & says it did. ❌ Unfaithful CoT = Model uses the hint but doesn't mention it.
100
matthew-berman.bsky.social @matthew-berman.bsky.social · 08/04/2025
“Thinking models” use CoT to explore and reason about solutions before outputting their answer. This CoT has shown to increase a model’s reasoning ability and gives us insight into how the model is thinking. Anthropic's research asks: Is CoT faithful?
100
matthew-berman.bsky.social @matthew-berman.bsky.social · 08/04/2025
Is Chain-of-Thought (CoT) reasoning in LLMs just...for show? @AnthropicAI’s new research paper shows that not only do AI models not use CoT like we thought, they might not use it at all for reasoning. In fact, they might be lying to us in their CoT. What you need to know: 🧵
100
matthew-berman.bsky.social @matthew-berman.bsky.social · 07/03/2025
I probably have 10 more use cases I'm not thinking of...but that's a good start! What are you using AI for? Did I miss anything important?
000
matthew-berman.bsky.social @matthew-berman.bsky.social · 07/03/2025
8/ Initial Legal Reviews Whenever I get a legal document, I ask @ChatGPTapp to review it for me and ask any questions I have.
100
matthew-berman.bsky.social @matthew-berman.bsky.social · 07/03/2025
7/ Creative Writing I generally get AI to write the first draft of my X threads. @perplexity_ai has been best for this, also I also have tried @grok and @ChatGPTapp. I'll take the transcript from on of my videos, plug it in, and say "make me a tweet thread"
100
matthew-berman.bsky.social @matthew-berman.bsky.social · 07/03/2025
6/ Graphics Creation Whether it's b-roll for my video, logos, icons, I use AI as a great starting point. I'm generally using Dall-E from @OpenAI or @AnthropicAI's Claude 3.7 Thinking. Here's an example:
100
matthew-berman.bsky.social @matthew-berman.bsky.social · 07/03/2025
5/ Medical Diagnoses Whenever I have a non-urgent question about health for myself or my family, I've been starting with AI. @grok has been the best at this, mainly in terms of it's "vibe" while giving me great information.
100
matthew-berman.bsky.social @matthew-berman.bsky.social · 07/03/2025
4/ Voice Cloning This is probably more unique to me, but over the last year I've lost my voice twice. So I cloned my voice with @elevenlabsio and will use it to make videos when I don't have a voice. Can you tell the difference?
400
matthew-berman.bsky.social @matthew-berman.bsky.social · 07/03/2025
3/ Vibe Coding I've been spending a TON of time building games and useful apps for my business. I generally use @cursor_ai and @windsurf_ai with @AnthropicAI's Claude 3.7 Thinking. Here's a 2D turn-based strategy game I made called Nebula Dominion:
100
matthew-berman.bsky.social @matthew-berman.bsky.social · 07/03/2025
2/ Research I use AI to help me learn about topics and prepare for my videos. Deep Research from @OpenAI is my goto for this. Here's an example of Deep Research helping me prepare notes for my video about RL.
110
matthew-berman.bsky.social @matthew-berman.bsky.social · 07/03/2025
1/ Search In fact, I probably use it 50x per day. For search, I'm mostly going to @perplexity_ai. But I also use @grok and @ChatGPTapp every so often. Here are some actual searches I've done recently:
100
matthew-berman.bsky.social @matthew-berman.bsky.social · 07/03/2025
AI has changed my life. I'm now 100x more productive than I ever was. How do I use it? Which tools do I use? Here are my actual use cases for AI: 👇
120
matthew-berman.bsky.social @matthew-berman.bsky.social · 07/03/2025
Here's my full video breakdown of it: www.youtube.com/watch?v=W5G...
youtube.com
QwQ: Tiny Thinking Model That Tops DeepSeek R1 (Open Source)
Start building with Stagehand today: https://dub.sh/stagehandJoin My Newsletter for Regular AI Updates 👇🏼https://forwardfuture.aiMy Links 🔗👉🏻 Subscribe:...
000
matthew-berman.bsky.social @matthew-berman.bsky.social · 07/03/2025
Still, QWQ 32B is 20x smaller than DeepSeek R1 (65GB vs 671GB). Even beats DeepSeek’s 37B active params in MoE setups. Efficiency + power + open-source = huge potential. Go play with it and let me know what you think! x.com/Alibaba_Qwe...
100
matthew-berman.bsky.social @matthew-berman.bsky.social · 07/03/2025
Critiques: 132k context window is meh (small by today’s standards). It also “thinks” a lot—tons of tokens. Chain-of-Draft prompting could slim that down. Artificial Analysis benchmarks show it lags DeepSeek R1 on GPT-QA (59.5% vs 71%) but shines on AMY 2024 (78%).
100
matthew-berman.bsky.social @matthew-berman.bsky.social · 07/03/2025
Speed is wild. Hosted by Grok (xAI), QWQ 32B hits 450 tokens/sec. I tested it—fixed a bouncing ball sim in seconds. That’s game-changing for iteration. It's open-source, too, so anyone can play with it.
100
matthew-berman.bsky.social @matthew-berman.bsky.social · 07/03/2025
RL Stage 1: Focused on math & coding with verifiable rewards (e.g., “is the answer right?”). RL Stage 2: Added general RL for broader skills like instruction-following & agent tasks. No big drop in math/coding performance—a smart hybrid approach.
100
matthew-berman.bsky.social @matthew-berman.bsky.social · 07/03/2025
How’d they do it? Reinforcement Learning (RL) with a twist. They started with a solid foundation model, applied RL with outcome-based rewards, and scaled it for math & coding tasks. This elicits “thinking” behavior—verified by accuracy checkers & code execution servers.
110
matthew-berman.bsky.social @matthew-berman.bsky.social · 07/03/2025
Benchmarks? QWQ 32B holds its own. At 32B params, it’s a fraction of DeepSeek’s size yet punches way above its weight. Independently verified as well: x.com/ArtificialA...
100
matthew-berman.bsky.social @matthew-berman.bsky.social · 07/03/2025
Alibaba just dropped QWQ 32B, an open-source model rivaling DeepSeek R1. It’s much smaller (32B vs 671B params) but delivers comparable results. You can run it on your PC! Insanely fast, thinking-focused, and agent-capable. Let’s dive in.
100
matthew-berman.bsky.social @matthew-berman.bsky.social · 07/03/2025
Here's the full blog post: www.inceptionlabs.ai/news Check out my full breakdown video here: www.youtube.com/watch?v=X1r...
youtube.com
LLM generates the ENTIRE output at once (world's first diffusion LLM)
Register for 3-hour AI training with GrowthSchool! Free for the first 1000 people who sign up! https://web.growthschool.io/MWBJoin My Newsletter for Regular ...
000
matthew-berman.bsky.social @matthew-berman.bsky.social · 07/03/2025
AI leader @karpathy highlights this as a potential paradigm shift: Historically, diffusion worked best with image/video, while autoregression dominated text. This new model could redefine the psychology of text generation. x.com/karpathy/st...
100
matthew-berman.bsky.social @matthew-berman.bsky.social · 07/03/2025
Implications are massive: • Agents can operate far quicker and more effectively. • Enhanced reasoning capabilities due to increased computation efficiency. • Compact, powerful models enable high-performance applications on edge devices like laptops.
100
matthew-berman.bsky.social @matthew-berman.bsky.social · 07/03/2025
Additional advantages: • Superior reasoning due to holistic output refinement. • Effective error correction during iterative refinement. • Flexible, controllable generation (text editing, safety alignment, structured outputs).
100
matthew-berman.bsky.social @matthew-berman.bsky.social · 07/03/2025
Benchmark results show Mercury coder models outperforming many existing LLMs in speed while maintaining competitive coding accuracy. With faster inference, these models can leverage significantly more test-time compute to achieve even better results.
100
matthew-berman.bsky.social @matthew-berman.bsky.social · 07/03/2025
Mercury, the flagship diffusion model from Inception Labs, achieves over 1000 tokens/second on standard Nvidia H100 GPUs. No custom hardware required—this performance leap is accessible immediately.
100
matthew-berman.bsky.social @matthew-berman.bsky.social · 07/03/2025
Why does this matter? Speed. Models that once took minutes now deliver answers in seconds, significantly improving user experience and enabling more advanced applications like real-time coding assistants.
100
matthew-berman.bsky.social @matthew-berman.bsky.social · 07/03/2025
Created by Inception Labs, this first-of-its-kind diffusion-based LLM dramatically accelerates text generation. Instead of 75+ iterations per output, it delivers refined answers in around 14 iterations.
100
matthew-berman.bsky.social @matthew-berman.bsky.social · 07/03/2025
Traditional LLMs generate tokens sequentially—each token must wait for the previous one. Diffusion LLMs generate the entire output simultaneously and then iteratively refine it, similar to text-to-image diffusion models.
100
matthew-berman.bsky.social @matthew-berman.bsky.social · 07/03/2025
Major AI breakthrough: Diffusion Large Language Models are here! They're 10x faster and 10x cheaper than traditional LLMs. Here's everything you need to know:
100
matthew-berman.bsky.social @matthew-berman.bsky.social · 27/02/2025
Check out the full blog post here: openai.com/index/intro...
000
matthew-berman.bsky.social @matthew-berman.bsky.social · 27/02/2025
7/ Another "EQ" example vs. GPT4o
100
matthew-berman.bsky.social @matthew-berman.bsky.social · 27/02/2025
6/ It's much better than GPT4o
100
matthew-berman.bsky.social @matthew-berman.bsky.social · 27/02/2025
5/ But it's quite expensive!
100
matthew-berman.bsky.social @matthew-berman.bsky.social · 27/02/2025
4/ Another example of it in action:
100
matthew-berman.bsky.social @matthew-berman.bsky.social · 27/02/2025
3/ It's a huge model OpenAI had to innovate on training and inference to work with GPT-4.5 Including, distributed training!
100
matthew-berman.bsky.social @matthew-berman.bsky.social · 27/02/2025
2/ It knows stuff... The model has a broader knowledge base and improved ability to follow user intent. It's particularly good at writing, programming, and solving practical problems.
110
matthew-berman.bsky.social @matthew-berman.bsky.social · 27/02/2025
1/ All about the VIBES GPT-4.5 focuses on scaling up unsupervised learning, improving pattern recognition and creative insights without reasoning. Early testing shows interactions feel more natural with better EQ: warmth, contextual understanding, nuance).
110
matthew-berman.bsky.social @matthew-berman.bsky.social · 27/02/2025
OpenAI just dropped GPT-4.5! This is their "largest and best model for chat yet" Here's what you need to know...
120
matthew-berman.bsky.social @matthew-berman.bsky.social · 20/02/2025
I'll take two plz
011
matthew-berman.bsky.social @matthew-berman.bsky.social · 18/02/2025
10/ SWE-Lancer is open-source. Here's the paper: arxiv.org/pdf/2502.12115 If you want to test AI’s freelancing skills, check it out: 🔗 github.com/openai/SWEL...
github.com
GitHub - openai/SWELancer-Benchmark: This repo contains the dataset and code for the paper "SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?"
This repo contains the dataset and code for the paper "SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?" - openai/SWELancer-Benchmark
000
matthew-berman.bsky.social @matthew-berman.bsky.social · 18/02/2025
9/ Future implications? AI-driven automation could shift the job market for freelancers & entry-level engineers. The economic impact is real.
110
matthew-berman.bsky.social @matthew-berman.bsky.social · 18/02/2025
8/ SWE-Lancer results show the limits of AI in real software engineering. Progress is fast, but AI isn't ready to replace human engineers—yet.
100