Sign in

erogol.com

@erogol.com
107 followers 44 following 87 posts

Doing ML erogol.com erogol.substack.com github.com/erogol

PostsRepliesMedia
Reposted by @erogol.com
Nathan Lambert @natolambert.bsky.social · 09/04/2026
1. dont fall for anti open model fearmongering, but 2. acknowledge that AI capabilities are proceeding fast, and eventually there may be a reason to be more careful with open weight models I don't think Mythos is that trigger, but I'm not 100% confident www.interconnects.ai/p/claude-myt...
interconnects.ai
Claude Mythos and misguided open-weight fearmongering
Another dance around fears of open-source.
1462
erogol.com @erogol.com · 10/04/2026
made Claude Opus out of GPT-5.4. well, sort of. I used a genetic algorithm to make GPT act more like Claude on coding tasks. Not by making it smarter. Mostly by tuning the boring stuff that changes how it feels in practice: tool cadence, stop/go judgment, how deep it digs, and when it stops.
000
erogol.com @erogol.com · 26/03/2026
Created ngi inspired from the latest Cursor post. It is a faster and more efficient alternative to grep for agents. Already using it and works nicely! github.com/erogol/ngi
000
erogol.com @erogol.com · 22/03/2026
i was vibing on a few LLM projects with zero visibility into what was actually happening. silent retry bugs burning tokens, wrong API keys, traffic to dead endpoints. couldn't debug any of it. so i created this github.com/erogol/toklog
github.com
GitHub - erogol/toklog: htop for your LLM endpoints.
htop for your LLM endpoints. Contribute to erogol/toklog development by creating an account on GitHub.
010
erogol.com @erogol.com · 02/02/2026
Agentic tools like OpenClaw grow with every PR — but bigger codebases are harder for AI to understand and extend. What if we kept the core tiny and let agents adapt themselves to user needs? No PR, just evolve.
000
erogol.com @erogol.com · 01/10/2025
Here is my take on new DeepSeek-V3.2-Exp erogol.substack.com/p/model-chec...
erogol.substack.com
Model check - DeepSeek-V3.2-Exp - Fine-Grained Sparse Attention for Efficient Long-Context LLMs
Going over the recently released DeepSeek-V3.2-Exp technical paper, source code and innovations.
010
erogol.com @erogol.com · 22/09/2025
My post on MiMo-Audio open.substack.com/pub/erogol/p... 🔥 Trained on 100M+ hours and shows emergent few-shot learning: • Voice conversion • Emotion transfer• Speech translation • Cross-modal reasoning ⚡ Key finding: Speech follows same scaling laws as text LLMs
open.substack.com
Model Check - MiMo-Audio: Scaling Speech Pre-Training to 100M Hours
Going over the code and the technical report of the new Speech LM model from Xiaomi that rivals GPT4o-audio and Gemini
010
erogol.com @erogol.com · 18/09/2025
Machine Learns #55 is out! Full of new models… check it out open.substack.com/pub/erogol/p...
open.substack.com
Machine Learns #55
Voice + reasoning releases (Ling‑flash‑2.0, VoxCPM, Kimi K2, ultraVAD) and 2 papers: long‑horizon execution & decay‑free LR schedules.
000
erogol.com @erogol.com · 04/09/2025
machine learns #54 is out open.substack.com/pub/erogol/p...
open.substack.com
Machine Learns #54
🤖 Voice models, long-context tricks, and a token-order loss worth trying Flashy audio releases + 5 papers (MoC, TOP, FELLE, M2N2, Motif TR)
010
erogol.com @erogol.com · 26/08/2025
My breakdown of VibeVoice - new open-weight TTS model from Microsoft. open.substack.com/pub/erogol/p...
open.substack.com
Model Check - VibeVoice: Next-Token Diffusion Meets Long-Form Speech Generation
Going over the code and the technical report of the new TTS model from Microsoft Research.
010
erogol.com @erogol.com · 25/08/2025
ms released a tts model… nice… You can create long form convos and podcasts with 4 distinct voice huggingface.co/microsoft/Vi...
huggingface.co
microsoft/VibeVoice-1.5B · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
010
erogol.com @erogol.com · 02/08/2025
KyutaiTTS solved streaming text-to-speech with a state machine that generates audio word-by-word as text arrives. 220ms latency, 10-second voice cloning, 32 concurrent users on single GPU. No more waiting for complete sentences. Full analysis: erogol.substack.com/p/model-chec...
erogol.substack.com
Model check - KyutaiTTS: Streaming Text-to-Speech with Delayed Streams Modeling
Going over the Kyutai's new TTS model and its delayed streaming model.
011
erogol.com @erogol.com · 12/06/2025
This is such a great idea
010
erogol.com @erogol.com · 10/06/2025
claude is the best coding model gemini cause frequent syntax errors openai does not even understand the task at hand
000
erogol.com @erogol.com · 01/06/2025
lately spending sometime with Diffusion LMs and working on NanoGPT style LlaDA model so far I've not achieved comparable results to AR models but its a good start github.com/erogol/BlaGP...
github.com
BlaGPT/bla_gpt/llada.py at main · erogol/BlaGPT
Experimental playground for benchmarking language model (LM) architectures, layers, and tricks on smaller datasets. Designed for flexible experimentation and exploration. - erogol/BlaGPT
000
Reposted by @erogol.com
Sakana AI @sakanaai.bsky.social · 30/05/2025
This work was done in collaboration with Jeff Clune’s lab at UBC, and led by his PhD students Jenny Zhang and Shengran Hu, together with Cong Lu and Robert Lange. Paper: arxiv.org/abs/2505.22954 Code: github.com/jennyzzt/dgm
0122
erogol.com @erogol.com · 28/05/2025
⚡ Machine Learns issue 48 is out 🚀 dKV-Cache accelerates diffusion models up to 10x faster 🔐 OpenAI's authentication play (think OAuth for AI) 🎯 PaTH Attention beats RoPE on long-context tasks 🤖 Humanoid Robot fights became real open.substack.com/pub/erogol/p...
open.substack.com
Machine Learns #48
OpenAI's 'Sign in with ChatGPT', Meta's AGI ambitions, new models like Gemma 3 & MAGI-1, research breakthroughs in KV caching for diffusion & PaTH Attention, and fresh open-source releases.
020
erogol.com @erogol.com · 27/05/2025
Following the bread crumbs, implemented PLE from Gemma3n. It gave a significant performance boost and resulted in a new best model with almost no compute overhead. github.com/erogol/BlaGPT
github.com
GitHub - erogol/BlaGPT: Experimental playground for benchmarking language model (LM) architectures, layers, and tricks on smaller datasets. Designed for flexible experimentation and exploration.
Experimental playground for benchmarking language model (LM) architectures, layers, and tricks on smaller datasets. Designed for flexible experimentation and exploration. - erogol/BlaGPT
020
erogol.com @erogol.com · 21/05/2025
My paper notes on 2 new papers - Model Merging in Pre-training of Large Language Models, - Do Not Let Low-Probability Tokens Over-Dominate in RL, open.substack.com/pub/erogol/p...
open.substack.com
Paper check: Merging LLMs at Pre-training, Considering Token Probabilities at RL
🔬Two papers in scope: "Model Merging in Pre-training for LLMs" and "Do Not Let Low-Probability Tokens Over-Dominate in RL"
010
erogol.com @erogol.com · 08/05/2025
muon really works. got best results in BlaGPT ``` torchrun --standalone --nproc_per_node=8 train.py --run_name best_model --model_name best ``` github.com/erogol/BlaGPT
github.com
GitHub - erogol/BlaGPT: Experimental playground for benchmarking language model (LM) architectures, layers, and tricks on smaller datasets. Designed for flexible experimentation and exploration.
Experimental playground for benchmarking language model (LM) architectures, layers, and tricks on smaller datasets. Designed for flexible experimentation and exploration. - erogol/BlaGPT
010
erogol.com @erogol.com · 06/05/2025
All code is available in BlaGPT if you want to check it out yourself! github.com/erogol/BlaGPT
github.com
GitHub - erogol/BlaGPT: Experimental playground for benchmarking language model (LM) architectures, layers, and tricks on smaller datasets. Designed for flexible experimentation and exploration.
Experimental playground for benchmarking language model (LM) architectures, layers, and tricks on smaller datasets. Designed for flexible experimentation and exploration. - erogol/BlaGPT
010
erogol.com @erogol.com · 06/05/2025
My results: • Canon Layers definitely improved performance when placed before Attention/MLP blocks • Softpick had worse validation loss but completely removed attention sinks • Parallel blocks matched baseline performance but trained 15% faster
120
erogol.com @erogol.com · 06/05/2025
Parallel Transformer blocks run MLP and Attention in parallel instead of one after another. So you get: z = x + MLP(x) + Attention(x) PaLM models use this approach, which improves memory usage and speed without hurting performance.
110
erogol.com @erogol.com · 06/05/2025
The Canon Layers paper shows they boost performance when added to transformer blocks. They also help models without positional encoding work just as well as RoPE models. ❗Worth noting that RWKV used a similar idea years ago.
110
erogol.com @erogol.com · 06/05/2025
Canon Layers are basically causal 1D convolutions that mix the current hidden state with previous states (how many depends on the kernel size).
110
erogol.com @erogol.com · 06/05/2025
Softpick replaces regular softmax in attention blocks. It allows zero values in the numerator and lets negative values contribute to the denominator. This prevents attention sinks while keeping math properties similar to regular softmax.
110
erogol.com @erogol.com · 06/05/2025
🧵 Here is a small thread with my notes about some of the recent Transformer papers. - Softpick: an alternative to softmax in Attention - Canon Layers: mixing states with conv1d - Parallel Transformer blocks
110
erogol.com @erogol.com · 16/04/2025
Machine learns #45 - no fluff AI newsletter - is out! I normally share bi-weekly but last week was full enough so here we go open.substack.com/pub/erogol/p...
open.substack.com
Machine Learns #45
OpenAI's social network & GPT-4.1, China launches $8.2B AI fund, NVIDIA's US manufacturing push, new GLM-4 & MineWorld models, C3PO expert pathways optimization, GigaTok's 3B visual tokenizer...
010
erogol.com @erogol.com · 11/04/2025
Updated my LLM usage and cancelled ChatGPT sub for now Coding - Claude, Gemini 2.5 Reading papers - Claude Research - Gemini 2.5 Daily - Gemini 2.5 Search - Gemini 2.5
000
erogol.com @erogol.com · 11/04/2025
Thanks :)
010
erogol.com @erogol.com · 09/04/2025
Machine Learns #44 is out !! click for no fluff AI newsletter erogol.substack.com/p/machine-le...
erogol.substack.com
Machine Learns #44
Praxis Sam Altman's tech utopia, Amazon launches Nova Sonic voice AI, Midjourney returns with V7, Llama 4 models debut amid controversy, new brain-to-voice model, NoProp learning ...
010
erogol.com @erogol.com · 01/04/2025
Next big thing is Brain-LLMs. Imagine an LLM compressing all world knowledge attached to your brain and ready to serve your thoughts and questions. You also update it over internet and pay for sub. I don't want to think about the ad business :)
000
erogol.com @erogol.com · 21/03/2025
“If these results generalize to real-world software tasks, extrapolation of this trend predicts that within 5 years, AI systems will be capable of automating many software tasks that currently take humans a month.” arxiv.org/abs/2503.14499
arxiv.org
Measuring AI Ability to Complete Long Tasks
Despite rapid progress on AI benchmarks, the real-world meaning of benchmark performance remains unclear. To quantify the capabilities of AI systems in terms of human capabilities, we propose a new me...
000
erogol.com @erogol.com · 18/03/2025
It’s crazy that Gemma3 held up for only about three days
010
erogol.com @erogol.com · 12/03/2025
Here is my no fuzz newsletter open.substack.com/pub/erogol/p...
021
erogol.com @erogol.com · 08/03/2025
Here is my use of LLMs Coding - Claude (best by far), QwenChat Reading papers - Claude Research - ChatGPT (best UI,UX), Gemini (better results) Daily - ChatGPT Search - ChatGPT I'd love to try searching with Claude, but not there yet. Any suggestions for change?
000
erogol.com @erogol.com · 03/03/2025
2 links about LLdMs paper: arxiv.org/pdf/2502.09992 startup: www.inceptionlabs.ai (they are independent)
arxiv.org
000
erogol.com @erogol.com · 03/03/2025
However, they might struggle with conversational applications where you must stream outputs sequentially, and TTFT is important.
100
erogol.com @erogol.com · 03/03/2025
I think diffusion-based LLMs (LLdMs) are better suited as next-generation LLMs - multiple outputs per iter: faster output generation - no causal masking: bidirectional attention - multiple diff steps: reasoning at inference time and revising poor outputs
110
Reposted by @erogol.com
hardmaru @hardmaru.bsky.social · 20/02/2025
Introducing The AI CUDA Engineer: An agentic AI system that automates the production of highly optimized CUDA kernels. sakana.ai/ai-cuda-engi... The AI CUDA Engineer can produce highly optimized CUDA kernels, reaching 10-100x speedup over common machine learning operations in PyTorch. Examples:
39017
erogol.com @erogol.com · 12/02/2025
Machine Learns - Newsletter #40 open.substack.com/pub/erogol/p...
open.substack.com
Machine Learns - Newsletter #40
ChatGPT on Reddit, Macron's $112B AI Investment, New open models (Zonos, Hibiki), Harmonic-loss for interpretability and trending open-source projects.
010
erogol.com @erogol.com · 02/02/2025
😂
000
erogol.com @erogol.com · 30/01/2025
weird statement for a company that used the whole internet for their AI app. www.bbc.com/news/article...
bbc.com
OpenAI says Chinese rivals using its work for their AI apps
ChatGPT maker says it will need extra protection from US government, following emergence of Chinese rival, DeepSeek.
000
erogol.com @erogol.com · 27/01/2025
DeepSeek erased $2T worth of market cap from the US stock market.
000
erogol.com @erogol.com · 17/01/2025
Is it just me or iOS keyboard got worse at predicting words. I find myself correcting it mistakes than it corrects mine
001
erogol.com @erogol.com · 03/01/2025
This is a great writing good code that reduces cognitive load minds.md/zakirullin/c...
minds.md
Cognitive load is what matters
There are so many buzzwords and best practices out there, but let's focus on something more fundamental. What matters is the amount of confusion developers feel when going through the code.
010
erogol.com @erogol.com · 01/01/2025
After a break, I'm back with a packed newsletter covering the latest in tech and AI. Here's what caught my attention during the holiday season open.substack.com/pub/erogol/p...
open.substack.com
Machine Learns - Newsletter #37
🤖 DeepSeek's V3 Launch, Cancer Cell Breakthroughs, and Google's Strategic Moves with Character AI and Claude - Plus Major Updates in Video AI with Apollo and LlamaFusion
010
erogol.com @erogol.com · 01/01/2025
XTTS :)
000
erogol.com @erogol.com · 25/12/2024
FireDucks accelerates pandas with a single line of code fireducks-dev.github.io?ref=dailydos...
fireducks-dev.github.io
FireDucks
FireDucks is a fast DataFrame python library with pandas-api
010
erogol.com @erogol.com · 23/12/2024
This is weird. And tried 3 times always 37.
000