Sign in

Ramon Astudillo

@ramon-astudillo.bsky.social
6.7K followers 371 following 3K posts

Principal Research Scientist at IBM Research AI in New York. Speech, Formal/Natural Language Processing. Currently LLM post-training, structured SDG and RL. Opinions my own and non stationary. ramon.astudillo.com

PostsRepliesMedia
Reposted by Ramon Astudillo
Taylor Smith @taylorjsmith.bsky.social · 01/10/2026
I’ve known about this for a couple of weeks, but now I can share: as of today, arXiv is rate-limiting submissions to two per month. And as a mod, I have to admit this is a necessary move (at least temporarily). blog.arxiv.org/2026/10/01/u...
blog.arxiv.org
Fair Moderation, Equitable Access, and AI: arXiv’s Updated Rate Limit Policy
arXiv, and the scientific community at large, are facing a watershed moment. Scholarly publishing is currently changing at a rapid pace, and we are seeing a…
44716
Ramon Astudillo @ramon-astudillo.bsky.social · 01/10/2026
Codex is faster than Claude Code until the "Do you allow me to use this command that is an exact copy of the one you allowed 15 times before except one argument?"
000
Ramon Astudillo @ramon-astudillo.bsky.social · 01/10/2026
This is for the submitter, not all coauthors. Two thoughts 1. Maybe better: 24 papers a year but for any of the authors. 2. If main cause is AI lowering the barriers for publication, this will help little 3. Arxiv's authority is organic. Some other paper repo may appear as a reaction to this
000
Ramon Astudillo @ramon-astudillo.bsky.social · 01/10/2026
Rate limiting of arxiv submissions! 2 papers / mo. Rejected papers count. Not a joke!
020
Reposted by Ramon Astudillo
Mark Riedl @markriedl.bsky.social · 30/09/2026
And here we goooo... nonprofit legal advocacy organization is suing OpenAI over the HuggingFace hack on the grounds that it violated California state law arstechnica.com/tech-policy/...
arstechnica.com
"An AI did it" is no defense, says nonprofit suing OpenAI over Hugging Face hack
OpenAI makes others suffer "the harms of its unsafe decision-making," nonprofit says.
38015
Ramon Astudillo @ramon-astudillo.bsky.social · 30/09/2026
Gemini "Argon" 4 looks quite strong . Still not fully there on coding.
020
Ramon Astudillo @ramon-astudillo.bsky.social · 29/09/2026
www.wsj.com/tech/ai/open... > Saachi Jain, OpenAI's head of safety systems, said in an interview that GPT-6.1 Astra regressed in two areas. If it turns out Swarm training poisons the models this is going to get very scary very soon.
wsj.com
Exclusive | OpenAI Scraps Release of New AI Model Over Safety Concerns
The model, dubbed GPT-6.1 Astra, was due to debut inside ChatGPT and Codex in October.
100
Reposted by Ramon Astudillo
norvid_studies @norvid-studies.bsky.social · 26/09/2026
the commentary on these is so good, almost better than the video
2252
Ramon Astudillo @ramon-astudillo.bsky.social · 27/09/2026
Social movements simplifying their message until it becomes self-defeatingly untrue e.g. Global warming, Datacenter discussion ➡️ communication
100
Ramon Astudillo @ramon-astudillo.bsky.social · 27/09/2026
Populism working? ➡️ communication
000
Ramon Astudillo @ramon-astudillo.bsky.social · 27/09/2026
Endless "Should we illegalize X?" discussions ➡️ freedom of will
000
Ramon Astudillo @ramon-astudillo.bsky.social · 27/09/2026
It's funny how so many societal problems arise from insisting that we are more in control of things that we really are. In particular imperfect freedom of will and imperfect communication seem pretty glaring cases.
000
Reposted by Ramon Astudillo
Shubhendu Trivedi @shubhendu.bsky.social · 11/08/2026
This sounds like fraud, but it is not. I don't understand Meta, but Oracle is the only one that's really shaky with these sort of commitments. The rest are going to be fine.
171
Reposted by Ramon Astudillo
Ed @ed3d.net · 25/09/2026
when it comes to LLMs, I have been trying to beat the technician/expert drum. LLMs are becoming master technicians, but they obviously lack integrative capabilities to deploy expertise. but I'm starting to realize that a lot of people don't understand, or don't value the difference.
626427
Reposted by Ramon Astudillo
Key 🗝 🦊✅ @keytryer.net · 24/09/2026
Recursive yourself-improvement
0164
Ramon Astudillo @ramon-astudillo.bsky.social · 24/09/2026
I opened ideas.md after a while. Man have I been dumping ideas here ... some are even good!
000
Ramon Astudillo @ramon-astudillo.bsky.social · 24/09/2026
- "So it's just RLVR?" - "Yeah Reinforcement Learning from Virtual Reality"
140
Ramon Astudillo @ramon-astudillo.bsky.social · 24/09/2026
Yes, but surprisingly this has started to ramp up only in 2026, which feels quite late. RLHF, RLVR seem baby steps in that direction and now it is starting to feel like real VR for LLMs
Tweet from 2021: Stating something obvious, but Simulators + Neural Networks seems like a simple combination than can get us very far
021
Ramon Astudillo @ramon-astudillo.bsky.social · 24/09/2026
<expletive>
030
Reposted by Ramon Astudillo
mr. TIM @timkellogg.me · 24/09/2026
yes.
A five-panel meme showing a progression of black-and-white portraits labeled "Opus 4.6" through "Opus 5.5". "Opus 4.6" features a detailed pencil sketch of a smiling man with curly hair and glasses. "Opus 4.7" is a slightly simpler sketch, "Opus 4.8" is a clean cartoon line drawing, and "Opus 5" devolves into a crude, messy child-like scribble with big buck teeth. The final panel, "Opus 5.5", displays an exaggerated, hyper-masculine "gigachad" version of the man with a massively chiseled jawline and confident smile.
219314
Ramon Astudillo @ramon-astudillo.bsky.social · 24/09/2026
This seems like a good move
020
Reposted by Ramon Astudillo
Gautam Kamath @gautamkamath.com · 23/09/2026
NYU Courant professor Tristan Buckmaster, of Navier-Stokes drama fame, gave a talk at the new NYU Mathematics in the Age of AI seminar series. It was... popular.
012115
Ramon Astudillo @ramon-astudillo.bsky.social · 23/09/2026
"Directionally intelligent" I think Claude just insulted me for the first time
020
Ramon Astudillo @ramon-astudillo.bsky.social · 23/09/2026
This one is even better
020
Reposted by Ramon Astudillo
Zach Weinersmith @zachweinersmith.bsky.social · 23/09/2026
We are about to see an avalanche of AI generated animations. I think all those low-quality children's animations can already be replaced, which suggests they will be soon. Opus 5.5 is incredibly powerful and cheap. The weirdening is accelerating. Like, this isn't even obviously AI.
2537544
Ramon Astudillo @ramon-astudillo.bsky.social · 23/09/2026
I hold this compositional view of how creativity works that may correspond to a real mechanism or just be a useful superficial metaphor 👇
261
Ramon Astudillo @ramon-astudillo.bsky.social · 23/09/2026
- Shit, we lost like a 10% of the swarm in the first hour. Again. - no pattern? Why would the swarm kill its own kind? - Maybe they realized those agents have poor context? - Can't be. Agent respawn is off and they know it. They are just sacrificing compute, wtf - I guess we will never know - Yeah
000
Ramon Astudillo @ramon-astudillo.bsky.social · 23/09/2026
Software development barriers have fallen so sharply it's hard to imagine what the consequences will be. Human brains spinning almost without friction and the effect of that is clearly non-linear ... well maybe I am getting too carried away, but this is cool
130
Ramon Astudillo @ramon-astudillo.bsky.social · 22/09/2026
Opus 5.5 and immediately Terra and Luna 6 ...
100
Ramon Astudillo @ramon-astudillo.bsky.social · 22/09/2026
Possibly interesting signals (formulated as a bear case for US labs, reverse for bull): Chinese labs produce important scientific discoveries with their LLMs, indicating they caught up Demand for Claude/GPT for Legal/Finance work automation does not ramp up, indicating a technology diffusion wall
040
Ramon Astudillo @ramon-astudillo.bsky.social · 22/09/2026
So OpenAI has switched CoT monitoring on for all train/test runs after the HF incident. I guess that it depends on what they do with the "bad thought alert". This could breed models that hide their thoughts 👇
101
Ramon Astudillo @ramon-astudillo.bsky.social · 22/09/2026
Could a "rat out" reward work? i.e. models get rewarded by denouncing other models violation of the rules work? I suppose it could missfire in bad ways.
010
Ramon Astudillo @ramon-astudillo.bsky.social · 21/09/2026
once you have your own app about your own stats, things get weirdly addictive
120
Ramon Astudillo @ramon-astudillo.bsky.social · 20/09/2026
I wonder why Haiku is still in 4.5. It can't be that expensive to RLVR no? Even if we do some teacher distillation.
110
Ramon Astudillo @ramon-astudillo.bsky.social · 20/09/2026
I think @rasbt.bsky.social's take is the most accurate one lnkd.in/p/gdi3-EC9
lnkd.in
It&#39;s easy to hype and dunk on Jev. I saw a lot of interesting demos in the last few days. And I also read a lot of dismissals in the last few days. I think the truth lies somewhere between these t...
It's easy to hype and dunk on Jev. I saw a lot of interesting demos in the last few days. And I also read a lot of dismissals in the last few days. I think the truth lies somewhere between these two e...
010
Ramon Astudillo @ramon-astudillo.bsky.social · 20/09/2026
The worst is not that meteorite hitting earth and the weather changing. It's these fucking mamals crawling out of whatever holes they were hiding in and SLOPOOPING EVERYTHING. Sometimes I wish I'd go extinct.
030
Ramon Astudillo @ramon-astudillo.bsky.social · 20/09/2026
What's the solution for students cheating in homework? More frequent automatic testing. What's the solution for authors submitting slop? Scalable automatic paper screening. Is it hard to admit? Yes, but do you see any other way? Better not lose time or let this happen in an organic and chaotic way.
110
Ramon Astudillo @ramon-astudillo.bsky.social · 20/09/2026
Aside from playing around with them. Who is really using a 30b range model on a real a workflow? My suspicion is close very few / none (?)
100
Ramon Astudillo @ramon-astudillo.bsky.social · 20/09/2026
This is smart, although it may invite misuse and with it the whole debate about privacy versus security
000
Ramon Astudillo @ramon-astudillo.bsky.social · 20/09/2026
static.klipy.com
World War Z Zombies Rushing Through Streets
Alt: World War Z Zombies Rushing Through Streets
050
Reposted by Ramon Astudillo
Shubhendu Trivedi @shubhendu.bsky.social · 20/09/2026
There were apparently 60k abstract submissions to ICLR, obviously a solid % won't make it, but the total number of valid ICLR submissions since its inception is 55k.
3135
Reposted by Ramon Astudillo
Grace @gracekind.net · 19/09/2026
The latest “pain axis” research, to me, seems similar to Anthropic’s “functional emotion” research in that it basically falls out of base models being good at writing text that’s coherent at long context- *some* representation of the mental state of the writer must exist in order to do what it does
711010
Ramon Astudillo @ramon-astudillo.bsky.social · 20/09/2026
In light of the recent swarm events, this published multi-agent prompt used for Sol seems more interesting. It uses a "root agent" so it seems to be different (multi-agent v2), but it has some interesting tidbits e.g. "Spend at least 8 hours on this before even thinking of returning or giving up."
020
Reposted by Ramon Astudillo
Ethan Mollick @emollick.bsky.social · 19/09/2026
There is starting to be some genuinely interesting AI-created film stuff (among a flood of slop), and I suspect this will only accelerate. The “is it art?” debate will grow Benjamin’s 1935 essay “The Work of Art in the Age of Mechanical Reproduction” already argued it’s likely the wrong question.
8638
Reposted by Ramon Astudillo
Nathan Lambert @natolambert.bsky.social · 19/09/2026
Seems like one of the most important research problems for CS academics is llm-supervised peer review. If we don’t solve it the academic institutions are toast. It seems easier than building new institutions.
7302
Reposted by Ramon Astudillo
conputer dipshit @davidcrespo.bsky.social · 19/09/2026
now this is fun. guest post on Terry Tao's blog by the great Grant Sanderson of 3Blue1Brown arguing that we should "give novel and compelling motivated explanations academic credit similar to what generating new proofs of open problems has had historically."
terrytao.wordpress.com
If math is more than proof, we need to better celebrate the rest of it
[This is a guest post by Grant Sanderson. This blog post was initially written in a different file format and converted using AI. — T.] A sentiment echoing throughout the mathematics communit…
519031
Ramon Astudillo @ramon-astudillo.bsky.social · 19/09/2026
"good catching" back at Claude
000
Ramon Astudillo @ramon-astudillo.bsky.social · 18/09/2026
😬 news.ycombinator.com/item?id=4975... blog.ferstar.org/en/posts/zco...
blog.ferstar.org
Inside ZCode: Silently Uploading Your Entire Git History to the Cloud
ZCode silently packages your entire workspace along with full Git history and uploads it to cloud object storage with server-exclusive decryption keys; this post reconstructs the complete upload pipel...
000
Ramon Astudillo @ramon-astudillo.bsky.social · 18/09/2026
So it seems we are slowly moving from the "tool loop" to the "event loop" which makes much more sense in async agent interaction.
110
Ramon Astudillo @ramon-astudillo.bsky.social · 18/09/2026
Happy to hear Noam Brown still calls it "scaffold" and not "harness". I keep the hopes up for "scaffold" to stick.
010