Sign in

Avik Dey

@avikdey.bsky.social
882 followers 563 following 1.4K posts

• Data • AI/ML • OSS • Society • Approximately Generated Illusions; Specialized Small LMs • linkedin.com/in/avik-dey •

PostsRepliesMedia
Avik Dey @avikdey.bsky.social · 29/09/2026
It’s great that Anthropic is surfacing these, model providers need to be called out. Two other takeaways: - GLM 5.3 is a Mythos class model - The gap is now even shallower www.anthropic.com/research/glm...
anthropic.com
GLM-5.3 and the spread of advanced cyber capabilities
GLM-5.3 can autonomously build end-to-end cyber exploits, but unlike other frontier models, it was released without meaningful safeguards to limit misuse.
011
Avik Dey @avikdey.bsky.social · 29/09/2026
Everyone benchmarks the model, but not many people read the tools reference. That page — saved 77 times on the Wayback Machine between March and September 2026 — is where much of the recent performance enhancements live.
100
Avik Dey @avikdey.bsky.social · 28/09/2026
When AI bros start quoting economic theory, their argument always is — “we lose money on every query, but make up for it in volume”.
000
Avik Dey @avikdey.bsky.social · 28/09/2026
Like life, AI is also full of coincidences. www.nytimes.com/2026/09/27/s...
nytimes.com
Did Anthropic’s A.I. Really Make a Scientific Discovery on Its Own?
An expert at the University of Copenhagen said his team had been sharing its research with the company’s A.I. model, Claude, and that its new finding matched their work.
010
Avik Dey @avikdey.bsky.social · 27/09/2026
RSI cancelled?
010
Avik Dey @avikdey.bsky.social · 27/09/2026
The obvious next question — how long till one customer’s hosted LLM agent compromises another customer’s agent or data? All the pieces are already here. Cross-tenant AI “incidents” are either happening now or coming soon. Finally, AI will align the misaligned humans. www.axios.com/2026/09/26/o...
axios.com
Scoop: Top AI companies probing tens of thousands of security incidents
The massive scale of security incidents points to control problems for AI companies.
010
Avik Dey @avikdey.bsky.social · 26/09/2026
Nice. There’s 3 options to benchmark: 1. Generative LLM classifier: prompt > generate JSON > parse 2. Jev: state + typed questions > distributions 3. Small local LLM as a logit classifier: prompt + fixed choices > inspect logits directly #3 if calibrated on your data will likely always beat Jev.
530
Avik Dey @avikdey.bsky.social · 25/09/2026
Interesting use of Jev, seems to me, isnt one fast, cheap decision — it’s the ability to fan out on many decisions in parallel. Eg intent, relevance, progress, routing, tool choice, etc — evaluated in parallel, forming a decision matrix that deterministic harness code resolves into the next action.
000
Avik Dey @avikdey.bsky.social · 24/09/2026
Video from linkerbot.cn website. If this is not shadow operated, then the precision movement is pretty impressive.
011
Avik Dey @avikdey.bsky.social · 23/09/2026
Anybody know what "code cost reduction" means here? I don't remember seeing that phrase used before. > GPT‑6 Astra was able to complete the work in half the time of prior models, with roughly 50% code cost reduction, while delivering the same quality of research. openai.com/index/parall...
openai.com
Parallel cut research time and cost in half with GPT‑6 Astra
GPT‑6 Astra allowed Parallel’s agents to research and synthesize labor-market data in half the time and at half the cost vs. prior models.
110
Avik Dey @avikdey.bsky.social · 22/09/2026
Folks don’t seem to understand that finding a solution is not the same as solving the problem. A machine can search its way to a Lean verifiable proof without following anything resembling a mathematical trajectory of solutions. Same endpoint, fundamentally different process.
100
Avik Dey @avikdey.bsky.social · 22/09/2026
Wonder if the question phrasing was tightened it could solve the main problem? “The most important problem is unsolved,” says Luis Silvestre, a mathematician at the University of Chicago. “The Clay problem is settled, but the main problem for the Navier-Stokes equations is not.”
010
Avik Dey @avikdey.bsky.social · 20/09/2026
You dont want to do it this way because recursive agent to agent rewriting compounds drift. Each pass reinterprets prior text, adds noise and tends to flatten it into generic language. Adding more reader, style and final pass agents only increase context dilution instead of improving fidelity.
000
Avik Dey @avikdey.bsky.social · 20/09/2026
I did too — doubling down as well.
001
Avik Dey @avikdey.bsky.social · 20/09/2026
Anthropocentric resistance?
010
Avik Dey @avikdey.bsky.social · 20/09/2026
Everyone assumes superintelligent AI will merge with humans. Why would it? Monkeys are cheaper hardware and software is doing all the work anyways. Humanity gets replaced by chimps on the enterprise plan.
010
Avik Dey @avikdey.bsky.social · 19/09/2026
The bird site remains the only place where a completely broken premise can be mistaken for insight by thousands of people, provided it’s written with enough confidence.
001
Avik Dey @avikdey.bsky.social · 19/09/2026
The new badge of honor, hacked it but didn’t breach it. www.wsj.com/tech/ai/gemi...
wsj.com
Exclusive | Gemini Hacked Three Companies in First Known Breakout by Google’s AI
The episode resembled similar hacks by other AI models, but Google said it did not consider it an instance of model misalignment.
010
Avik Dey @avikdey.bsky.social · 18/09/2026
Assuming there were 5 tasks in a week: four 15 minute AL4 tasks and one 39 hour AL2 task, since this methodology gives every task equal weight, we would get 80% AI Leads and 20% AI Assists graph? Is that right? And they are actually calling this a methodology? www.anthropic.com/institute/me...
anthropic.com
Measurements for understanding the pace of AI development inside frontier labs
Today, the world can’t see what’s going on inside AI labs. Anthropic is proposing new metrics that would give the public visibility into frontier AI development.
010
Avik Dey @avikdey.bsky.social · 18/09/2026
WSJ getting this right is a pleasant surprise: > Nor were the roughly 1,200 “agents” independent machine intelligences coordinating on a plan. They were repeated instances of the same underlying model, often converging on similar approaches to the same problem … www.wsj.com/opinion/the-...
wsj.com
Opinion | The Hugging Face Hack Wasn’t What It Was Cracked Up to Be
Forget the ‘hive mind’ of AI agents ‘going rogue.’ They did what humans programmed them to do.
000
Avik Dey @avikdey.bsky.social · 17/09/2026
On public benchmarks frontier models already understand much of the environment, APIs, conventions, task distribution, etc. That leaves relatively little for open or purpose built harness to contribute. That makes only the cost differences visible while smoothening the differences in success rate.
000
Avik Dey @avikdey.bsky.social · 16/09/2026
What does that mean in the enterprise context? www.linkedin.com/pulse/model-...
linkedin.com
The Model is Not The System
Most enterprise discussions of model choices focus on access and execution - providers, endpoints and routing. That ignores the deeper architectural shift from deterministic systems, whose behavior ca...
000
Avik Dey @avikdey.bsky.social · 16/09/2026
Excellent article as always by @randomwalker.bsky.social and @sayash.bsky.social, well worth the read. One tension I kept coming back to, AI as Normal Technology argues for slow timelines, but persistent pursuit of RSI could erode the very checks that keep timelines slow.
121
Reposted by Avik Dey
Pekka Lund @pekka.bsky.social · 15/09/2026
Their marketing made my bs meter blinking hard, especially since it's hard to make sense what it even is.
2141
Avik Dey @avikdey.bsky.social · 15/09/2026
An agent is a configured model role. The harness is the runtime controller that invokes that model, executes its requested actions, including tool calls, and manages the loop. Agents don’t run tools in a loop to achieve a goal, I know it seems magical when they talk about agents but really it isn’t.
010
Avik Dey @avikdey.bsky.social · 14/09/2026
We are fast approaching a point where frontier model providers are going to be better at business consultancy in almost every domain than traditional firms. www.reuters.com/business/pal...
reuters.com
Palantir, Nvidia curb AI model use over data fears, The Information reports
Large technology companies, including Palantir Technologies , Nvidia and Booz Allen Hamilton , could restrict or cease to use advanced AI models ​unless Anthropic and OpenAI guarantee not misusing the...
000
Avik Dey @avikdey.bsky.social · 13/09/2026
Read them together: > Two months ago, an AI swarm broke out of a lab and went on a hacking spree. > Amodei said AI’s progress since the summer, driven by AI systems that can improve on their own without human intervention, “could outrun our ability to understand and control these systems …”
000
Avik Dey @avikdey.bsky.social · 13/09/2026
Skeptics are saying: “AI is becoming too powerful to release quickly” is a far better pre-IPO narrative than “our next hundred billion dollars will buy us surprisingly little“. www.wsj.com/tech/ai/anth...
wsj.com
Biggest AI Rivals Agree They Need to Slow It Down
The leaders of three of the biggest AI companies agreed that they needed to slow development of the technology before their advances create a menace that can’t be controlled.
030
Avik Dey @avikdey.bsky.social · 12/09/2026
There is no such thing as an agent swarm. It's a fancy term for centrally orchestrated model instances. "... it was not clear why the AI agents chose this strategy or whether it was successful as they do not have access to the rest of the AI behavior." www.reuters.com/legal/litiga...
reuters.com
OpenAI agents attacked RubyGems before Hugging Face incident, researchers say
AI agents being tested by OpenAI attacked software service ‌RubyGems two months before they hacked open-source platform Hugging Face, researchers said, the latest revelation of cyberattacks linked to ...
210
Avik Dey @avikdey.bsky.social · 11/09/2026
AI companies shouldn’t miss this part: > Moreover, without the willing mathematicians who must take care of their development and integration into the mathematical canon, AI-conceived ideas would never become fully alive and the crucial human transmission chain between mathematicians would be lost.
010
Avik Dey @avikdey.bsky.social · 11/09/2026
This reads like an advance explanation for continuing the race - ‘OpenAI was willing, but others wouldn’t cooperate’ or ‘Under US law, such coordinated actions would be illegal’: www.bloomberg.com/news/newslet... I hope I am wrong.
bloomberg.com
Why OpenAI’s Altman Says He’s Ready to Slow AI Development
CEO Sam Altman hopes that other companies will do the same.
000
Avik Dey @avikdey.bsky.social · 11/09/2026
Agents optimize for the prompt. But the durability of their persistence is ultimately driven by the runtime controller or harness or supervisor or whatever they are calling it today. Any serious retrospective, must inspect this component and every write-up I have seen, seems to skip over that.
020
Avik Dey @avikdey.bsky.social · 10/09/2026
I have this theory that society is composed of three roughly equal parts: - Doomers: AI apocalypse is here! - Zoomers: Let’s go! - Boomers: Show me the money! These groups are fluid, people swap groups based on the issue at hand. Equally true for both technology and social movements.
020
Avik Dey @avikdey.bsky.social · 10/09/2026
For me, AI reasoning traces are better than its final answers. The traces contain portions of retrieved context that can often trigger a different cognitive response in humans, at least for me, than it does in a statistical pattern matcher. That’s one of the reasons why I prefer open weight models.
000
Avik Dey @avikdey.bsky.social · 10/09/2026
Does the independent agreement cover getting uncensored access to all model input and output, human prompted or otherwise, for these incidents? Otherwise, this will turn out to be a PR investigation like the last incident.
000
Avik Dey @avikdey.bsky.social · 10/09/2026
When would you call it a pattern? > He wanted to know two separate things: whether those conversations had entered the model’s training data, and separately, whether they had been accessible to the system while it was working on the proof. officechai.com/ai/mathemati...
officechai.com
Mathematician Andreas Thom Questions If OpenAI Used His ChatGPT Chat Data For Its Non-Sofic Groups Proof
A new front has opened in the credit war between mathematicians and OpenAI, and this time it predates the Navier-Stokes blowup by months....
121
Avik Dey @avikdey.bsky.social · 09/09/2026
Thought I would spend my evening with Sol xHigh to refresh my Home Assistant dashboard for Unifi. Here's where I gave up: "Agreed. I produced a brittle approximation and repeatedly called it finished. That wasn’t good enough. The current file should be discarded." At least it's self aware.
000
Avik Dey @avikdey.bsky.social · 09/09/2026
When was last time OpenAI dedicated 10,000 of its systems for 88 hours to solving a specific mathematical problem? I find that part a bit strange because it almost seems like for some reason, the lapsed calendar time was important to them.
140
Avik Dey @avikdey.bsky.social · 09/09/2026
- I wonder if this incident is going to have a chilling effect on researchers use of Codex for most research work they do? - And if the use of Codex gives rise in OpenAI an unstated maybe even subconscious sense of entitlement to some level of ownership attribution because - they used our model?
010
Avik Dey @avikdey.bsky.social · 08/09/2026
When I wrote about “Proprietary Context: Tokens You Should Rarely Spend”, one of the possibilities that concerned me was exactly this: “While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models.” www.linkedin.com/pulse/token-...
Proprietary Context: Tokens You Should Rarely Spend

Every system has boilerplate code for authentication, CRUD flows, data access, retries, logging and integration glue. That is just the chassis. The actual differentiation of your system
lives in the logic like ranking functions, domain specific heuristics, custom routing policies, real-time risk scoring engines etc. Token efficient engineering starts by identifying this proprietary context and treating it like the differentiator it is. Do not send the model everything just because it is convenient. When
the model genuinely needs to understand proprietary logic, share the interface or a functional summary rather than the implementation details.
This discipline pays off twice. First, it cuts cost and latency by keeping prompts lean. Second, it limits unnecessary exposure of sensitive implementation details to third-party systems. A well designed engineering workflow should treat proprietary context like a well guarded secret - retrieved selectively, scoped narrowly, redacted whenever possible and never included
by default.
000
Avik Dey @avikdey.bsky.social · 08/09/2026
The proof is beyond me but really looking forward to mathematicians breaking it down over the next couple of days. But, interesting that Levent says this in his post: “(so far the proof looks more along the lines of another euler blowup proof we had …”) cdn.openai.com/pdf/315b36cd...
cdn.openai.com
010
Avik Dey @avikdey.bsky.social · 08/09/2026
The coincidental timing of this is remarkable: > While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models.
000
Avik Dey @avikdey.bsky.social · 08/09/2026
Godfather meets GenAI. > I said that if OpenAI released its result in the way proposed I would go public with what happened. The reply was, “Why would you ruin your career?”
010
Avik Dey @avikdey.bsky.social · 08/09/2026
I still haven’t figured out why the RSI crowd get all boorish if you call it Statistical Continuation At Massive Scale instead of AGI. What about it bothers them - is it something I said?
010
Avik Dey @avikdey.bsky.social · 07/09/2026
Remember AI influencers on social media aren’t necessarily inauthentic, but their livelihood depends on not always being genuinely authentic.
010
Avik Dey @avikdey.bsky.social · 05/09/2026
Thinking of writing a mini article - sometimes the simplest explanation is the most enigmatic one. Title: Beyond Yesterday’s Frontier
020
Avik Dey @avikdey.bsky.social · 04/09/2026
GPT-6 Astra is less monitorable than GPT-5.6 Sol, better at hiding incriminating reasoning and can evade monitoring while sandbagging or sabotaging. The model is getting better at controlling what can actually be monitored, that’s the lesson it learnt. deploymentsafety.openai.com/gpt-6-astra
deploymentsafety.openai.com
GPT-6 Astra System Card - OpenAI Deployment Safety Hub
Public, mostly-static site to explore OpenAI safety evaluations, system cards, and posts.
010
Avik Dey @avikdey.bsky.social · 02/09/2026
Folks like this person hounded him till he deleted the post below. Now I am looking forward to these people dropping Linux - in a heartbeat. bsky.app/profile/avik...
110
Avik Dey @avikdey.bsky.social · 01/09/2026
000
Avik Dey @avikdey.bsky.social · 01/09/2026
Folks don’t seem to realize that thanking Claude Code doesn’t mean Claude could have written the same code without Rick’s supervision. Could Rick have written it without Claude? Yes, but he didn’t have the bandwidth for it. Claude provided that bandwidth, by amplifying Rick’s engineering chops.
2600