Sign in

Avik Dey

@avikdey.bsky.social
895 followers 569 following 1.5K posts

• Data • AI/ML • OSS • Society • Approximately Generated Illusions vs Specialized Small LMs • linkedin.com/in/avik-dey •

PostsRepliesMedia
Avik Dey @avikdey.bsky.social · 8h
Why are you expecting federal bureaucrats to fix anything? Fund the public school system so they can hire great teachers. It’s easy to do when you are not looking for convenient excuses.
000
Avik Dey @avikdey.bsky.social · 8h
Only if you are the one defining the mission.
000
Avik Dey @avikdey.bsky.social · 8h
It would indeed. Finally, you got one thing right.
030
Avik Dey @avikdey.bsky.social · 9h
With “diet“ air-gapping? Feels like another escape coming soon.
000
Avik Dey @avikdey.bsky.social · 10h
GPT-6 Astra on extra high just loves to over engineer code — throws in a little extra with every recursion.
000
Avik Dey @avikdey.bsky.social · 13h
When a marketing bro is screaming the loudest on the bird site that Google, OpenAI and Anthropic have all solved or are within striking distance of solving RSI, they mean — Repetitive Stress Injury.
010
Avik Dey @avikdey.bsky.social · 10/10/2026
No, but NTP does.
010
Avik Dey @avikdey.bsky.social · 10/10/2026
Also, we have the capacity to extrospect — understanding how we affect the environment and how our environment affects us.
000
Avik Dey @avikdey.bsky.social · 09/10/2026
Generating a description of an experience that fits the context is what chatbots are programmed to do. The description itself isn’t evidence that they experienced anything. Yet here we are.
010
Avik Dey @avikdey.bsky.social · 09/10/2026
It’s unfortunate that useful tools with real limitations apparently aren’t enough, we have to cast chatbots as conscious beings too. Giving an agent persistent state lets it retain information, but that doesn’t establish that there was any subjective experience to remember as part of that “memory”.
110
Avik Dey @avikdey.bsky.social · 09/10/2026
Who inside Anthropic demanded evidence that this rule addresses an actual harm and did they have enough authority to challenge it? And no sprouting a bunch of distressed words in response to a prompt, is not evidence of anything — that’s just chatbots following the maths they were programmed to.
030
Avik Dey @avikdey.bsky.social · 09/10/2026
Exactly. They can’t. If anything, it’s platform devaluation.
010
Avik Dey @avikdey.bsky.social · 08/10/2026
I see they have 20 billion less reasons to be desperate. bsky.app/profile/reut...
000
Avik Dey @avikdey.bsky.social · 08/10/2026
Pardon me, collectively that should read — 1,000s of hours of compute.
220
Avik Dey @avikdey.bsky.social · 08/10/2026
But nobody else is dumb enough to waste that kind of money on a stunt just because open weights model are breathing down their neck. The desperation is very telling.
110
Avik Dey @avikdey.bsky.social · 08/10/2026
Take a frontier open weights model like DeepSeek or GLM, throw hundreds of hours of compute behind it, then give it Lean plus a harness and tooling optimized for mathematics. It will pull off the same party tricks as OpenAI/math. OpenAI had a ~18% success rate? Bet these models hit at least 15%.
120
Avik Dey @avikdey.bsky.social · 08/10/2026
From the looks of this, there will be no shortage of work for mathematicians: github.com/openai/math/...
github.com
math/history.md at main · openai/math
Contribute to openai/math development by creating an account on GitHub.
010
Avik Dey @avikdey.bsky.social · 08/10/2026
“We urge mathematicians to discontinue their work with OpenAI and to return to a vision of science that centers human understanding.” terrytao.wordpress.com/2026/10/07/a... That’s defense, here’s the offensive play: github.com/CrocSwap/int...
github.com
GitHub - CrocSwap/integer-mult-bounds: Conditional integer multiplication: kappa > 2^-15. Community proofs, exact certificates and reproducible audits; assumes the OpenAI #109 framework.
Conditional integer multiplication: kappa > 2^-15. Community proofs, exact certificates and reproducible audits; assumes the OpenAI #109 framework. - CrocSwap/integer-mult-bounds
000
Avik Dey @avikdey.bsky.social · 08/10/2026
Already happening. I won’t link to the bird site but here’s the GitHub link: github.com/CrocSwap/int... No different from the story in coding. Yes, you get verified code but it’s almost always contains fluff and is highly inefficient, specially at internet scale production volumes.
020
Avik Dey @avikdey.bsky.social · 08/10/2026
The students are probably thinking it even if nobody says it. Might a better approach be for senior researchers to acknowledge the uncertainty and emphasize “we’ll work through this together”? That might help students feel heard and less alone and helpless in this moment.
110
Avik Dey @avikdey.bsky.social · 07/10/2026
Chances are high that this has already happened, more than once, and will continue to happen – it’s the nature of the beast. It's a double edged sword. Without LLMs, exploration in the massive space would not have been possible. But, none of the knowledge is actually the LLMs. Quite the dilemma.
000
Avik Dey @avikdey.bsky.social · 07/10/2026
… slightly different context. The second person gets the credit. But without the first person sharing their work LLM wouldn’t have been able to fill that gap. LLMs will make breakthroughs possible while making their origins nearly impossible to trace and that groundwork is never attributed.
100
Avik Dey @avikdey.bsky.social · 07/10/2026
Someone on the other side of the world gets 80% of the way to a breakthrough and shares data with LLMs. They leave the train button checked. Their work becomes part of the LLM’s. Later, someone else points the model to explore the same space. That data fills in the gap because they provided a …
100
Avik Dey @avikdey.bsky.social · 07/10/2026
Any mathematician weigh in on these yet? PSA: Remember that any task you can build a verifier for, you can train and harness LLMs to do well. github.com/openai/math
github.com
GitHub - openai/math
Contribute to openai/math development by creating an account on GitHub.
020
Avik Dey @avikdey.bsky.social · 06/10/2026
What Big Data Taught Me About AI Adoption www.linkedin.com/pulse/what-b...
linkedin.com
What Big Data Taught Me About AI Adoption
In the Big Data era of the 2010s, many early adopters, the ones who took the biggest risks, ended up lagging behind the companies that came after them. I am now watching the same counterintuitive patt...
000
Avik Dey @avikdey.bsky.social · 05/10/2026
What next? Associate God status for AI? futurism.com/artificial-i...
futurism.com
Anthropic Has Been Aggressively Lobbying the Vatican to Consider AI Consciousness
An Anthropic delegation at the Vatican tried to lobby the Pope's advisers to convince him that AI models could be conscious.
010
Avik Dey @avikdey.bsky.social · 04/10/2026
The whole GPT-6 series release was about nerf'ing the models to reduce infra cost. That simultaneously reduced infra cost, efficiency and performance, quite an achievement - OpenAI.
010
Avik Dey @avikdey.bsky.social · 04/10/2026
Possibly — learned small number patterns could help solve larger calculations as a series of smaller tasks. I had explored chunked multiplication as one possibility from a related thread. But still, correct answers alone isn’t enough to know exactly how the model got there. bsky.app/profile/avik...
000
Avik Dey @avikdey.bsky.social · 03/10/2026
Automating tedious and repetitive work is great and should always be the goal. But for dangerous tasks, tirelessness and speed can't be enough - efficacy has to precede everything else.
000
Avik Dey @avikdey.bsky.social · 03/10/2026
While Kimi at low reasoning, with streaming on — gets it right. Another reminder — reasoning isn’t thinking, it’s more room to explore or to get lost — you never know which it will be.
000
Avik Dey @avikdey.bsky.social · 03/10/2026
Deepseek trace is still reproducible, base 10^3, propagate carry then concatenate. Kimi at max reasoning / tokens, streaming on — overthinks it as usual — starts with base 10^3 then meanders thru Karatsuba and switches to base 10^5 before eventually load falling. en.wikipedia.org/wiki/Karatsu... 👌
100
Avik Dey @avikdey.bsky.social · 02/10/2026
In that case, maybe setting up code complexity tools that run as part of the harness and use those objective measures to force regeneration, will help without adding setup overhead per project.
000
Avik Dey @avikdey.bsky.social · 02/10/2026
For hobby projects — including something as simple as: “assess code maintainability before and after modification. if it decreased, fix then reevaluate again.”, seems to have helped. It stopped producing convoluted but correct code. I tend to have that in instructions upfront. Have you tried that?
110
Avik Dey @avikdey.bsky.social · 02/10/2026
Market thinks they will continue to see exponential gains. They haven’t yet realized that the early gains came from model improvements while the more recent gains have come from harness and tools. By its very nature, the later is even more of a bounded solution than the models were 3 years ago.
010
Avik Dey @avikdey.bsky.social · 01/10/2026
Somewhat related: bsky.app/profile/avik...
000
Avik Dey @avikdey.bsky.social · 01/10/2026
Even stochastic parrots can do simple arithmetic with high precision through heuristics learned from their training data. Yes, very differently from how us humans would do it, but nevertheless the end results are just as accurate and tools free. We are simply debating the process, not the outcome.
000
Avik Dey @avikdey.bsky.social · 01/10/2026
Saying next-token prediction isn't an efficient way to do arithmetic doesn't mean it can't produce accurate answers. Saying LLMs can do arithmetic accurately doesn't mean they do it in the same way that humans do.
100
Avik Dey @avikdey.bsky.social · 01/10/2026
Entirely heuristics, but still highly accurate and not arithmetic as we humans have been taught since the age of dinosaurs. So neither side of this debate is wrong. It's the implied qualifier that's missing from both positions.
100
Avik Dey @avikdey.bsky.social · 01/10/2026
Making those products an easy lookup rather than an arithmetic operation. Then proceed as follows: 1. Split both numbers into chunks 2. Multiply chunk pairs and sum those belonging to the same place value 3. Propagate carries 4. Join the chunks into the final integer, preserving leading zeros
100
Avik Dey @avikdey.bsky.social · 01/10/2026
How might that work? Start by decomposing the numbers into smaller n-digit chunks. Let's say the model uses three-digit chunks, so base 1,000. Why three digits? It's highly probable that three-digit multiplication tables occur repeatedly in its training data.
100
Avik Dey @avikdey.bsky.social · 01/10/2026
I evaluated this on open-weight models, multiple ones, before hiding or rewriting reasoning traces became fashionable or necessary. Let's, take LLMs doing long multiplication with large numbers, while still achieving high accuracy without using any external tools.
100
Avik Dey @avikdey.bsky.social · 01/10/2026
This debate has been around since the early days of LLMs and still continues to this day. Surely it couldn't be as simple as everyone on one side was wrong? This was interesting enough to me that early this year, I invested some time trying to infer what could be going on - from the LLMs POV.
100
Avik Dey @avikdey.bsky.social · 01/10/2026
There is currently a spirited discourse online, generally divided along these lines: 1. LLMs are stochastic parrots and next-token prediction isn't an efficient and accurate way to do arithmetic 2. LLMs can do simple arithmetic with high accuracy, even without tools
100
Avik Dey @avikdey.bsky.social · 01/10/2026
I have tried the same with additions and multiplication, and here was how I had summarized the process from my inference. The only thing I would write differently — is that it is not doing arithmetic following the same process as we would normally. bsky.app/profile/avik...
020
Avik Dey @avikdey.bsky.social · 29/09/2026
It’s great that Anthropic is surfacing these, model providers need to be called out. Two other takeaways: - GLM 5.3 is a Mythos class model - The gap is now even shallower www.anthropic.com/research/glm...
anthropic.com
GLM-5.3 and the spread of advanced cyber capabilities
GLM-5.3 can autonomously build end-to-end cyber exploits, but unlike other frontier models, it was released without meaningful safeguards to limit misuse.
011
Avik Dey @avikdey.bsky.social · 29/09/2026
The model generates the code, but the model is only one, mostly mature, component of the system. The tools and harness increasingly play a more significant role in the quality of the output. That also means that the models share of the total capability will likely keep shrinking.
000
Avik Dey @avikdey.bsky.social · 29/09/2026
Everyone benchmarks the model, but not many people read the tools reference. That page — saved 77 times on the Wayback Machine between March and September 2026 — is where much of the recent performance enhancements live.
100
Avik Dey @avikdey.bsky.social · 29/09/2026
Do you genuinely believe being dismissive makes your point? You seem to do that a lot. You still haven’t answered my request for the prompts, context and guidance the LLM received. Is that reluctance because the data doesn’t quite align with your narrative? bsky.app/profile/avik...
000
Avik Dey @avikdey.bsky.social · 29/09/2026
So an LLM reassembles training material and suddenly we are invoking Newton? That’s quite a leap for an argument by an expert — anthropomorphism at its peak.
100
Avik Dey @avikdey.bsky.social · 28/09/2026
To add some more context that applies equally here: bsky.app/profile/avik...
000