mr. TIM @timkellogg.me · 51mInterview from the VP at OpenAI responsible for Jalepeño, their custom AI chip text: morethanmoore.substack.com/p/interview-... video: youtu.be/8s7uYtCM1bcmorethanmoore.substack.comInterview with Richard Ho, OpenAIVP Hardware for Jalapeño 170
mr. TIM @timkellogg.me · 10husing chatgpt dot, and i think i get it 1. it’s the new chatgpt 2. chatgpt work/codex already works great so they’re moving away from subsidies 3. dot is the new loss leader 0140
mr. TIM @timkellogg.me · 11hContext Language Models New agent architecture where the LLM can edit its own context it seems to have emergent capabilities, creates its own memory management & organization algorithms, and coordinates multi agents github.com/facebookrese... 101329
mr. TIM @timkellogg.me · 13hSol 6.1 on ARC-AGI-3 curves backwards as effort levels go up. this is a bit of an artifact of ARC-AGI-3, where scoring higher equates to fewer moves, ending the game with fewer turns 1220
mr. TIM @timkellogg.me · 16hGemini 4 Argon Same price as Sol 6.1, rolling out to selected cyberdefenders blog.google/innovation-a... 6432
mr. TIM @timkellogg.me · 18hi agree with the ai companies on a lot, but this distillation thing is just like the record companies vs the public 2150
mr. TIM @timkellogg.me · 22hGoogle’s ops is so good that Gemini was down a day ago and nobody even noticed 3470
mr. TIM @timkellogg.me · 23hAnthropic would greatly benefit from partnering with some neocloud to offer some open weights models as part of subscriptions & enterprise i don’t see this happening, but they should do it rather than write attack “research”, actually influence how GLM is deployed & consumed 2331
mr. TIM @timkellogg.me · 30/09/2026talking to some people in South Dakota. General sentiment is - AI bad - it’ll take decision making away from people - electricity prices, noise (not water use or land use) - no mention of jobs - no mention of x-risk 2421
mr. TIM @timkellogg.me · 30/09/2026i prefer to see these graphs on a log scale for cost effort vs cost should be a logarithmic function, and when drawn on a log axis it shows up as a straight line for that line: * slope — efficiency when scaling up to multiagents * straightness — predictable scaling * relative position — the usual 1281
mr. TIM @timkellogg.me · 30/09/2026Enterprise and Business Codex users can use GLM-5.3 Flash and Kimi K3 natively through the Codex B2B Marketplace via a Baseten <> OpenAI partnership x.com/baseten/stat...x.comBaseten (@baseten) on XAnnouncing our partnership with OpenAI 1110
mr. TIM @timkellogg.me · 29/09/2026Anthropic releases a detailed research advertisement for GLM-5.3 as an alternative to Fable in cybersecurity www.anthropic.com/research/glm...anthropic.comGLM-5.3 and the spread of advanced cyber capabilitiesGLM-5.3 can autonomously build end-to-end cyber exploits, but unlike other frontier models, it was released without meaningful safeguards to limit misuse. 51044
mr. TIM @timkellogg.me · 29/09/2026yes, if the Decisions API is multimodal (likely) then its a big improvement over Jev, even if its worse 4531
mr. TIM @timkellogg.me · 29/09/2026OpenAI is offering ultrafast mode (8x faster, 300 tok/s) for Astra to use it, you need to subscribe to a new usage tier — Pro 500 — 25x usage of Plus (presumably $500/mo) chatgpt.com/pricing/chatgpt.comPricing | ChatGPTPaid plans (Plus, Pro, Business, and Enterprise) are priced per user per month. We offer monthly plans for Plus, Pro, and Business and annual plans for Business and Enterprise. 1130
mr. TIM @timkellogg.me · 29/09/2026OpenAI Decisions API A Jev-clone on Luna. 150ms, pricing unknown thenewstack.io/openai-decis...thenewstack.ioOpenAI answers TypeSafe's Jev with a Decision API built on LunaOpenAI's Decision API, built on its small Luna model, returns predefined answers with confidence scores in 150 milliseconds. Pricing is still unknown. 4483
mr. TIM @timkellogg.me · 29/09/2026Dots: A GrokBot / OpenClaw type experience in Codex powered by Astra openai.com/index/how-we...openai.comHow we build safety, security, and privacy into dotsLearn how OpenAI builds safety, security, and privacy into dots—always-on agents with controls, secure workspaces, and checks before they act. 050
mr. TIM @timkellogg.me · 29/09/2026GPT-6.1 Sol Close-to-Astra at 20% the price openai.com/index/introd... 4611
mr. TIM @timkellogg.me · 29/09/2026i see this as a good thing, despite that it’s definitely bad in the short term subscriptions created a bubble such that you couldn’t afford not to be on a sub but subs also had stricter rules, e.g. anthropic kicking off openclaw users this is a move toward Subscription = Capped API 1110
mr. TIM @timkellogg.me · 29/09/2026reminder that Anthropic still hasn’t caught up to OpenAI. 6 > 5.5 6962
mr. TIM @timkellogg.me · 29/09/2026also, $34B of the $42B shortfall was an accounting quirk caused by the rapid increase in valuation Ed is arguing in regards to operating costs, which is hardly even relevant here 3181
mr. TIM @timkellogg.me · 29/09/2026NanoGPT pretraining runs now take 39.9 seconds for a 124M model(!!) now, 124M is *tiny* so it might not seem relevant. But much of the gains in LLM pretraining are data quality and a great way to find quality data is to train on it & measure the lift 3928
mr. TIM @timkellogg.me · 29/09/2026Today is DevDay, and OpenAI is bringing back the $200 subscription, but at half the value you get from it also, I’m curious about (d), new things that don’t detract from usage (Jev?) 4231
mr. TIM @timkellogg.me · 29/09/2026who called it “midtraining” instead of just “training”? @dorialexander.bsky.social there must be an answer 460
mr. TIM @timkellogg.me · 29/09/2026OpenAI scrapped their plan to launch Astra 6.1 over concerns around deceptive behavior www.wsj.com/tech/ai/open...wsj.comExclusive | OpenAI Scraps Release of New AI Model Over Safety ConcernsThe model, dubbed GPT-6.1 Astra, was due to debut inside ChatGPT and Codex in October. 10677
Reposted by mr. TIMriley @riguh.bsky.social · 28/09/2026Sonnet 5.5 is out! But if you’re on Amazon, only if you’re prepared to let them run it anywhere in the world. Not awesome for the many companies that enforce data sovereignty constraints. docs.aws.amazon.com/bedrock/late...docs.aws.amazon.comAnthropic - Amazon BedrockThe following Anthropic models are available in Amazon Bedrock: 141
mr. TIM @timkellogg.me · 28/09/2026progressive disclosure is when you put your pronouns in your bio 1350
Reposted by mr. TIMsoftmacs @softmacs.bsky.social · 28/09/2026Anthropic is just serving up a single model with a pay-what-you-want pricing scheme really 0173
mr. TIM @timkellogg.me · 28/09/2026Claude 5.5 Sonnet is live and it’s roughly Opus 5.5 but cheaper www.anthropic.com/claude-sonne... 2020623
mr. TIM @timkellogg.me · 28/09/2026this statement: “{OpenAI|Anthropic} released cheaper model X because they are short on compute” does not make sense if you believe in the Jevons Paradox 6511
mr. TIM @timkellogg.me · 28/09/2026OpenAI’s dev day is on Tuesday. They have “too much to launch”, meaning they’ll likely spill over the launches into Monday & Wednesday. I think we can expect another GPT-6 model, possibly Aeon (highly persistent), along with new products that make use of the persistent behavior 3260
mr. TIM @timkellogg.me · 28/09/2026given the existence of things like embeddings and PCA, we should be able to plot the vibe shift here “AI go rogue” is a very different conversation from “stochastic parrot”. things are shifting 1170
mr. TIM @timkellogg.me · 28/09/2026seems inevitable tbqh software engineer is (actually) hard, and too many people never operated at the level you have to operate now so yeah, laziness.. 4551
mr. TIM @timkellogg.me · 28/09/2026seems likely that Sol-6 is a much smaller model than Sol-5.6 also seems likely that Opus-5.5 is smaller (a lot?) than Opus-5 i think labs are finding ways to replicate a lot of that big model feel in smaller models, which is a good sign for that cognitive core dream 6821
mr. TIM @timkellogg.me · 27/09/2026i think GPT-6 Sol is a good model, but i don’t like coding with it, at least not directly 5110
mr. TIM @timkellogg.me · 27/09/2026one of the dangers in RSI is you don’t have to make models that people like when you sell models, they have to be aligned enough for people to actually buy them with RSI, the goal is entirely different 2233
mr. TIM @timkellogg.me · 27/09/2026i would say, in most contexts it is probably unwise to focus on risks you’ve never experienced e.g. in software engineering i’ve developed a bit of a gag reflex whenever an engineer is solving a problem that doesn’t exist yet trouble is the severity of x-risks puts them in a unique position 6571
Reposted by mr. TIMEugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 26/09/2026"These people are really annoying and/or weird" is not actually an argument against their point 9945
mr. TIM @timkellogg.me · 27/09/2026this by no means solves alignment, but a whole lot of misaligned behaviors stem from reward hacking, so this is a very big deal 5464
mr. TIM @timkellogg.me · 27/09/2026RLMs & Program Agents I'm experimenting with this idea, Program Agents, inside of DeepSeek Harness (DSH) They're like RLMs, except without the LLM. Agents are writing such large blocks of code, what if they just never exited? github.com/tkellogg/dsh... 4292
mr. TIM @timkellogg.me · 26/09/2026New from ArtificialAnalysis: The price of compute per KG set to surpass the price of cocaine in 2029 5404
mr. TIM @timkellogg.me · 26/09/2026is openai releases Jevstra at dev day on tuesday, that’ll be quite a statement about the level of RSI we’re at 1110
mr. TIM @timkellogg.me · 26/09/2026GDM employee quit with a rather thoughtful essay robert.ocallahan.org/2026/09/good... 1563