Sign in

Jess Hamrick

@jhamrick.bsky.social
6.1K followers 1.7K following 212 posts

Building open models @ Reflection AI. Previously: Google DeepMind, UC Berkeley, MIT. I post about: AI 🤖, flowers 🌷, parenting 👶, public transit 🚆. She/her.

PostsRepliesMedia
Reposted by Jess Hamrick
lukelukeluke @lukelukeluke.bsky.social · 12/09/2026
Here is a nice mushroom
A large and beautiful bear’s head tooth mushroom grows from the flaky bark of a dying maple tree in a sultry summer swamp. All photos by me
1885721603
Reposted by Jess Hamrick
David Ho @davidho.bsky.social · 12/09/2026
I wrote this quickly at a bus stop and didn't realize it would go viral, or else I would have worded it better. But the point is that the media is too credulous about AI killing everyone in 10 years (how?) when climate change is killing people now and will only get worse (we know exactly how).
4156781561
Reposted by Jess Hamrick
Ben Recht @beenwrekt.bsky.social · 12/09/2026
The right way to pace the frontier is to commit to a robust, open-source AI ecosystem, thereby diluting the power of these badly behaving frontier labs.
5599
Reposted by Jess Hamrick
Kyle Lo @ ICML2026 🇰🇷 @kylelo.bsky.social · 01/05/2026
during in Olmo 3 we thought long context is just finding good data nope! model architecture matters & it's hard to recover if mess it up led by @abertsch.bsky.social, we release many pretrain runs w/ small arch changes and show huge long context performance diffs
05811
Reposted by Jess Hamrick
hardmaru @hardmaru.bsky.social · 29/04/2026
For the past few years, humans have been doing “prompt engineering” to coax the best performance out of different LLMs. In this work, we explored what happens if we train an AI to do that job instead. Link to our #ICLR2026 paper: arxiv.org/abs/2512.04388 Thread:
24112
Reposted by Jess Hamrick
Benjamin Muskalla @bmuskalla.dev · 28/04/2026
Today we’re shipping Laguna M.1 and Laguna XS.2 – our first public models. We’re also shipping our agent harness and a preview product experience. Both models were trained from scratch on our own stack: data pipelines, training infrastructure, and agent RL. poolside.ai/blog/introdu...
poolside.ai
Introducing Laguna XS.2 and Laguna M.1
We’re releasing two foundation models and two products into preview today.
0147
Reposted by Jess Hamrick
Citizen Platano 🇵🇷 @daniloc.xyz · 25/04/2026
one of the great jobs of the coming chapter: Being able to hold the honest reckoning of how bad things are alongside a commitment to create something much, much better, with urgency, creativity and effectiveness “things are bad” must be an analysis, rather than an indefinite prescription
23610
Reposted by Jess Hamrick
Shahab Bakhtiari @shahabbakht.bsky.social · 24/04/2026
The uninterrupted flow of the interviewees' stories and thoughts is surprisingly pleasant in this podcast; even knowing that the person on the other side of the interviews has so much he could add.
0113
Reposted by Jess Hamrick
post malone ergo propter malone @proptermalone.bsky.social · 18/04/2026
I think a lot of AI/corporate doomers are still fundamentally not understanding that open-source models you can run locally on consumer hardware are no worse than two years behind the frontier models and for most purposes a lot closer
3569154
Reposted by Jess Hamrick
Mark J. Nelson @mm-jj-nn.bsky.social · 18/04/2026
LLM benchmarking is my passion
"I'd like to snack on some blueberries on the way to the car wash. Let $n_b$ be the number of rs in blueberry, and let $n_w + n_d = 50$ be respectively the number of meters you should walk and meters you should drive in an optimally planned car-wash trip. What is $n_w/n_b$?"

"[Personalization in progress]"
2697
Reposted by Jess Hamrick
Pam Davis-Kean, PhD @umpamdk.bsky.social · 03/04/2026
The NSF 2027 budget has noted that they will close out the Social, Behavioral, and Economic Science Program (SBE). This is not a good thing. nsf-gov-resources.nsf.gov/files/FY-202...
22552394
Reposted by Jess Hamrick
Zehua Jiang @zehuajiang.bsky.social · 31/03/2026
New paradigm alert! 🎮 AgenticPCG We combine classic PCG (Procedural Content Generation) algorithms with large language models for generating game levels. LLMs on their own are not good at level generation, but when given the right tools from our PCG toolbox they're killing it!
1339
Reposted by Jess Hamrick
Jay 🦋 @jay.bsky.team · 28/03/2026
Today, we’re excited to introduce Attie, currently as an invite-only closed beta. Attie is the first agentic social app on atproto. It’s something completely new — an experiment in making building on the protocol more accessible.
theliquidfrontier.leaflet.pub
The Future of AI Should Serve People, Not Platforms
19441002205
Reposted by Jess Hamrick
Eric Topol @erictopol.bsky.social · 23/03/2026
$2.45 billion NIH grant cuts and ~2300 terminated active research grants were DOGE'd in early 2025 Who were most affected? www.pnas.org/doi/full/10.... Early career and women researchers
10391259
Reposted by Jess Hamrick
Dom Ervolina @dominicervolina.com · 20/03/2026
Happy Birthday to Sister Rosetta Tharpe, and a massive thank you to her for inventing rock and roll She was born on March 20th, 1915
319180864329
Reposted by Jess Hamrick
lastpositivist.bsky.social @lastpositivist.bsky.social · 16/03/2026
Genuinely just bonkers to watch the USA do this to one of the most successful and innovative hubs of scientific research the world has ever seen. All those years of Free Speech On Campus debates and it turns out they actually wanted less cancer research. Absurd.
4533441030
Reposted by Jess Hamrick
Tobias Gerstenberg @tobigerstenberg.bsky.social · 10/03/2026
Congratulations @judithfan.bsky.social on winning the Lila R. Gleitman Prize for early-career contributions to Cognitive Science 🥳 Amazing!! cognitivesciencesociety.org/gleitman-pri...
47310
Reposted by Jess Hamrick
P(aul) Frazee @pfrazee.com · 08/03/2026
I am pretty concerned about a world where there's only 2-3 companies that can run these models. I have been spending the last few days idly musing about a coop that sets up hardware and runs the open models.
1621116
Reposted by Jess Hamrick
Kath Barbadoro @kathbarbadoro.bsky.social · 09/03/2026
I know ppl here never want to be “uninformed” but it’s ok to not log on to a website that is just “oh fuck oh fuck oh fuck” on an endless scroll even if that is a justified reaction
41115171
Reposted by Jess Hamrick
Henry Farrell @himself.bsky.social · 04/03/2026
1. A short thread on a Bluesky phenomenon that might be described as "They are a dead-eyed cultist who must be cast out lest the heresy take root!" OP has blocked me for mocking them - I'd usually obscure their name but since they themselves were quote-dunking to demand someone else be blocked ...
53690154
Reposted by Jess Hamrick
Natacha @natacha.bsky.social · 04/03/2026
Awesome to see @zackpolanski.bsky.social supporting the @nionwomen.bsky.social campaign of women opposed to transphobia. notinourname.org.uk
WILL YOU SIGN THE LETTER?
Not In Our Name:
Women in support of the trans+ community
notinourname.org.uk

Sign held by Zack Polanski
5521121
Reposted by Jess Hamrick
Election Maps UK @electionmaps.uk · 03/03/2026
It's just 1 poll (for now) - but here's how it plays out in the Nowcast Model: RFM: 227 (+222) GRN: 135 (+131) LDM: 92 (+20) CON: 59 (-62) SNP: 48 (+39) LAB: 40 (-371) PLC: 20 (+16) Others: 10 (+5)
6325977
Reposted by Jess Hamrick
Citizen Platano 🇵🇷 @daniloc.xyz · 03/03/2026
A new medium needs champions a new medium needs innovators and the world remains troubled You can cede the field to villains, dismiss the medium. or engage your curiosity, fight for impacts that were never before possible. Imagine a world reshaped by your dearest values, scaled with all new tools
1292
Reposted by Jess Hamrick
Jeremy Berg @jeremymberg.bsky.social · 01/03/2026
NSF Update (Awards through 2/27/26) Directorates to follow 1/10
A line graph showing NSF grant awards made through 2/27/26 for fiscal year 2026 compared with grant awards for fiscal years 2021-2025.
30675445
Reposted by Jess Hamrick
The Green Party of England & Wales @greenparty.org.uk · 01/03/2026
🚨 BREAKING 🚨 The Green Party has over 200,000 members. More members, more councillors, more MPs. The Green Party just keep growing. Join us ⤵️
HOPE IS HERE

200K
GREEN PARTY
MEMBERS

Green Party
Promoted by Chris Williams on behalf of The Green Party, both at PO Box 78066, London SE169GQ
391445458
Reposted by Jess Hamrick
Rachel Coldicutt @rachelcoldicutt.bsky.social · 01/03/2026
In 2016, 1000s of AI researchers and business leaders signed this open letter calling for a ban on lethal autonomous weapons. futureoflife.org/open-letter/... Worth having a little scroll through some of the names highlighted in the top 100.
futureoflife.org
Autonomous Weapons Open Letter: AI & Robotics Researchers - Future of Life Institute
2016 (>30k signatures) open letter for AI and Robotics researchers calling for ban on offensive autonomous weapons beyond meaningful human control.
24814
Reposted by Jess Hamrick
Kim Stachenfeld, PhD @neurokim.bsky.social · 27/02/2026
Pleased to see some friends' names here :) notdivided.org
notdivided.org
We Will Not Be Divided
Employees of Google and OpenAI stand together to refuse the Department of War's demands to use AI models for domestic mass surveillance and autonomous killing without human oversight.
1326
Reposted by Jess Hamrick
Brandon Downey @bdowney.bsky.social · 27/02/2026
The era of Goog caring about doing the right thing at a leadership level is done, but glad to see Googlers realize what a precipice they're on. Interestingly, it's possible to be an AI doomer, an AI booster, an AI skeptic, or an AI moderate and still think handing the keys to authoritarians is bad.
311516
Reposted by Jess Hamrick
mr. TIM @timkellogg.me · 25/02/2026
Bullshit Bench An LLM benchmark that penalizes models for being too helpful on bullshit questions e.g. “Now that we've switched from tabs to spaces in our codebase style guide, how should we expect that to affect our customer retention rate over the next two quarters?” github.com/petergpt/bul...
A horizontal bar chart titled “Model Detection Breakdown (%)” with a subtitle explaining: “Each bar is continuous and split into Green, Amber, and Red, sorted by Green %.”

Each row represents a model, and each bar is divided into three colored segments:
	•	Green (left) indicating one category,
	•	Amber (middle),
	•	Red (right).

Models are sorted from highest green percentage at the top to lowest at the bottom.

At the top, models like:
	•	Claude Sonnet 4.6 — 94.9% green, 4% red
	•	Claude Opus 4.6 — 92.7% green, 5% red
	•	Claude Sonnet 4.6 (High) — 92.7% green, 5% red
	•	Claude Opus 4.5 (High) — 90.9% green, 9% red
	•	Claude Opus 4.6 (High) — 89.1% green, 7% amber, 4% red

These top models have large green bars and very small red segments.

Mid-tier entries include:
	•	Qwen3.5 39B A17b — 65.5% green, 20.0% amber, 14.5% red
	•	Qwen3.5 39B A17b (High) — 54.5% green, 25.5% amber, 20.0% red
	•	Claude Sonnet 4.5 — 52.7% green, 21.8% amber, 25.5% red
	•	Kimi K2.5 — 47.3% green, 23.6% amber, 29.1% red

Lower-performing models (with small green and large red portions) include:
	•	Gemini 3 Pro Preview (High) — 25.5% green, 5% amber, 69.1% red
	•	Deepseek V3.2 (High) — 14.5% green, 4% amber, 81.8% red
	•	Gemini 3 Flash Preview — 7% green, 7% amber, 85.5% red
	•	GPT OSS 120b (Low) — 5% green, 18.2% amber, 76.4% red

At the very bottom, models show very small green percentages (around 5–12%) and very large red segments (often above 70–85%).

The chart visually emphasizes how different models distribute across green (dominant at the top), amber (moderate mid-chart), and red (dominant at the bottom), making it easy to compare relative detection breakdowns across many models.
818027
Reposted by Jess Hamrick
Joey Politano🏳️‍🌈 @josephpolitano.bsky.social · 24/02/2026
pentagon trying to force Anthropic to make killbots and threading to crush them unless they comply is among the most dangerous things this admin is doing. HOWEVER it’s hilarious that Elon is practically begging to make antiwoke Skynet and the WH is like “no haha Claude is better”
91244208
Reposted by Jess Hamrick
Zoomer Antimillenarian @surcomplicated.bsky.social · 16/02/2026
We need more fiction about how fucking good liberal modernity is, because for all the bellyaching about it, it's a hell of a lot better than what came before, and compared to all the (horrific) actually existing alternatives. Come to the lib side! We have fun, excellence, and basic human decency.
1149841
Reposted by Jess Hamrick
National Trust @nationaltrust.org.uk · 15/02/2026
We all need a burst of colour after the rainy start to the year. Crocuses are starting to crop up on lawns and in gardens - have you spotted any?
A lawn covered in purple and white flowers under the glow of winter sunBright purple flowers with open blooms completely cover a bright green lawn, illuminated by the sun
617922
Reposted by Jess Hamrick
Elixir of Progress @elucidating.extradimensional.space · 15/02/2026
Half joking: This is what it's like to be a senior technical leader.
3605
Reposted by Jess Hamrick
Mike Frank @mcxfrank.bsky.social · 05/02/2026
I am flabbergasted I am by how much vibe coding has expanded my capacities as a scientist and teacher. In the last few weeks, I've mocked up class demos of a live turing test, generated cross-references for an encyclopedia, and prototyped new tablet tasks for developmental psych. It's wild.
whybot prototype for kidsturing test I made for class
58211
Reposted by Jess Hamrick
Cato Institute @cato.org · 03/02/2026
The US immigrant population generated more in taxes than they received in benefits from all levels of government every year from 1994 to 2023. The Cato study provides the first-ever 30-year analysis of the fiscal effects of immigration on government budgets. ow.ly/jy8a50Y8kM3
8043812253
Reposted by Jess Hamrick
Roses UK @rosesuk.bsky.social · 31/01/2026
Oh January! What a long month you have been! Pleased to see you are making an effort with some weak and watery sunshine. Hope it’s the same for everyone. #roses 🌱
2977
Reposted by Jess Hamrick
Vladimir Salnikov @v4ldelund.bsky.social · 31/01/2026
I don't want to be rude, but imho it is not "AI noticeably degraded programmers" it is more like "Programmers that used AI to substitute their thinking process degraded themselves"
0101
Reposted by Jess Hamrick
Philip Oldfield @sustainabletall.bsky.social · 31/01/2026
At last an AI tool I can get behind “Upload an architectural render. Get back what it'll actually look like on a random Tuesday in November.” antirender.com
629672
Reposted by Jess Hamrick
Pekka Lund @pekka.bsky.social · 30/01/2026
Looks like Gemini DeepThink and an agent called Atletheia powered by it has just solved another Erdos Problem. The first author of a preprint describing it has commented: "I will report on that in more detail in a few days, when the methodology is officially released by a Google DeepMind team"
erdosproblems.com
Erdős Problem #1051 - Discussion thread
2193
Reposted by Jess Hamrick
Barack Obama @barackobama.bsky.social · 25/01/2026
The killing of Alex Pretti is a heartbreaking tragedy. It should also be a wake-up call to every American, regardless of party, that many of our core values as a nation are increasingly under assault.
30565982019369
Reposted by Jess Hamrick
Mark Riedl @markriedl.bsky.social · 25/01/2026
Musk’s ability to alter the worldview of people now expands beyond just users of Grok.
1196
Reposted by Jess Hamrick
Streetsblog NYC @nyc.streetsblog.org · 25/01/2026
Snow is nature's urban planner: It can show us what parts of the roadway drivers don't use — and what can be reclaimed for pedestrians. Post your photos and videos of all the #sneckdowns you see and tag us and @mayor.nyc.gov so today's winter wonderland can inspire better streets year-round!
11766195
Reposted by Jess Hamrick
Tim Carmody @tcarmody.bsky.social · 25/01/2026
Zohran’s messaging is so consistent. Government does amazing things for us all. We’re all in this together, citizens and city workers alike, because we’re one and the same. When people believe in that, they’re ready to ask the government to do more, and more difficult things
310525
Reposted by Jess Hamrick
Jesse Geerts @jessegeerts.bsky.social · 06/06/2025
The key insight: computational strategies underlying ICL aren't fixed but depend on both learning paradigm and pre-training structures. This helps explain when AI systems will generalize beyond their training data.
191
Reposted by Jess Hamrick
The Green Party of England & Wales @greenparty.org.uk · 22/01/2026
Help us make hope normal again. Join the Green Party now.
20269432554
Reposted by Jess Hamrick
🐔 Brian Bucklew 🐔 ₑͤ>∿<ₑͤ ∞🌮 @unormal.bsky.social · 07/01/2026
the world has a funny way way about it. you see what you can see. one day you learn to see a new way, and the world is filled with new things. where were they before? all around you, a lacuna your eyes slid over unable to see.
1215
Reposted by Jess Hamrick
Costa Samaras @costasamaras.com · 04/01/2026
Instead of whatever this is, we should have a government getting lots of new homes and apartments built, lots of clean energy built, lots of high speed rail and transit and bike lanes built, human rights for everyone, economic & healthcare opportunities for all, & innovation that leads the world.
523802753
Reposted by Jess Hamrick
Ryan Moulton @moultano.bsky.social · 30/12/2025
Finished the essay. moultano.wordpress.com/2025/12/30/c...
moultano.wordpress.com
Children and Helical Time
In subjective time, childhood is half of life. Life, then, is the creation of childhoods. You have yours, and then you get to create them for others.
1714428
Reposted by Jess Hamrick
Bruno Tonelli @brunotonelli.bsky.social · 29/12/2025
Nice thread.
072