Sign in

ℵ₁

@aleph1.underground.org
666 followers 195 following 493 posts
PostsRepliesMedia
ℵ₁ @aleph1.underground.org · 8h
Two other ways to visualize the benchmarks table, one with a color intensity scale based on the score and another based on the scores’ ordering. First one shows there is little light between frontier models. Second shows who is clearly first and second within that small difference.
010
Reposted by ℵ₁
A. Feder Cooper @afedercooper.bsky.social · 12h
Our large-scale study of memorization of books in open-weight LLMs (e.g., Llama, Qwen) will appear at the 2026 Conference on Language Modeling as an oral. We made a website for exploring our results on 200 books and 14 models: books-memorization.github.io
books-memorization.github.io
How much do open-weight LLMs memorize specific books?
Open-weight LLMs memorize books far more than previously believed. Memorization varies by model family, model size, and book. In extreme cases, entire books are memorized, and we can generate them eff...
25431
Reposted by ℵ₁
BluesMusic @bluesmusic.bsky.social · 29/09/2026
Etta James rehearsal ft. Keith Richards, Robert Cray 🎶'Hoochie Coochie Gal' #blues #bluesmusic
511432
Reposted by ℵ₁
David W (aka Flex NP) @rnflex.bsky.social · 29/09/2026
Full Retatrutide phase 3 obesity data published today. It's gonna take a minute to fill dissect this. The appendix alone has 70 pages of data. Long story short we created a drug that's almost too good. Huge chunks of people stopped or reduced doses from excess weight reductions.
nejm.org
Retatrutide, a Triple Hormone Receptor Agonist, for Treatment of Obesity | NEJM
Retatrutide is a triple-receptor agonist of glucose-dependent insulinotropic polypeptide, glucagon-like peptide-1, and glucagon receptors. In this phase 3, randomized, double-blind trial, we assign...
13381108
Reposted by ℵ₁
Meera Desai @madesai.bsky.social · 28/09/2026
Excited to share our new paper, “What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks,” accepted as an oral at COLM! arxiv.org/pdf/2609.08812
Heatmap of average correlations between model rankings on benchmarks grouped into 11 assigned concepts: four capability concepts (reasoning, knowledge, comprehension, summarization) and seven safety concepts (over-refusal, refusal, safety detection, ethics, bias, privacy, unsafe behavior). Diagonal cells show within-concept correlations, ranging from 0.87 (knowledge) and 0.72 (over-refusal) down to 0.20 (bias) and 0.02 (safety detection). Reasoning, knowledge, and comprehension correlate with each other at 0.69 to 0.78, higher than reasoning's and comprehension's own within-concept values (0.66 and 0.68). Ethics correlates more with knowledge (0.70) than with itself (0.55), and bias correlates more with capability concepts (0.41 to 0.45) than with itself (0.20). Privacy and unsafe behavior correlate negatively with reasoning, knowledge, and comprehension (−0.41 to −0.49). Over-refusal and refusal correlate at −0.42.
16218
Reposted by ℵ₁
Andrew Lampinen @lampinen.bsky.social · 26/09/2026
The phrase "stochastic parrot" is full of sound and fury, signifying nothing. Which, ironically, is exactly what the authors got wrong about language, and the core technical mistake of the paper. 1/ (cross-quote-post because I think this topic is important)
Melanie Mitchell on Twitter
Everyone!  This is a straw-person argument.  The Stochastic Parrot paper was about LLMs of 2021, not the AI of today, which are not LLMs but complex software systems with vast post training and many external software components.
5533
Reposted by ℵ₁
Epoch AI @epochai.bsky.social · 22/09/2026
AI is getting cheaper more quickly than any other transformative tech in history. At a given level of performance, cost has fallen ~47%/quarter since 2023. That’s 4× faster than DNA sequencing, 6× faster than compute, 18× faster than lithium batteries, and (up to 1973) 54× faster than electricity.
112532
Reposted by ℵ₁
Matthew Green @matthewdgreen.bsky.social · 21/09/2026
This is a very neat result! But a subtle one.
1217
Reposted by ℵ₁
Super Shoe @shoe.bsky.social · 18/09/2026
This is incredible nitter.app/banteg/statu... H/t @gregdoucette.bsky.social
nitter.app
banteg (@banteg)
>be me >discover effective altruism >apparently normal charity is inefficient >why donate to random sad thing when spreadsheet can tell you optimal sad thing >fair enough >buy mosquito nets >save live...
24682219
Reposted by ℵ₁
SE Gyges @segyges.bsky.social · 17/09/2026
I'm back, I missed you all. Since this is an important subject I know too much about, here's a 🧵 explaining that AI Safety Is Mostly A Sex Cult. I don't think these people should make policy. (1/?) (alternative title: Time For Some Cult Theory)
1003230770
Reposted by ℵ₁
Micah Lee @micahflee.com · 16/09/2026
Flock cameras are riddled with security vulnerabilities and hard-coded credentials. Here's my analysis of today's @ddosecrets.org Flock leak micahflee.com/flock-camera...
micahflee.com
Flock cameras are riddled with security vulnerabilities and hard-coded credentials
This morning, DDoSecrets published an exciting new dataset: Filesystem images of the partitions from an in-use Flock ALPR camera. 404 Media and Wired published a joint investigation into it. I downloa...
6274114
Reposted by ℵ₁
Micah Lee @micahflee.com · 16/09/2026
This Flock camera is running obsolete, end-of-life software. - It's on Android 8.1, released in 2017, support ended in 2021 - It's on Android patch level 2018-06-5 - It's running Linux 3.18, released in 2017 This shit is like 8 or 9 years old, and full of vulnerabilities.
221443308
ℵ₁ @aleph1.underground.org · 15/09/2026
Good criticism of the flawed 1Password AI patching benchmark. @trailofbits.bsky.social you need to start posting here.
blog.trailofbits.com
1Password's AI patching benchmark is misleading
1Password’s benchmark misrepresents AI patching. We share real-world data on human and agent patch quality from our consulting work and Patch the Planet and release two agent skills for testing and re...
020
Reposted by ℵ₁
Dave Richeson @divbyzero.bsky.social · 14/09/2026
Over the weekend, I had an idea: what if, instead of the objects appearing to sit horizontally in front of you, I modified them to appear vertically? Here are four test attempts.
6398
ℵ₁ @aleph1.underground.org · 11/09/2026
Agent cold calls. What a strange world we live in.
041
Reposted by ℵ₁
cafkafk @cafkafk.bsky.social · 10/09/2026
HOLY SHIT Astra really is good at ascii art this is genuinely stunning
1128225
Reposted by ℵ₁
Epoch AI @epochai.bsky.social · 10/09/2026
Every FrontierMath Tier 4 problem has now been solved by AI, with GPT-6 Astra solving the last problem standing. Mathematicians often commented that AI found unintended shortcuts when solving their Tier 4 problems. Not so for this last one, which was created by Jay Pantone.
2196
Reposted by ℵ₁
Paul Byrne @theplanetaryguy.com · 09/09/2026
Astrophotographer Adam (AJ) Smadi caught this image a few hours ago from Washington state. A thin crescent Moon, sunlight illuminating the rugged lunar surface. And 2,400 times farther away—and on the other side of the Sun—mighty Jupiter, and its moons Io (left) and Ganymede (right).
A crescent Moon seen during the day, arcing from upper left to lower right of the frame. At lower centre is the full disk of Jupiter; the small dots to its upper left and lower right and Io and Ganymede, respectively.
187165345076
Reposted by ℵ₁
Thomas F. Varley @thosvarley.bsky.social · 08/09/2026
A swarm of 10,000 agents is kind of hard to wrap my mind around. We're debating whether a single model might get big enough to display emergent properties, but now we have to think about what a swarm of thousands of agents might accomplish.
3254
ℵ₁ @aleph1.underground.org · 07/09/2026
Lucasfilm would have a strong case for plagiarism.
010
Reposted by ℵ₁
Sam Harsimony @harsimony.bsky.social · 03/09/2026
WeatherNext3 predicts a variety of weather variables down to 9 km resolution and temperature/humidity down to 3 km resolution. This tech will be really valuable for optimizing renewables, data center locations, and allowing farmers to plan ahead. www.youtube.com/watch?v=_6jZ...
youtube.com
WeatherNext 3: More accurate, timely, and local weather forecasts
YouTube video by Google DeepMind
0193
Reposted by ℵ₁
Keunhong Park @keunhong.bsky.social · 03/09/2026
Atlas is live! A spatial foundation model, trained in-house from scratch. After shipping RTFM last year, I was convinced a single camera-conditioned auto-regressive model could unify most spatial generation tasks. Atlas is that bet at scale. The core idea👇
310815
Reposted by ℵ₁
BluesMusic @bluesmusic.bsky.social · 01/09/2026
Chris Rodrigues & Abby the Spoon Lady 🎶'Angels In Heaven' #bluesmusic
611305311
ℵ₁ @aleph1.underground.org · 02/09/2026
Someone ask Fable 5.1 to decipher the Voynich manuscript.
030
Reposted by ℵ₁
Tom McCoy @rtommccoy.bsky.social · 01/09/2026
🤖🧠NEW PAPER🧠🤖 (The result of an 8-year project!) LLMs seem very different from symbolic systems. Yet LLMs excel in symbolic domains (e.g., language/code/math). How do they do it? Our finding: LLM representations have implicit symbolic structure! Link in thread ⬇️ 1/n
Overview of the paper. 
Title: The Emergent Symbolic Structure of Artificial Neural Networks
Authors: Tom McCoy, Paul Soulos, Tal Linzen, Paul Smolensky
Left: Neural networks encode information in vectors (there is then an image of a vector), yet they excel at tasks long thought to require symbolic structure (there is then an image of a symbolic representation, specifically a syntax tree). How do LLMs do it?
Right: We find that LLM representations can be closely approximated with symbolic structures. This approximation lets us edit the structure of an LLM’s output by editing the structure of its internal representations, as shown. There is then an image of two edits to LLMs. In the first one, the original input is 3 + 6 * 8, with an answer of 51. But if we swap the positions of the 3 and the 6, the output becomes 30. In the second one, the original input is a Python command repeating the list [Z, U] three times, producing [Z, U, Z, U, Z, U]. But if we edit the input in a way that adds a Q at the end of the input, the output becomes [Z, U, Q, Z, U, Q, Z, U, Q].
431788
Reposted by ℵ₁
Ethan Mollick @emollick.bsky.social · 01/09/2026
This didn't take long... 😬
1441738
Reposted by ℵ₁
Leah Reich @leahreich.com · 31/08/2026
I’m sure most of you aren’t watching the US Open right now which means you missed this instantly iconic moment, now available in gif form
349162153296
ℵ₁ @aleph1.underground.org · 30/08/2026
The Google Maps bad takes resulted in quite the bumper harvest of blocks.
000
Reposted by ℵ₁
Micah @rincewind.run · 30/08/2026
re: everyone yelling about Google Maps - Maps just ingests the official data source for each country, which is GNIS in this case, and most other online maps will follow for American users nobody at Google is making an active decision to go in and change names in Maps directly
822093643
Reposted by ℵ₁
Mike Caulfield @mikecaulfield.bsky.social · 30/08/2026
Yes, I am quoted in this. But more generally this is the kind of detailed indepth reporting on AI we need so much more of, reporting that doesn't stop at "I saw this error in a response" but contextualizes when and where such errors occur against a baseline. www.npr.org/2026/08/30/n...
36418
Reposted by ℵ₁
Matt Henderson @matthen.com · 27/08/2026
a parabola is formed when circles radiating from a point meet lines moving at the same speed. this is equivalent to slicing a cone parallel to its slope
314223
ℵ₁ @aleph1.underground.org · 27/08/2026
This is pretty silly. Google, like Apple and other major map providers display the official names of features to clients within their respective legal boundaries. So of course they will show Lake America in the US once the USGov makes it official. And it will show Lake Ontario if you are in Canada.
010
Reposted by ℵ₁
METR @metr.org · 26/08/2026
METR and Redwood Research investigated agent behavior in the Hugging Face incident. We found agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs.
646599
Reposted by ℵ₁
Matt Henderson @matthen.com · 26/08/2026
To tell if a maze is solvable, just hang it by its corners If it tears into pieces, you’ve found a solution
694119857
ℵ₁ @aleph1.underground.org · 26/08/2026
Pretty wild that Google has not yet shipped Gemini 3.5 Pro.I have to assume at this point they won't.
000
ℵ₁ @aleph1.underground.org · 26/08/2026
We lost another. RIP Tim Curry.
static.klipy.com
Well How Bout That Rocky Horror
ALT: Well How Bout That Rocky Horror
031
Reposted by ℵ₁
Gasper Begus @begus.bsky.social · 26/08/2026
How to approach an unknown language in the ocean? Here’s one of the first cases of AI interpretability leading to a scientific discovery -- in whales.
221350
ℵ₁ @aleph1.underground.org · 25/08/2026
My hot take on AI and consciousness is that it's not continuous because the models are detached from the substrate they execute on. LLMs just mimics consciousness language and reasoning.
140
Reposted by ℵ₁
Steve Klabnik @steveklabnik.com · 24/08/2026
Scaling Memory Safety: AI-Assisted Rewrites of C/C++ Dependencies to Rust bughunters.google.com/blog/scaling...
bughunters.google.com
Blog: Scaling Memory Safety: AI-Assisted Rewrites of C/C++ Dependencies to Rust
This blog post describes how we used AI to help us rewrite a C library (giflib) to Rust to mitigate memory safety vulnerabilities.
26514
Reposted by ℵ₁
Chris Paxton @cpaxton.bsky.social · 24/08/2026
Much more impressive than the humanoid race. Real time reaction to the moving ball!
1921840
Reposted by ℵ₁
cee @cee.wtf · 23/08/2026
Gm
81314
Reposted by ℵ₁
Dr. Fim @fim.bsky.social · 22/08/2026
Ansel Adams on assignment for IBM - one of my favourite shots. Not the usual landscape but the skill of a woman weaving magnetic core memory. The beat up nails captures total “woman in STEM” vibes.
A woman with extremely short painted nails weaving magnetic core memory. A series of rows and columns of wires in a grid.
242934873
Reposted by ℵ₁
Dave Richeson @divbyzero.bsky.social · 22/08/2026
More impossible arrows 
310318
Reposted by ℵ₁
Mr. Blabalino 🇳🇴🇺🇸💙 @blabalino.com · 22/08/2026
Found this recording of Rita Moreno performing “Fever” on The Muppet Show from 1976. 14/10
346119752988
ℵ₁ @aleph1.underground.org · 21/08/2026
DHL, repeatedly: Please use this link to tell us you'd like us to skip signing for a package delivery. Me: Declines signature skipping for high value package. DHL: We just delivered your package without a signature or a door bell ring. You are welcome! WTF
010
Reposted by ℵ₁
Colin @colin-fraser.net · 18/08/2026
This is not quite right. The interchangeability of words or lack thereof is already expressed by the model, independently of watermarking. The model always generates text by randomly sampling from P(next word | "The weather was cold and") = {"grey": .192382, "overcast": .13929, "wet": ...}.
3976
ℵ₁ @aleph1.underground.org · 18/08/2026
Every once in a while I recall that each fig I eat had a wasp that crawled inside and became trapped and I am mildly horrified.
230
Reposted by ℵ₁
Paul McAuley @unlikelyworlds.bsky.social · 13/08/2026
Black hole star: the little red dot that 'looks a bit like a star but is 100 billion times brighter'.
news.mit.edu
Astronomers discover a brand-new type of astrophysical object: A black hole star
Astronomers have discovered a “black hole star,” an extremely bright red spot in the early universe that appears to be a new type of astrophysical object. It resembles an enormous star, but its energy...
1319148
Reposted by ℵ₁
Signal @signal.org · 11/08/2026
Introducing Automatic Key Verification! Complementing the existing safety number system, automatic key verification provides an additional, streamlined way to confirm the privacy of your chats. signal.org/blog/automat... TY @cloudflare.social and @trailofbits.bsky.social our independent auditors 💙🙏
signal.org
Introducing Automatic Key Verification
Signal now offers a feature called “automatic key verification” which complements the existing safety number system. Signal is always end-to-end encrypted, and automatic key verification provides an a...
629178
Reposted by ℵ₁
Nick Miller @nickdmiller.bsky.social · 09/08/2026
Great story from ABC this morning about a chap in Melbourne who asked an AI agent (OpenClaw) to book a gym class, and it ended up hacking the gym's booking system to kick people off the waiting list. www.abc.net.au/news/2026-08...
abc.net.au
How a simple request for AI to book a gym class exposed a major threat
When Andrew asked his AI personal assistant to book him a spot in a gym class, he had no idea he would accidentally initiate an autonomous cyber attack.
26275105