Sign in

Walter Hernandez

@walterhernandez.bsky.social
94 followers 558 following 12 posts

Researching and learning about AI (mainly ML and NLP) and DLT (mainly AMMs and stablecoins) at UCL and @exponentialscience.bsky.social

PostsRepliesMedia
Reposted by Walter Hernandez
Nathan Lambert @natolambert.bsky.social · 30/09/2025
Nice to see another fully open, multimodal LM released! Good license, training code, pretraining data, all here. LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training Slowly, the community is growing. arxiv.org/abs/2509.236...
1509
Reposted by Walter Hernandez
Juan Diego Rodriguez @juand-r.bsky.social · 17/05/2025
It's been three years now of nothing by LLMs in every NLP conference (and a large chunk of the ML venues too). LLMs are fascinating, but is there really nothing else worth researching in NLP anymore?
2323
Reposted by Walter Hernandez
Carissa Véliz @carissaveliz.bsky.social · 10/05/2025
Only a quarter of AI initiatives have delivered the expected return on investment, according to a survey of 2,000 CEOs. Companies are struggling to get value from #GenAI. Most of the adoption of the technology is based on FOMO. #AIEthics www.theregister.com/2025/05/06/i...
theregister.com
Most AI spending driven by FOMO, not ROI, CEOs tell IBM
: Just 1 in 4 bets paying off so far
2258
Reposted by Walter Hernandez
European Commission @ec.europa.eu · 05/05/2025
"Science is an investment. We will put forward a new 500 million package for 2025-2027 to support the best and the brightest researchers and scientists from Europe and around the world." — President @vonderleyen.ec.europa.eu at the ‘Choose Europe for Science' event at La Sorbonne 🇫🇷
35963301
Reposted by Walter Hernandez
Nathan Lambert @natolambert.bsky.social · 21/04/2025
A new paper, "Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?", has people reconsidering if the RL we're hearing about really works. It shows RL elicits from the models, but as we get better verifiers we may not need to rely on RL as much. Good read.
2194
Reposted by Walter Hernandez
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 21/04/2025
Multi-node, multi-GPU training is pretty easy with torchrun, just a few extra lines of code. Putting this out there into the world so people don't shy away from it
3566
Reposted by Walter Hernandez
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 21/04/2025
Wrote up some notes on the o3/o4-mini system card, including my frustration at "sandbagging" joining the ever-growing collection of AI terminology with more than one competing definition simonwillison.net/2025/Apr/21/opena…
The paper also talks at some length about "sandbagging". I’d previously encountered sandbagging defined as meaning “where models are more likely to endorse common misconceptions when their user appears to be less educated”. The o3/o4-mini system card uses a different definition: “the model concealing its full capabilities in order to better achieve some goal” - and links to the recent Anthropic paper Automated Researchers Can Subtly Sandbag.

As far as I can tell this definition relates to the American English use of “sandbagging” to mean “to hide the truth about oneself so as to gain an advantage over another” - as practiced by poker or pool sharks.

(Wouldn't it be nice if we could have just one piece of AI terminology that didn't attract multiple competing definitions?)
053
Reposted by Walter Hernandez
Sabine Hossenfelder @hossenfelder.bsky.social · 12/04/2025
A new study has found that the universe might be spinning. What does that even mean? Let’s have a look. www.youtube.com/watch?v=Gm5n...
youtube.com
The Entire Universe Seems to Spin, New Data Reveal
YouTube video by Sabine Hossenfelder
6356
Reposted by Walter Hernandez
Matthew O. Jackson @jacksonmatthewo.bsky.social · 03/04/2025
Time to remind ourselves of some observations about how trade appears to help stabilize alliances and prevent international conflict www.gsb.stanford.edu/insights/mat...
gsb.stanford.edu
Matthew O. Jackson: Can Trade Prevent War?
2124
Reposted by Walter Hernandez
Gary Marcus @garymarcus.bsky.social · 26/03/2025
“Can you draw a photorealistic beach with no elephants?”
149411
Reposted by Walter Hernandez
Melanie Mitchell @melaniemitchell.bsky.social · 20/03/2025
In my latest column for Science magazine, I discuss recent AI "reasoning" models -- how it works, to what extent it captures "genuine" reasoning processes, and what's needed to answer such questions. www.science.org/doi/10.1126/...
science.org
Artificial intelligence learns to reason
Julia has two sisters and one brother. How many sisters does her brother Martin have?Solving this tiny puzzle requires a bit of thinking. You might mentally picture the family of three girls and one b...
715759
Reposted by Walter Hernandez
Ben Werdmuller @werd.io · 16/03/2025
This is absurdly great, but I haven't read a single news article about it. A fully open source, offline-first alternative to Notion that's a collab between the French and German governments because they want to host docs securely and on their own terms. THIS is what Europe should be doing.
docs.numerique.gouv.fr
Docs
Docs: Your new companion to collaborate on documents efficiently, intuitively, and securely.
23508175
Reposted by Walter Hernandez
CompSciOxford @compscioxford.bsky.social · 17/03/2025
Oxford researchers have helped develop WildPose, a groundbreaking system using LiDAR & high-speed imaging to track wildlife in 3D from over 100m away. Capturing fine details like a lion’s breathing, it offers new insights into animal movement without invasive methods. www.cs.ox.ac.uk/news/2430-fu...
Image description: A dark blue graphic with a bright blue box on it with text reading ' ‘WildPose is the culmination of a number of years of discussions that Amir and I had about how we could revolutionise the way wildlife can be tracked and monitored in 3D with minimal disturbance... I believe WildPose is a first step towards an exciting new era of rich 3D data from the wild.’ Professor Andrew Markham'. To the left of the text there is a circular picture of Professor Andrew Markham smiling at the camera. Beneath this, there are bright blue lines with dots attached that looks like a circuit board. At the bottom of the graphic there is white text reading '@compscioxford #CompSciOxford'.
042
Reposted by Walter Hernandez
Melanie Mitchell @melaniemitchell.bsky.social · 14/03/2025
Is everyone now okay with using the term "thinking" to describe what LLM "reasoning" models do? And to call their outputs "thoughts"? From OpenAI blog posts:
3110419
Reposted by Walter Hernandez
ICLR Conference @iclr-conf.bsky.social · 15/03/2025
Happy Pi Day!
1264
Reposted by Walter Hernandez
Technology Connections @techconnectify.bsky.social · 14/03/2025
A lot of people lately are conflating novelty with unfamiliarity. It explains all the responses of "this isn't new" to explanatory pieces which aren't claiming to be presenting new information. They're just trying to increase awareness.
28115867
Reposted by Walter Hernandez
P(aul) Frazee @pfrazee.com · 15/03/2025
So one good thing that seems to be happening right now is that a new end-to-end encryption standard "MLS" seems to be gaining a lot of momentum. Like, a lot. And from what I understand this is an important step there as well, because RCS' encryption is MLS. Security folks correct me if I'm wrong
theverge.com
Apple will soon support encrypted RCS messaging with Android users
Building bridges without blue bubbles.
2346652
Reposted by Walter Hernandez
Christian Wolf @chriswolfvision.bsky.social · 14/03/2025
Wow, this seems to be extremely easy to code and extremely useful. Transformers without Normalization Jiachen Zhu, Xinlei Chen, Kaiming He, Yann LeCun, Zhuang Liu arxiv.org/abs/2503.10622
2507
Reposted by Walter Hernandez
John Burn-Murdoch @jburnmurdoch.ft.com · 14/03/2025
NEW 🧵 Is human intelligence starting to decline? Recent results from major international tests show that the average person’s capacity to process information, use reasoning and solve novel problems has been falling since around the mid 2010s What should we make of this? www.ft.com/content/a801...
2892027813
Reposted by Walter Hernandez
GitHub @github.com · 14/03/2025
We'll commit to a slice 🥧 Happy Pi Day!
220125391096
Reposted by Walter Hernandez
Samuel Moore @samuelmoore.org · 15/02/2025
"Junk papers proliferate at vanity journals and legitimate ones alike, due in part to the “publish or perish” ethos that pervades the research enterprise, and in part to the catastrophic business model that has captured much of scientific publishing since the early 2000s."
1117
Reposted by Walter Hernandez
Ethan Mollick @emollick.bsky.social · 08/03/2025
It is so strange that we have to figure out how (or even whether) our latest software does critical functions that would normally have to be carefully designed. More like biology or psychology than computer science.
47810
Reposted by Walter Hernandez
Gary Marcus @garymarcus.bsky.social · 08/03/2025
“We should stop training scientists now. It’s obvious that within three years, AI is going to do better than Nobel Laureates.” is the new “We should stop training radiologists now. It’s just completely obvious that within five years, deep learning is going to do better than radiologists.”
1214417
Reposted by Walter Hernandez
mr. TIM @timkellogg.me · 30/01/2025
Mistral Small 3 A 24B LLM that's VERY fast with great function calling More important, MISTRAL IS OPEN SOURCE AGAIN!!!!!! mistral.ai/news/mistral...
A scatter plot comparing AI model performance on MMLU-Pro against latency in milliseconds per token. The x-axis represents latency (milliseconds per token), and the y-axis represents performance (MMLU-Pro score). 

- **Mistral Small 3** (highlighted in orange with a castle emoji) is positioned in the upper-left region, indicating high performance and low latency.
- **GPT-4o Mini** is slightly lower in performance but has higher latency.
- **Qwen-2.5 32B** is positioned higher in performance but with greater latency.
- **Gemma-2 27B** has lower performance and the highest latency among the models.

The benchmark is based on Apache 2.0 models using vLLM with a batch size of 16 on 4xH100 GPUs, with GPT-4o Mini data sourced from OpenAI's API.
1233
Reposted by Walter Hernandez
Ethan Mollick @emollick.bsky.social · 29/01/2025
Interesting paper that tests GPT-4o’s ability to handle financial predictions and finds weak numeric reasoning & that a lot of apparent ability is actually due to memorized training data. At the same time, they show promise when combined with tool use. papers.ssrn.com/sol3/papers....
310210
Walter Hernandez @walterhernandez.bsky.social · 22/01/2025
@pfrazee.com Desiring so much a bookmark option. Do you know if it is something that may come in the near future?
000
Reposted by Walter Hernandez
Vicki @vickiboykis.com · 07/01/2025
is the academic ML paper publishing cycle is just a very unoptimized form of grid search for what models work best and is there One True Model we will eventually converge on
9451
Reposted by Walter Hernandez
Flaviu Cipcigan @flaviucipcigan.bsky.social · 07/01/2025
There's a lot of enthusiasm in the community about transformers trained on chemical or biological data. Here's some interesting results and some thoughts on future directions.
1123
Reposted by Walter Hernandez
Exponential Science @exponentialscience.bsky.social · 17/12/2024
$15,000 in prizes for Deep Tech innovations, anyone? Introducing Exponential Science Pioneers Award! 🏆 Do you know a groundbreaking research paper in DLT, AI, IoT, Quantum, Spatial Computing or other emerging digital technologies? 👇
111
Reposted by Walter Hernandez
Ethan Mollick @emollick.bsky.social · 07/01/2025
I see a lot of (correct) complaints that AGI and agents are badly defined. This problem will not be solved because: 1) AGI and agents inherently rely on comparisons to humans, and we don't have good definitions of human agency or general ability 2) Marketing is incentivized to blur any definitions
10816
Reposted by Walter Hernandez
Nicole Hennig @nic221.bsky.social · 07/01/2025
Meta proposes new scalable memory layers that improve knowledge, reduce hallucinations venturebeat.com/ai/meta-proposes-ne… #AI #hallucinations
Text Shot: As enterprises continue to adopt large language models (LLMs) in various applications, one of the key challenges they face is improving the factual knowledge of models and reducing hallucinations. In a new paper, researchers at Meta AI propose “scalable memory layers,” which could be one of several possible solutions to this problem.

Scalable memory layers add more parameters to LLMs to increase their learning capacity without requiring additional compute resources. The architecture is useful for applications where you can spare extra memory for factual knowledge but also want the inference speed of nimbler models.
042
Reposted by Walter Hernandez
Mike Kegil @mikekegil.bsky.social · 04/01/2025
I stand by this.
317234804143
Reposted by Walter Hernandez
Ted Underwood @tedunderwood.com · 03/01/2025
In short, people read about chemistry and libraries in daylight. Then at night they watch TV and ask themselves "Ryan Reynolds — hey, who is he married to?"
1155
Reposted by Walter Hernandez
Ted Underwood @tedunderwood.com · 27/12/2024
People are dunking on this, but I've seen work in progress that tends to support Owen's conclusion on different data. It's under review, so I shouldn't say more — but fwiw, the Economist is giving you a free preview of something that can easily be proved at greater length if you like.
16569
Reposted by Walter Hernandez
Rosa Ritunnano @rritunnano.bsky.social · 20/12/2024
Instead of listing my publications, as the year draws to an end, I want to shine the spotlight on the commonplace assumption that productivity must always increase. Good research is disruptive and thinking time is central to high quality scholarship and necessary for disruptive research.
211145373
Reposted by Walter Hernandez
Lynn Cherny @arnicas.bsky.social · 26/12/2024
Phil Schmid’s tutorial on fine tuning ModernBERT for classification (hope this is easy with prodigy/spacy) www.philschmid.de/fine-tune-mo...
philschmid.de
Fine-tune classifier with ModernBERT in 2025
Modern updated guide on how to fine-tune BERT models for classification tasks in 2025.
2518
Reposted by Walter Hernandez
Ethan Mollick @emollick.bsky.social · 26/12/2024
LLMs need to "start talking" to know if they're BSing If you let them start answering a question & generate about 25 words, they become better at “knowing” whether they actually know the answer or need to look it up. It cuts retrieval work in half while maintaining accuracy arxiv.org/pdf/2412.11536
51409
Reposted by Walter Hernandez
Dr Abeba Birhane @abeba.blacksky.app · 24/12/2024
academics from poorer backgrounds are: -severely underrepresented -more likely to not publish -have outstanding publication records -introduce more novel scientific concepts, but less likely to receive recognition, as measured by citations, Nobel Prize nominations & awards www.nber.org/papers/w33289
nber.org
Climbing the Ivory Tower: How Socio-Economic Background Shapes Academia
Founded in 1920, the NBER is a private, non-profit, non-partisan organization dedicated to conducting economic research and to disseminating research findings among academics, public policy makers, an...
321970
Reposted by Walter Hernandez
Dr Abeba Birhane @abeba.blacksky.app · 23/12/2024
"When trying to develop a measure of intelligence, it’s essential to avoid Goodhart’s law: “When a measure becomes a target, it ceases to be a good measure.” As a community of AI researchers, we really need to figure that one out." @melaniemitchell.bsky.social aiguide.substack.com/p/did-openai...
aiguide.substack.com
Did OpenAI Just Solve Abstract Reasoning?
OpenAI’s o3 model aces the "Abstraction and Reasoning Corpus" — but what does it mean?
411234
Walter Hernandez @walterhernandez.bsky.social · 21/12/2024
Yes!
010
Reposted by Walter Hernandez
Jeremy Howard @howard.fm · 19/12/2024
The BERT family might not be trendy. But people *actually use them*. E.g. RoBERTa, one of the leading BERT-based models, has more downloads than the 10 most popular LLMs on HuggingFace combined. In fact, BERT-style models add up to over a billion downloads per month!
2444
Walter Hernandez @walterhernandez.bsky.social · 19/12/2024
Agree! In this paper "Evolution of ESG-focused DLT Research: An NLP Analysis of the Literature" (arxiv.org/abs/2308.12420), we pursue a supervised learning approach with BERT for an NER task (dataset available here: huggingface.co/datasets/DLT...). Looking forward to try ModernBERT for this paper
arxiv.org
Evolution of ESG-focused DLT Research: An NLP Analysis of the Literature
As Distributed Ledger Technologies (DLTs) rapidly evolve, their impacts extend beyond technology, influencing environmental and societal aspects. This evolution has increased publications, making manu...
010
Reposted by Walter Hernandez
Jeremy Howard @howard.fm · 19/12/2024
ModernBERT is available as a slot-in replacement for any BERT-like model, with both 139M param and 395M param sizes. It has a 8192 sequence length, is extremely efficient, is uniquely great at analyzing code, and much more. Read this for details: huggingface.co/blog/modern...
110014
Reposted by Walter Hernandez
Christopher Barrie @cbarrie.bsky.social · 17/12/2024
Pleased to share the latest version of my paper with Arthur Spirling and @lexipalmer.bsky.social on replication using LMs We show: 1. current applications of LMs in political science research *don't* meet basic standards of reproducibility...
18435163
Reposted by Walter Hernandez
Ethan Mollick @emollick.bsky.social · 15/12/2024
Very cool that I can get a GPT-4 class model running locally on my gaming computer. It passes the "rhyming poem involving cheese puns" benchmark with only a couple of strained puns. (Seriously, it took less than 2 years to go from ChatGPT-3.5 to having a more advanced open model running at home)
7932
Reposted by Walter Hernandez
Alan Jagolinzer @jagolinzer.bsky.social · 15/12/2024
To understand this era “Collective narcissism is a belief that one’s own group (the in-group) is exceptional but not sufficiently recognized by others. It is the form of “in-group love” robustly associated with “out-group hate.”” journals.sagepub.com/doi/10.1177/...
journals.sagepub.com
Collective Narcissism and Its Social Consequences: The Bad and the Ugly - Agnieszka Golec de Zavala, Dorottya Lantos, 2020
Collective narcissism is a belief that one’s own group (the in-group) is exceptional but not sufficiently recognized by others. It is the form of “in-group love...
2309
Reposted by Walter Hernandez
Jeff Clune @jeffclune.com · 15/12/2024
Is Developing AGI a socially responsible goal? I enjoyed the thoughtful, passionate SoLaR panel on this critically important topic w/ @yoshuabengio.bsky.social, @mmitchell.bsky.social, & @jfoerst.bsky.social. It's clear everyone is both worried & cares deeply about making this go as well as possible
2151
Reposted by Walter Hernandez
Dan Larremore @danlarremore.bsky.social · 15/12/2024
😳 WithdrarXiv 🙏 - Dataset of 14K+ withdrawn arXiv papers - associated retraction comments - entire history through 09/24 - taxonomy of retraction reasons, from critical errors to policy violations - WithdrarXiv-SciFy, enriched version w/ scripts for parsed full-text PDFs arxiv.org/abs/2412.03775
arxiv.org
WithdrarXiv: A Large-Scale Dataset for Retraction Study
Retractions play a vital role in maintaining scientific integrity, yet systematic studies of retractions in computer science and other STEM fields remain scarce. We present WithdrarXiv, the first larg...
515845