Sign in

Dario

@dariogargas.bsky.social
676 followers 1.1K following 184 posts

Senior AI researcher at BSC. Random thinker at home.

PostsRepliesMedia
Dario @dariogargas.bsky.social · 10/02/2026
Open position for an AI Research Area Director at BSC, one of the largest European supercomputing facilities! Apply if you are into: * Open research 🤲 * Responsible science 🤔 * State-of-the-art AI advances 🧠 * Young & active teams 🛝 GPUs are on us. Barcelona is waiting 😉 www.bsc.es/join-us/job-...
bsc.es
69_26_DIR_AII_RAD
Reference: 69_26_DIR_AII_RAD Job title: Research Area Director – Artificial Intelligence (RE4/R4) About BSC The Barcelona Supercomputing Center - Centro Naciona
040
Dario @dariogargas.bsky.social · 03/06/2025
Are you into Chip Design, EDA or just like to do RTL code for fun? Check out the largest benchmarking of LLMs for Verilog generation: TuRTLe 🐢 It includes 40 open LLMs evaluated on 4 benchmarks, following 5 tasks. And its only growing! huggingface.co/spaces/HPAI-... arxiv.org/abs/2504.01986
072
Dario @dariogargas.bsky.social · 22/05/2025
So many healthcare LLMs, and yet so little information! Check out this table summarizing contributions, and find more details in our latest pre-print: arxiv.org/abs/2505.04388
020
Dario @dariogargas.bsky.social · 21/05/2025
The Aloe Beta preprint includes full details on data & training setup. Plus four different evaluation methods (including medical expert). Plus a risk assessment of healthcare LLMs. Two years of work condensed in a few pages, figures and tables. Love open research! huggingface.co/papers/2505....
000
Dario @dariogargas.bsky.social · 06/05/2025
We just opened two MLOps Engineer positions at @bsc-cns.bsky.social Our active and young research team needs someone to help sustain and improve our services, including HPC clusters, automated pipelines, artifact managements and much more! Are you up for the challenge? www.bsc.es/join-us/job-...
bsc.es
350_25_CS_AIR_RE2
Reference: 350_25_CS_AIR_RE2 Job title: Research Engineer - AI Factory (RE2) About BSC The Barcelona Supercomputing Center - Centro Nacional de Supercomputación (BSC-
000
Dario @dariogargas.bsky.social · 06/05/2025
Last week our team presented this at NAACL. Check out the beautiful poster they put together 😍
010
Dario @dariogargas.bsky.social · 11/04/2025
Working on a project for evaluating embryo quality using in-vitro fertilization data. A random forest using morphokinetic features of embryo evolution visually annotated by experts, and a CNN directly using static images get similar performance. Separately AND together. I find it surprising...
100
Dario @dariogargas.bsky.social · 09/04/2025
There are quite a lot of researchers who a so preoccupied with whether or not they could get the funding, they don't stop to think if they should. Being chased by dinosaurs and writing grants. Same thing.
010
Dario @dariogargas.bsky.social · 04/04/2025
How expensive 🫰 is it to get the best LLM performance? How much cash needs to burn 💸 to get reliable responses? Pareto optimal plots answer these questions. Our research shows it is economically feasible and scalable to achieve O1 level performance at a fraction of the cost. buff.ly/ji1VHiV
120
Dario @dariogargas.bsky.social · 01/04/2025
Our LLM safety project, Egida, reached 2K downloads 😀 It includes +60K safety questions expanded with jailbreaking prompts. The four models trained (and released) show strong signs of safety alignment and generalization capacity. Check out the 🤗 HF page and the paper for details! buff.ly/kxFVyl2
000
Dario @dariogargas.bsky.social · 01/04/2025
Today we release the TuRTLe leaderboard! 🐢 Are you in the Chip Design or EDA business? Wanna know which LLMs are best for the task? By integrating 4 benchmarks, TuRTLe evaluates: * Syntax * Functionality * Synthesizability * Power, Performance and Area metrics huggingface.co/spaces/HPAI-...
huggingface.co
TuRTLe Leaderboard - a Hugging Face Space by HPAI-BSC
A Unified Evaluation of LLMs for RTL Generation.
010
Dario @dariogargas.bsky.social · 03/03/2025
MIR is Spain's medical entrance exam. Best students reach an estimate accuracy of +90. Two or three every year. We took MIR, '20-'24 to test open LLMs. Llama 3.1 based models, like Aloe, reach +80 in accuracy. Deepseep R1 reaches +88. Boosted by a RAG system, 92. buff.ly/4bbbXMw buff.ly/4hLrhBV
buff.ly
HPAI-BSC/CareQA · Datasets at Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
100
Dario @dariogargas.bsky.social · 28/02/2025
After listening to the latest @fallofcivspod.bsky.social episode about the Mongolian Empire, by @paulcooper34.bsky.social , I realized Mongols and the Fremen from Dune share remarkable similarities. Skilled warriors adapted to harsh environments, taking over a society they don't want to adopt.
110
Dario @dariogargas.bsky.social · 21/02/2025
Human evaluation of LLMs is close to saturation. Models have been optimized so much for plausibility, that we are unable to tell good from bad. Only experts in expert domains can see a meaningful difference.
000
Dario @dariogargas.bsky.social · 21/02/2025
After a year working on LLM evaluation, our benchmarking paper is finally out (to be presented at NAACL 2025). Main lessons: * All LLM evals are wrong, some are slightly useful. * Goodhart's law. All the time. Everywhere. * Do lots of different evals and hope for the best.
buff.ly
Automatic Evaluation of Healthcare LLMs Beyond Question-Answering
Current Large Language Models (LLMs) benchmarks are often based on open-ended or close-ended QA evaluations, avoiding the requirement of human labor. Close-ended measurements evaluate the factuality…
130
Dario @dariogargas.bsky.social · 21/02/2025
Evaluating LLMs is a bit like paleontology. Trying to understand the behavior of very complex entities by observing only noisy and partial evidence. How do paleontologists deal with the uncertainty and frustration? Do they also feel like doing alchemy instead of science?
000
Dario @dariogargas.bsky.social · 19/02/2025
Wisdom from my 6y old daughter: "A king is a just person disguised as king."
000
Dario @dariogargas.bsky.social · 19/02/2025
Over and over again I keep finding @sarahooker.bsky.social papers to reference. This time about ELO rankings. She's always 2-3 years ahead...
010
Dario @dariogargas.bsky.social · 18/02/2025
So many keywords around LLM training, its easy to get lost. For an incoming paper, did this little visual summary. Would you change anything?
Summary of LLM learning methods
000
Dario @dariogargas.bsky.social · 17/02/2025
5th International Workshop on Computational Aspects of Deep Learning (CADL) to be held in conjunction with ISC-HPC 2025. 10 days to go, and an award to be decided! Submit your paper and join us sites.google.com/view/cadl2025/
sites.google.com
CADL 2025
Advancing AI Through Efficient Computing Over the past decade, Deep Learning (DL) has revolutionized numerous research fields, transforming AI into a computational science where massive models are tra...
000
Dario @dariogargas.bsky.social · 12/02/2025
Only two weeks until the deadline! Submit your paper and see you in Germany :)
000
Dario @dariogargas.bsky.social · 10/02/2025
Bring it on. Totally prepared for another lockdown.
000
Dario @dariogargas.bsky.social · 03/02/2025
Tired of not having enough time for reading cause all the writing I have to do.
000
Dario @dariogargas.bsky.social · 30/01/2025
Historical quotes from deep learning: "Don't be a hero" Andrej Karpathy "Attention is all you need" Vaswani et al "We Have No Moat And neither does OpenAI" Google Engineer
230
Dario @dariogargas.bsky.social · 30/01/2025
Trying to put some order in LLM keywords for an incoming paper. Green concepts are in a different axis, and only partly overlap with elements in blue.
010
Dario @dariogargas.bsky.social · 29/01/2025
AI researchers today, feeling the If poem: "If you can bear to hear the truth you’ve spoken Twisted by knaves to make a trap for fools, Or watch the things you gave your life to, broken, And stoop and build ’em up with worn-out tools"
000
Dario @dariogargas.bsky.social · 27/01/2025
To the editors out there, are there any serious downsides to adding an emoji to the title? The stochastic parrots paper seems to be doing alright...
000
Dario @dariogargas.bsky.social · 24/01/2025
"Personally I would not want to be a member of any group where you either can't wear a hat or you have to wear a hat." George Carlin
000
Dario @dariogargas.bsky.social · 24/01/2025
Aloe 🌱: How I Learned to Stop Worrying and Love LLMs Finishing the journal paper with all the details right now! huggingface.co/collections/...
000
Dario @dariogargas.bsky.social · 23/01/2025
What a day! Met the #FaccT deadline 🥵 Got good reviews from #CVPR 😃 Got a short paper accepted at #NAACL'25🥳 Research excitement all around. Need a rest now.
000
Dario @dariogargas.bsky.social · 22/01/2025
In 2024, we released Aloe Alpha, our 🥇 LLM specialized in healthcare, which gets 10K downloads/month 🥰 6 months later, in a huge team effort, we released four Aloe Beta variants, better than Alpha in every way test. And yet, these get at most 1K downloads/month. Why? huggingface.co/collections/...
huggingface.co
Healthcare LLMs (Aloe family) - a HPAI-BSC Collection
Aloe is a SOTA and open family of Fine-tuned Healthcare LLMs, trained by HPAI at BSC.
000
Dario @dariogargas.bsky.social · 21/01/2025
If I had a time machine, I would travel 15 years back in time, and tell my young self, who is starting a career in AI research: "Time! The key to AI is not how to represent knowledge, but how to represent time-grounded knowledge!"
100
Dario @dariogargas.bsky.social · 21/01/2025
At ISC High Performance 2025, I'll be co-organizing the 5th International Workshop on Computational Aspects of Deep Learning (CADL). See: buff.ly/40qiDS4 Deadline: 28 Feb Topics: -Energy-efficient AI -Large-scale pre-training -Distributed learning approaches -Model optimization strategies
buff.ly
CADL 2025
Advancing AI Through Efficient Computing Over the past decade, Deep Learning (DL) has revolutionized numerous research fields, transforming AI into a computational science where massive models are tra...
110
Dario @dariogargas.bsky.social · 31/12/2024
Answering "More than you have" to the question "How much data is needed for an AI to solve my problem?"
000
Dario @dariogargas.bsky.social · 29/12/2024
Jesus...
000
Reposted by Dario
Iris van Rooij 💭 @irisvanrooij.bsky.social · 16/08/2024
🚨Our paper `Reclaiming AI as a theoretical tool for cognitive science' is now forthcoming in the journal Computational Brain & Behaviour. (Preprint: osf.io/preprints/ps...) Below a thread summary 🧵1/n #metatheory #AGI #AIhype #cogsci #theoreticalpsych #criticalAIliteracy
The idea that human cognition is, or can be understood as, a form of computation is a useful conceptual tool for cognitive science. It was a foundational assumption during the birth of cognitive science as a multidisciplinary field, with Artificial Intelligence (AI) as one of its contributing fields. One conception of Al in this context is as a provider of computational tools (frameworks, concepts, formalisms, models, proofs, simulations, etc.) that support theory building in cognitive science. The contemporary field of Al, however, has taken the theoretical possibility of explaining human cognition as a form of computation to imply the practical feasibility of realising human(-like or -level) cognition in factual computational systems; and, the field frames this realisation as a short-term inevitability. Yet, as we formally prove herein, creating systems with human(-like or -level) cognition is intrinsically computationally intractable.
22491167
Dario @dariogargas.bsky.social · 28/12/2024
All natural intelligence is based on detecting and reacting to change. Meanwhile all our AI systems are static, and AGI sounds like a joke.
000
Dario @dariogargas.bsky.social · 28/12/2024
Always remember, AI does not create, destroy, discover or is in any way responsible for anything. It's all on us. There's always people behind every marvel and awful thing achieved with AI.
000
Dario @dariogargas.bsky.social · 20/12/2024
Looking for fun? Wear a shirt all day that says "I'm recording this conversation and I'll use it to train an AI model"
130
Dario @dariogargas.bsky.social · 18/12/2024
Current AI is at the intersection between statistics and astrology
020
Dario @dariogargas.bsky.social · 13/12/2024
The LLM Aloe family grows 🌱 We release two new versions, based on Qwen 2.5, fine tuned for healthcare. 7B and 72B. Using 2M instructions, including a variety of medical tasks (summarization, explanation, diagnosis, text classification, treatment recommendation...)
buff.ly
HPAI-BSC/Qwen2.5-Aloe-Beta-72B · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
110
Dario @dariogargas.bsky.social · 09/12/2024
The first time you send a team member on an overseas trip alone for the first time, there's a mix of excitement, pride, and a touch of worry. This one was a leap. From BSC@Barcelona to Meta@San Francisco. www.linkedin.com/feed/update/...
linkedin.com
HPAI-BSC on LinkedIn: HPAI at AI at Meta's Global Open Source Innovation Summit! 🌟 We’re…
HPAI at AI at Meta's Global Open Source Innovation Summit! 🌟 We’re excited to share that our colleague Jordi B. represented the group and Barcelona…
000
Dario @dariogargas.bsky.social · 08/12/2024
Indeed, Jurassic Park was released +30 years ago!
020
Dario @dariogargas.bsky.social · 07/12/2024
It used to be, data gathering and preprocessing was the costly first step towards the adoption of AI. Now is benchmarking. If you are going to spend money and effort, do it on building strong tests for AI applications that matter to you
000
Dario @dariogargas.bsky.social · 07/12/2024
AI Daily Dose - Day 11: Loosely speaking, a bias is a pattern in the data that can be identified and used by a model. Without biases, ML models cannot operate, but under certain undesirable biases, ML can produce wrong or dangerous outputs. Only humans can separate between good and bad biases.
120
Dario @dariogargas.bsky.social · 06/12/2024
AI Daily Dose - Day 10: Specifying what an AI model must learnt from a set of data scales poorly when depending on human criteria. In Machine Learning (ML) models are only told how to learn, not what. ML models exploit large volumes of data to find patterns in accordance to their programing.
000
Reposted by Dario
The PS1 startup sound as a lesbian @janus.bsky.social · 06/12/2024
you are not a serious person if your fear of A.I. is “what if scary computer decide 2 extinct humans” and not the health insurance companies using it to sentence people to death
321239205
Dario @dariogargas.bsky.social · 05/12/2024
AI Daily Dose - Day 9: The famous quote "All models are wrong, some are useful" refers to the impossibility of perfectly capturing reality, due to its unlimited complexity and granularity. Instead, models should serve a specific purpose, and model reality to that end.
000
Dario @dariogargas.bsky.social · 05/12/2024
The confusion around LLMs capabilities (clarified below) is increased by the introduction of engineering solutions (e.g., o1) which include one or more LLMs combined through clever tricks, and which are no longer "just an LLM".
000
Dario @dariogargas.bsky.social · 05/12/2024
One of our research engineers is at Meta’s Global Open Source Innovation Summit this week, presenting our family of fine-tuned open healthcare LLMs: Aloe Check out the entire model family here: huggingface.co/collections/... New, even better Aloe models based on Qwen 2.5 coming next week ;)
000