Sign in

joelniklaus.bsky.social

@joelniklaus.bsky.social
34 followers 59 following 253 posts
PostsRepliesMedia
Reposted by @joelniklaus.bsky.social
Dominik Stammbach @dominsta.bsky.social · 28/01/2026
📣 Call for Contributions: LEXam-v2 – A Benchmark for Legal Reasoning in AI How well do today’s AI systems really reason about law? We’re building a global benchmark based on real law school & bar exams. 🧵 Full details, scope, and how to contribute in the thread 👇
164
joelniklaus.bsky.social @joelniklaus.bsky.social · 13/11/2025
You know LLMs have become mainstream when Mark Manson teaches you prompt engineering 😉 youtu.be/AUAHkhOldx8
youtube.com
How to Use ChatGPT to Change Your Life
Prompt PDF here: https://markmanson.net/aipromptsIn this video, I put AI to the test. Not as a productivity hack, but as a personal growth tool. Everyone’s u...
000
joelniklaus.bsky.social @joelniklaus.bsky.social · 12/11/2025
Cool analogy regarding training on the test task
100
joelniklaus.bsky.social @joelniklaus.bsky.social · 11/11/2025
pleias just released 75B tokens of synthetic data upsampled from 50K vital Wikipedia articles! Some thoughts below: - Interesting that they use such a deep architecture for such small models (64 layers for 56M and 80 layers for 321M parameters)
100
joelniklaus.bsky.social @joelniklaus.bsky.social · 10/11/2025
What does AGI actually mean? A who's who in AI spent 57 pages answering that. TLDR: AGI is defined through ten measurable cognitive domains using psychometric theory.
100
joelniklaus.bsky.social @joelniklaus.bsky.social · 08/11/2025
Need copyright-clean training data at scale? Check out the gold mine of the KL3M Data Project on the Hugging Face Hub! ALEA Institute provides 132+ million documents from 16 sources with substantial training resources:
100
joelniklaus.bsky.social @joelniklaus.bsky.social · 07/11/2025
If you're interested in legal retrieval, check out the amazing Massive Legal Embedding Benchmark (MLEB) by Isaacus! Very cool collection of retrieval datasets all available on the Hugging Face hub! Great work by Umar Butler, Abdur-Rahman Butler, Adrian Lucas Malec!
100
joelniklaus.bsky.social @joelniklaus.bsky.social · 06/11/2025
CourtListener supports semantic search now! Apparently they implemented hybrid search using their own fine-tuned ModernBERT model publicly available on the Hugging Face hub! Congrats to @michaeljaylissner and the Free Law Project for making this happen!
110
joelniklaus.bsky.social @joelniklaus.bsky.social · 05/11/2025
On-policy distillation matches RL performance at 2-10% of compute cost. RL gives sparse feedback and burns compute. Off-policy distillation is efficient but learns in the teacher's states, not the student's, causing compounding errors on long sequences.
100
joelniklaus.bsky.social @joelniklaus.bsky.social · 04/11/2025
The correlation between number of reads and edits across Wikipedia articles is 0! This means there is a significant number of articles that are highly read but almost never edited.
100
joelniklaus.bsky.social @joelniklaus.bsky.social · 03/11/2025
If you're exploring computational legal research or building legal AI systems, take a look at Jurisprudence on the Hugging Face Hub by Antoine Jeannot.
100
joelniklaus.bsky.social @joelniklaus.bsky.social · 02/11/2025
Reasoning models excel at math but struggle with simple requests like word limits during thinking TLDR: Models ignore user instructions while reasoning despite following them in final outputs.
100
joelniklaus.bsky.social @joelniklaus.bsky.social · 01/11/2025
Cool long-context eval by Artificial Analysis! AA-LCR is a set of 100 tough questions where you need to piece together answers from several real-world documents—sometimes really big ones—so you can’t just copy and paste the answers.
100
joelniklaus.bsky.social @joelniklaus.bsky.social · 31/10/2025
Very cool work by researchers from Massachusetts Institute of Technology.
100
joelniklaus.bsky.social @joelniklaus.bsky.social · 30/10/2025
Check out the Hugging Face inference endpoints, quick and simple inference for many great open models!
100
joelniklaus.bsky.social @joelniklaus.bsky.social · 29/10/2025
Hugging Face just got promoted to the BigTech club 😉 Thanks to Gian Sbetta and Edouard Treccani for inviting me to a great first AI Builders event in Zurich this evening! Had lots of great conversations with super interesting people!
100
joelniklaus.bsky.social @joelniklaus.bsky.social · 29/10/2025
Seeking students & open-source contributors to join legal AI projects at Hugging Face. If you’re curious about ML × legal tech, this could be a great way to learn + contribute. The legal domain is very rich in hard natural language problems from large scale retrieval to hallucinations.
100
joelniklaus.bsky.social @joelniklaus.bsky.social · 28/10/2025
Impressive collection of specialized models and datasets for French taxation and legal documents by Louis Brulé Naudet:
100
joelniklaus.bsky.social @joelniklaus.bsky.social · 27/10/2025
I just evaluated MiniMax M2 on GPQA-Diamond and LEXam-English in the "I don't know" setup. TLDR: It is very strong on GPQA, especially for its size, but underperforms on LEXam.
100
joelniklaus.bsky.social @joelniklaus.bsky.social · 26/10/2025
Just finished reading "The Ultra-Scale Playbook: Training LLMs on GPU Clusters". Great coverage of the important concepts with good explanations and nice interactive graphics! Thanks Nouamane Tazi, Ferdinand Mom, Haojun Zhao, Phuc Nguyen, Mohamed Mekkouri, Leandro Werra, Thomas Wolf!
100
joelniklaus.bsky.social @joelniklaus.bsky.social · 23/10/2025
LEXam Update: GPT-5 Takes the Top Spot We're excited to share our latest LEXam evaluation results: - GPT-5 claims the #1 position, outperforming Gemini 2.5 Pro and setting a new state-of-the-art for legal reasoning on LEXam!
100
joelniklaus.bsky.social @joelniklaus.bsky.social · 22/10/2025
GPT-5 and Claude can ace GPQA Diamond, but LEXam (a legal reasoning benchmark) exposes a critical flaw: they'd rather be confidently wrong than admit uncertainty. ⚙️ The Setup I evaluated ten frontier models on LEXam (English MC subset) using an "I don't know" (IDK) protocol.
100
joelniklaus.bsky.social @joelniklaus.bsky.social · 21/10/2025
Stop what you are doing and try out GEPA now! "GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning" presents such elegant ideas by a collection of amazing researchers! Here is a tldr of how it works:
143
joelniklaus.bsky.social @joelniklaus.bsky.social · 20/10/2025
What's special about October 27th 2025? Yoshua Bengio, the most-cited computer scientist in the world, is 1 week away from becoming the first ML researcher to hit 1 million citations! 🤯 At his current rate of 366 citations/day, he'll reach this unprecedented milestone around October 27th 🎯
100
joelniklaus.bsky.social @joelniklaus.bsky.social · 19/10/2025
Very cool and detailed Stanford University and Carnegie Mellon University study on sycophancy: "Sycophantic AI Decreases Prosocial Intentions and Promotes Dependence" Sycophancy, the phenomenon of excessively agreeing with or flattering users, is a pervasive issue in current LLMs. Findings:
144
joelniklaus.bsky.social @joelniklaus.bsky.social · 16/10/2025
Super cool new LLM system by Alex Zhang and Omar Khattab!
100
joelniklaus.bsky.social @joelniklaus.bsky.social · 15/10/2025
Very excited to announce that I am officially the second most cited researcher working on #good_shit according to Google Scholar trailing behind Lucas Beyer by only 99,062 citations 🎉😂
000
joelniklaus.bsky.social @joelniklaus.bsky.social · 14/10/2025
Cool tech report by Chroma: "Context Rot: How Increasing Input Tokens Impacts LLM Performance" Main Findings: - LLM performance drops as input length grows, even in simple tasks. - Semantic ambiguity and distractors accelerate this decline.
100
joelniklaus.bsky.social @joelniklaus.bsky.social · 08/10/2025
Visiting the Hugging Face HQ in Paris last week was a pleasure. Many thanks to the amazing LeRobot team for showing me around and letting me play. 😉 Check out their awesome work on the hub!
000
joelniklaus.bsky.social @joelniklaus.bsky.social · 07/10/2025
Just read this nice blog post "The Second Half".
100
joelniklaus.bsky.social @joelniklaus.bsky.social · 05/10/2025
I am looking forward to presenting our work "LEXam: Benchmarking Legal Reasoning on 340 Law Exams at the Data Science Talk Series @ UNT! Thanks Haihua Chen for the invitation! Abstract:
100
joelniklaus.bsky.social @joelniklaus.bsky.social · 02/10/2025
Cool to see context windows continuously being pushed: xAI's Grok-4-fast supports 2M tokens now!
100
joelniklaus.bsky.social @joelniklaus.bsky.social · 01/10/2025
Amazing structured analysis on LoRA by Thinking Machines Lab! This was really missing, I was thinking about doing something similar myself. Here are the findings that surprised me most: - LoRA rank 1 is enough for RL
100
joelniklaus.bsky.social @joelniklaus.bsky.social · 30/09/2025
Finally seeing some movement on the kind of evals that actually matter 🎯 GDPval is a benchmark that measures how well AI models perform on economically valuable, real-world tasks across major knowledge work occupations by comparing their outputs to expert human deliverables.
100
joelniklaus.bsky.social @joelniklaus.bsky.social · 29/09/2025
Very detailed explanation of vLLM by Aleksa! Blog post: www.aleksagordic.com/blog/vllm
aleksagordic.com
Inside vLLM: Anatomy of a High-Throughput LLM Inference System - Aleksa Gordić
From paged attention, continuous batching, prefix caching, specdec, etc. to multi-GPU, multi-node dynamic serving at scale.
000
joelniklaus.bsky.social @joelniklaus.bsky.social · 28/09/2025
Following up on OpenAI's recent paper "Why Language Models Hallucinate," I wondered how their proposed scoring changes would affect actual benchmark performance. So I ran the experiment myself 🧪
100
joelniklaus.bsky.social @joelniklaus.bsky.social · 24/09/2025
Getting 10x speedup in #vLLM is easier than you think 📈 I just discovered speculative decoding with ngram lookup and the results speak for themselves. Here's what you add to your vLLM serve command: speculative_config={ "method": "ngram", "num_speculative_tokens": 8,
100
joelniklaus.bsky.social @joelniklaus.bsky.social · 23/09/2025
Major model providers consistently overlook legal capabilities in their evaluations. Yet, law is a formidable challenge for LLM evaluation: high-stakes accuracy requirements, complex multi-step reasoning, and almost entirely text-based workflows. Below a few key evals worth tracking:
100
joelniklaus.bsky.social @joelniklaus.bsky.social · 22/09/2025
Constitutional AI gets a Swiss twist 🇨🇭 Found an interesting nugget in the commendably detailed #Apertus paper. Rather than relying on generic preference datasets, they developed principles based on Swiss constitutional values - neutrality, federalism, direct democracy, privacy protection.
120
joelniklaus.bsky.social @joelniklaus.bsky.social · 20/09/2025
I just discovered the Hugging Face playground, and it's incredibly useful for quick debugging! You can explore a variety of open-source models rapidly. It even supports markdown parsing and structured outputs.
100
joelniklaus.bsky.social @joelniklaus.bsky.social · 19/09/2025
Not thrilled to announce that our paper has not been accepted to NeurIPS! Sharing our story below for those interested:
110
joelniklaus.bsky.social @joelniklaus.bsky.social · 18/09/2025
Switzerland is #10 in working-age population adjusted Anthropic Claude usage worldwide. Usage in Israel is almost 3x compared to Switzerland. I wonder if people in Switzerland just use chatbots less overall or if they just don't use Claude that much.
100
joelniklaus.bsky.social @joelniklaus.bsky.social · 17/09/2025
Last week I ran some inference speed benchmarking with vLLM. Here are three tidbits I found interesting: 1. The model you select can make or break your speed! Qwen3 4B gets a throughput of 2570 output tokens/s whereas Gemma3 4B achieves 4670. That's an almost 2x speedup!
100
joelniklaus.bsky.social @joelniklaus.bsky.social · 14/09/2025
Just read OpenAI's paper "Why Language Models Hallucinate" and there's crucial insights buried in the technical details that affects everyone using these systems.
100
joelniklaus.bsky.social @joelniklaus.bsky.social · 12/09/2025
The new Qwen3-Next model is genuinely impressive. What stands out most is their focus on improving inference speed - they're achieving over 10x throughput improvements for contexts beyond 32K tokens 🚀
110
Reposted by @joelniklaus.bsky.social
Negar Foroutan @negarforoutan.bsky.social · 11/08/2025
In short, Parity-aware BPE=minimal overhead+clear fairness gains. If you care about multilingual robustness, tokenization is low-hanging fruit. Joint work with Clara Meister, @debjit-paul.bsky.social @joelniklaus.bsky.social @sinaahmadi.bsky.social @abosselut.bsky.social @ricosennrich.bsky.social
131
Reposted by @joelniklaus.bsky.social
Guilherme Penedo @guilherme.hf.co · 08/12/2024
Announcing 🥂 FineWeb2: A sparkling update with 1000s of 🗣️languages. We applied the same data-driven approach that led to SOTA English performance in🍷 FineWeb to thousands of languages. 🥂 FineWeb2 has 8TB of compressed text data and outperforms other datasets.
17619