Sign in

Yu Fan

@yu-fan-768.bsky.social
14 followers 31 following 0 posts
PostsRepliesMedia
Reposted by Yu Fan
AI4Law@ICML @ai4law.bsky.social · 06/04/2026
🚨 Call for Papers: AI for Law Workshop @icmlconf.bsky.social (July 10, Seoul 🇰🇷), welcomes submissions across three themes: ⚖️ AI for Legal Reasoning 📊 AI Evaluation for Law 🌍 AI for Access to Justice 📄 Full papers (≤8 pages) ⏰ Deadline: May 22 (AoE) 📄Submission via: openreview.net/group?id=ICM...
164
Reposted by Yu Fan
Dominik Stammbach @dominsta.bsky.social · 28/01/2026
📣 Call for Contributions: LEXam-v2 – A Benchmark for Legal Reasoning in AI How well do today’s AI systems really reason about law? We’re building a global benchmark based on real law school & bar exams. 🧵 Full details, scope, and how to contribute in the thread 👇
164
Reposted by Yu Fan
Alexander Hoyle @alexanderhoyle.bsky.social · 24/09/2025
Accepted to EMNLP (and more to come 👀)! The camera ready version is now online---very happy with how this turned out arxiv.org/abs/2507.01234
0145
Reposted by Yu Fan
Alexander Hoyle @alexanderhoyle.bsky.social · 17/07/2025
New preprint! Have you ever tried to cluster text embeddings from different sources, but the clusters just reproduce the sources? Or attempted to retrieve similar documents across multiple languages, and even multilingual embeddings return items in the same language? Turns out there's an easy fix🧵
Barchart of number of items in four clusters of text embeddings, with colors showing the distribution of sources in each cluster.

Caption: Clustering text embeddings from disparate sources (here, U.S. congressional bill summaries and senators’ tweets) can produce clusters where one source dominates (Panel A). Using linear erasure to remove the source information produces more evenly balanced clusters that maintain semantic coherence (Panel B; sampled items relate to immigration). Four random clusters of k-means shown (k=25), trained on a combined 5,000 samples from each dataset
2327