Sign in

Xiulin Yang

@xiulinyang.bsky.social
37 followers 45 following 7 posts

PhD student at Georgetown University | Computational linguist | previously at Saarland, Groningen, Oxford, & SDU

PostsRepliesMedia
Reposted by Xiulin Yang
Leonie Weissweiler @weissweiler.bsky.social · 01/10/2026
📢I'm hiring a PhD student to work on interpretability for Protein Language Models! 🗓️There are now two weeks left to apply, find out more on the LION Lab website: lionlabnlp.github.io/jobs/ai4pf/
lionlabnlp.github.io
PhD Student Position in Interpretability for Protein Language Models — LION Lab
Fully funded PhD student position at LION Lab, Leipzig University, on interpretability for protein language models.
186
Xiulin Yang @xiulinyang.bsky.social · 30/09/2026
It’s all about tokens! 👀
051
Reposted by Xiulin Yang
Tom McCoy @rtommccoy.bsky.social · 01/09/2026
🤖🧠NEW PAPER🧠🤖 (The result of an 8-year project!) LLMs seem very different from symbolic systems. Yet LLMs excel in symbolic domains (e.g., language/code/math). How do they do it? Our finding: LLM representations have implicit symbolic structure! Link in thread ⬇️ 1/n
Overview of the paper. 
Title: The Emergent Symbolic Structure of Artificial Neural Networks
Authors: Tom McCoy, Paul Soulos, Tal Linzen, Paul Smolensky
Left: Neural networks encode information in vectors (there is then an image of a vector), yet they excel at tasks long thought to require symbolic structure (there is then an image of a symbolic representation, specifically a syntax tree). How do LLMs do it?
Right: We find that LLM representations can be closely approximated with symbolic structures. This approximation lets us edit the structure of an LLM’s output by editing the structure of its internal representations, as shown. There is then an image of two edits to LLMs. In the first one, the original input is 3 + 6 * 8, with an answer of 51. But if we swap the positions of the 3 and the 6, the output becomes 30. In the second one, the original input is a Python command repeating the list [Z, U] three times, producing [Z, U, Z, U, Z, U]. But if we edit the input in a way that adds a Q at the end of the input, the output becomes [Z, U, Q, Z, U, Q, Z, U, Q].
431888
Reposted by Xiulin Yang
Victoria Bosch @initself.bsky.social · 28/08/2026
Are brains and artificial neural networks converging onto universal representations? There is a seductive idea making the rounds in NeuroAI / machine learning: train systems well enough, and they all converge on the same representation of reality (i.e. a unique world model). We have thoughts™ 1/n
cell.com
The Umwelt Representation Hypothesis: rethinking Universality
Recent studies reveal striking representational alignment between artificial neural networks (ANNs) and biological brains, leading to proposals that all sufficiently capable systems converge on univer...
818176
Xiulin Yang @xiulinyang.bsky.social · 28/08/2026
🍎🍊 How would you know if a language model is better at one language than another? Our #EMNLP2026 paper argues that only one metric can actually lead to fair crosslingual evaluation. This work is a collaboration with @wegotlieb.bsky.social & @catherinearnett.bsky.social! (1/5)
13212
Reposted by Xiulin Yang
Francesca Padovani @frap98.bsky.social · 20/05/2026
You can access the models through the repository: github.com/fpadovani/CA... Thank you to the amazing team: @xiulinyang.bsky.social @bbunzeck.bsky.social @arianna-bis.bsky.social @yevgenm.bsky.social @jumelet.bsky.social
github.com
GitHub - fpadovani/CAIT-Toolkit: Syntactic Annotation Toolkit for Child–Adult InTeractions (CAIT)
Syntactic Annotation Toolkit for Child–Adult InTeractions (CAIT) - fpadovani/CAIT-Toolkit
042