Sign in

Cui Ding

@cuiding.bsky.social
44 followers 82 following 19 posts
PostsRepliesMedia
Cui Ding @cuiding.bsky.social · 02/09/2026
Looking forward to reading it!
000
Cui Ding @cuiding.bsky.social · 29/08/2026
昨天是emotional的一天。中午Reinhold问我,什么时候会完成dessertation,以后有什么计划。我忙着说对不起,说我总是想证明自己,证明自己其实比profile中展示出来的更好,因为博士前两年…… Reinhold说:don't worry. People have already formed opinions. 一直以来都觉得要做到“人不知而不愠”好难好难。那刻忽然像有了答案,很释怀了。
000
Cui Ding @cuiding.bsky.social · 24/08/2026
记录一个尴尬但不失可爱的talk经历。开头和我导师Lena的插科打诨被听众以为是彩排好了的相声导入。
100
Reposted by Cui Ding
Computational Linguistics @ UZH @cl-uzh.bsky.social · 08/07/2026
What a blast! 🎉 We won one 🏆Test of Time Award and two 🏆🏆SAC Highlight Awards at #ACL2026! Congratulations to our award winners: @ricosennrich.bsky.social, @michellewastl.bsky.social, @vamvas.bsky.social, @sinaahmadi.bsky.social and their (co-)authors! Links to the winning papers in the comments.
Our award winners in front of the ACL backdrop
3104
Reposted by Cui Ding
Vilém Zouhar @zouhar.bsky.social · 06/07/2026
All my homies left for ACL/ICML and left me home managing a project *for which we're looking for coauthors* 🥸. Join us (also talk to @sethjsa.bsky.social @onadegibert.bsky.social @niyatibafna.bsky.social @patuchen.bsky.social @michellewastl.bsky.social @ayukh.bsky.social)
2104
Reposted by Cui Ding
michellewastl.bsky.social @michellewastl.bsky.social · 03/07/2026
Excited to be at #ACL2026NLP in San Diego with two papers! Come by to learn about cross-lingual semantic differences from work with @vamvas.bsky.social & @ricosennrich.bsky.social, and how LLMs respond to expressions of belief, with Kevin Du, @clarakuempel.bsky.social, and @alexwarstadt.bsky.social
093
Reposted by Cui Ding
Jan Chromý @janchromy.bsky.social · 11/05/2026
New paper out in Cognitive Science! We analyzed various predictors of local ambiguity resolution on a heterogeneous set of garden-path sentences, namely surprisal, frequency, plausibility and cloze scores (measuring the likelihood of misanalysis of the ambiguous region). 1/2
onlinelibrary.wiley.com
Cloze, Frequency, Surprisal, or Plausibility? A Comparative Analysis of Predictors for Local Ambiguity Resolution
This study investigated the cognitive mechanisms underlying the processing of garden-path sentences by examining the influence of verb/structural bias, cloze probability, surprisal, and plausibility....
131
Reposted by Cui Ding
Computational Linguistics @ UZH @cl-uzh.bsky.social · 17/04/2026
🗓️ SwissText 2026 keynote speakers announced & registration open! We are delighted to welcome Prof. Dr. Alexandra Birch and Dr. Valentina Pyatkin as our keynote speakers. 📋 Register here: ema.uzh.ch/RHK4W Early-bird rates available throughout April, with additional student discounts. #NLProc 1/3
ema.uzh.ch
SwissText 2026
10. Juni 2026 | UZH Campus Oerlikon
153
Reposted by Cui Ding
Rico Sennrich @ricosennrich.bsky.social · 23/03/2026
My lab is recruiting one PhD student and one post-doctoral researcher for a start as soon as this Summer! Apply by April 1 / March 31 to be among first candidates considered. jobs.uzh.ch/job-vacancie... jobs.uzh.ch/job-vacancie...
jobs.uzh.ch
UZH: PhD in Language AI / Natural Language Processing
You will be joining the Department of Computational Linguistics, which has 6 Research Groups and around 70 postdoctoral and student researchers in the areas of Text Technologies, Phonetics and Speech ...
046
Cui Ding @cuiding.bsky.social · 20/12/2025
So happy to have presented our project on **Individual signatures of spillover** at CPL2025. Grateful to my supervisor Lena Jaeger, for her great ideas and for her kind support! Thanks to my coauthors and the audience at CPL2025. The award was a democratic vote by the audience!
020
Reposted by Cui Ding
Tamar Regev @tamaregev.bsky.social · 15/12/2025
New preprint on prosody in the brain! tinyurl.com/2ndswjwu HeeSoKim NiharikaJhingan SaraSwords @hopekean.bsky.social @coltoncasto.bsky.social JenniferCole @evfedorenko.bsky.social Prosody areas are distinct from pitch, speech, and multiple-demand areas, and partly overlap with lang+social areas→🧵
tinyurl.com
A distinct set of brain areas process prosody--the melody of speech
Human speech carries information beyond the words themselves: pitch, loudness, duration, and pauses--jointly referred to as 'prosody'--emphasize critical words, help group words into phrases, and conv...
13313
Reposted by Cui Ding
michellewastl.bsky.social @michellewastl.bsky.social · 09/12/2025
Take a look at how we challenge state-of-the-art NLP systems to recognize token-level semantic differences across languages with our new SwissGov-RSD dataset! @vamvas.bsky.social @ricosennrich.bsky.social Paper: arxiv.org/pdf/2512.075... Dataset: huggingface.co/datasets/Zur... #NLProc
063
Reposted by Cui Ding
Yevgeni Berzak @whylikethis.bsky.social · 02/12/2025
#NeurIPS2025 Check out EyeBench 👀, a mega-project which provides a much needed infrastructure for loading & preprocessing eye-tracking for reading datasets, and addressing super exciting modeling challenges: decoding linguistic knowledge 👩 and reading interactions 👩+📖 from gaze! eyebench.github.io
092
Cui Ding @cuiding.bsky.social · 01/12/2025
Actually, we are also trying to working on a solution, but for Chinese, mainly.
010
Cui Ding @cuiding.bsky.social · 26/11/2025
Wow! Thank you so much!!
130
Reposted by Cui Ding
Computational Linguistics @ UZH @cl-uzh.bsky.social · 06/11/2025
UZH group picture at #EMNLP2025! If you're here, catch us for a chat!
041
Cui Ding @cuiding.bsky.social · 02/11/2025
Take-home Message 🔹 We formalize input quality in reading as mutual information. 🔹 We link it to measurable human behavior. 🔹 We show multimodal LLMs can model this effect quantitatively. Bottom-up information matters — and now we can measure how much it matters.
000
Cui Ding @cuiding.bsky.social · 02/11/2025
Key Result 2: Information from Models Using fine-tuned Qwen2.5-VL and TransOCR, we estimated the MI between images and word identity. MI systematically drops: Full > Upper > Lower — perfectly mirroring human reading patterns! 🤯
100
Cui Ding @cuiding.bsky.social · 02/11/2025
Key Result 1: Human Reading Reading times show a clear pattern: Full visible< Upper visible < Lower visible in both English & Chinese. 👉 Upper halves are more informative (and easier to read).
100
Cui Ding @cuiding.bsky.social · 02/11/2025
We model reading time as proportional to the number of visual “samples” needed to reduce uncertainty below a threshold ϕ. Higher mutual information → fewer samples → faster reading.
100
Cui Ding @cuiding.bsky.social · 02/11/2025
📊 We quantify this using mutual information (MI) between visual input and word identity. To test the theory, we created a reading experiment using the MoTR (Mouse-Tracking-for-Reading) paradigm 🖱️📖 We ran the study in both English and Chinese.
120
Cui Ding @cuiding.bsky.social · 02/11/2025
We propose a formal model where reading is a Bayesian update integrating top-down expectations and bottom-up evidence. When bottom-up input is noisy (e.g., words are partially occluded), comprehension becomes harder and slower.
110
Cui Ding @cuiding.bsky.social · 02/11/2025
👀Ever wondered how visual information quality affects reading and language processing? Our new #EMNLP2025 paper with @wegotlieb.bsky.social, Lena Jäger -- “Modeling Bottom-up Information Quality during Language Processing”, bridges psycholinguistics and multimodal LLMs. 🧠💡👇 arxiv.org/pdf/2509.17047
140
Reposted by Cui Ding
Kirill Semenov @kiryukhasemenov.bsky.social · 28/10/2025
Let's meet at #EMNLP and talk about multilingual knowledge benchmarks! ⚠️MLAMA is full of disfluent sentences ❓Reason: templated translation 💡Simple full-sentence translation improves factual retrieval up to 25% 🙌Remember to check your benchmarks with speakers! Link: arxiv.org/pdf/2510.15115
011
Reposted by Cui Ding
Hanxu Hu @hanxuhu.bsky.social · 21/10/2025
💥Introducing new paper: arxiv.org/pdf/2510.17715, QueST — train specialized generators to create challenging coding problems. From Qwen3-8B-Base ✅ 100K synthetic problems: better than Qwen3-8B ✅ Combining with human written problems: matches DeepSeek-R1-671B 🧵(1/5)
143
Reposted by Cui Ding
Henrik Singmann @singmann.bsky.social · 02/09/2025
Exciting #rstats news for Bayesian model comparison: bridgesampling is finally ready to support cmdstanr, see screenshot. Help us by installing the development version of bridgesampling and letting us know if it works for your model(s): pak::pkg_install("quentingronau/bridgesampling#44")
R code and output showing the new functionality:
``` r
## pak::pkg_install("quentingronau/bridgesampling#44")
## see: https://cran.r-project.org/web/packages/bridgesampling/vignettes/bridgesampling_example_stan.html
library(bridgesampling)

### generate data ###
set.seed(12345)
mu <- 0
tau2 <- 0.5
sigma2 <- 1
n <- 20
theta <- rnorm(n, mu, sqrt(tau2))
y <- rnorm(n, theta, sqrt(sigma2))

### set prior parameters ###
mu0 <- 0
tau20 <- 1
alpha <- 1
beta <- 1

stancodeH0 <- 'data {
  int<lower=1> n; // number of observations
  vector[n] y; // observations
  real<lower=0> alpha;
  real<lower=0> beta;
  real<lower=0> sigma2;
}
parameters {
  real<lower=0> tau2; // group-level variance
  vector[n] theta; // participant effects
}
model {
  target += inv_gamma_lpdf(tau2 | alpha, beta);
  target += normal_lpdf(theta | 0, sqrt(tau2));
  target += normal_lpdf(y | theta, sqrt(sigma2));
}
'
tf <- withr::local_tempfile(fileext = ".stan")
writeLines(stancodeH0, tf)
mod <- cmdstanr::cmdstan_model(tf, quiet = TRUE, force_recompile = TRUE)

fitH0 <- mod$sample(
  data = list(y = y, n = n,
              alpha = alpha,
              beta = beta,
              sigma2 = sigma2),
  seed = 202,
  chains = 4,
  parallel_chains = 4,
  iter_warmup = 1000,
  iter_sampling = 50000,
  refresh = 0
)
#> Running MCMC with 4 parallel chains...
#> 
#> Chain 3 finished in 0.8 seconds.
#> Chain 2 finished in 0.8 seconds.
#> Chain 4 finished in 0.8 seconds.
#> Chain 1 finished in 1.1 seconds.
#> 
#> All 4 chains finished successfully.
#> Mean chain execution time: 0.9 seconds.
#> Total execution time: 1.2 seconds.
H0.bridge <- bridge_sampler(fitH0, silent = TRUE)
print(H0.bridge)
#> Bridge sampling estimate of the log marginal likelihood: -37.73301
#> Estimate obtained in 8 iteration(s) via method "normal".

#### Expected output:
## Bridge sampling estimate of the log marginal likelihood: -37.53183
## Estimate obtained in 5 iteration(s) via method "normal".
```
2289
Reposted by Cui Ding
Shravan Vasishth @shravanvasishth.bsky.social · 31/08/2025
We are done with the ninth Statistical Methods for Linguistics and Psychology (SMLP) summer school, Potsdam, Germany. The tenth edition is planned for 24-28 August 2026.
0163
Reposted by Cui Ding
Tiago Pimentel @tpimentel.bsky.social · 31/07/2025
Honoured to receive two (!!) SAC highlights awards at #ACL2025 😁 (Conveniently placed on the same slide!) With the amazing: @philipwitti.bsky.social, @gregorbachmann.bsky.social and @wegotlieb.bsky.social, @cuiding.bsky.social, Giovanni Acampa, @alexwarstadt.bsky.social, @tamaregev.bsky.social
0223
Reposted by Cui Ding
Computational Linguistics @ UZH @cl-uzh.bsky.social · 30/07/2025
Congratulations to @sinaahmadi.bsky.social and co-authors for receiving an ACL 2025 Outstanding Paper Award for PARME: Parallel Corpora for Low-Resourced Middle Eastern Languages! aclanthology.org/2025.acl-lon...
Sina Ahmadi receiving award.
0156
Reposted by Cui Ding
Shravan Vasishth @shravanvasishth.bsky.social · 10/07/2025
Next week onwards, I'm teaching a five-day introductory course on Bayesian Data Analysis in Gent. Newly recorded video lectures to accompany the course are now online: vasishth.github.io/LecturesIntr...
vasishth.github.io
Shravan Vasishth's Intro Bayes course home page
0135
Reposted by Cui Ding
Kirill Semenov @kiryukhasemenov.bsky.social · 06/06/2025
📣Take part in 3rd Terminology shared task @WMT!📣 This year: 👉5 language pairs: EN->{ES, RU, DE, ZH}, 👉2 tracks - sentence-level and doc-level translation, 👉authentic data from 2 domains: finance and IT! www2.statmt.org/wmt25/termin... Don't miss an opportunity - we only do it once in two years😏
www2.statmt.org
Terminology Translation Task
032
Cui Ding @cuiding.bsky.social · 04/06/2025
Some of my colleagues are already very excited about this work!
020
Reposted by Cui Ding
Tiago Pimentel @tpimentel.bsky.social · 29/05/2025
If you're finishing your camera-ready for ACL or ICML and want to cite co-first authors more fairly, I just made a simple fix to do this! Just add $^*$ to the authors' names in your bibtex, and the citations should change :) github.com/tpimentelms/...
Inline citations with only first author name, or first two co-first author names.
48322
Reposted by Cui Ding
Yevgeni Berzak @whylikethis.bsky.social · 29/05/2025
👀 📖 Big news! 📖 👀 Happy to announce the release of the OneStop Eye Movements dataset! 🎉 🎉 OneStop is the product of over 6 years of experimental design, data collection and data curation. github.com/lacclab/OneS...
193
Cui Ding @cuiding.bsky.social · 14/05/2025
I am so proud of this work. My first NLP experience. I learned a lot from this amazing team!!!!
030
Reposted by Cui Ding
Ethan Gotlieb Wilcox @wegotlieb.bsky.social · 13/05/2025
⭐🗣️New preprint out: 🗣️⭐ “Using Information Theory to Characterize Prosodic Typology: The Case of Tone, Pitch-Accent and Stress-Accent” with @cuiding.bsky.social , Giovanni Acampa, @tpimentel.bsky.social , @alexwarstadt.bsky.social ,Tamar Regev: arxiv.org/abs/2505.07659
arxiv.org
Using Information Theory to Characterize Prosodic Typology: The Case of Tone, Pitch-Accent and Stress-Accent
This paper argues that the relationship between lexical identity and prosody -- one well-studied parameter of linguistic variation -- can be characterized using information theory. We predict that lan...
1125
Cui Ding @cuiding.bsky.social · 07/03/2025
The biggest advantage of MoTR over alternative methods is that it is very cheap and fast compared to its alternatives, while still provides very sensitive and accurate measurements. Our online data collection from 60 Russian speakers took less than 24 hours!!
010
Cui Ding @cuiding.bsky.social · 07/03/2025
Participants must move their mouse over the text to reveal the words, while their cursor movements are recorded (similar to how eye movements are recorded in eye tracking). See below for an example MoTR trial.
100
Cui Ding @cuiding.bsky.social · 07/03/2025
2- We use MoTR (Mouse Tracking for Reading) as a cheaper but reliable alternative to in-person eye tracking. MoTR is a new experimental tool, where participants screen is blurred except for a small region around the tip of the mouse pointer
120
Cui Ding @cuiding.bsky.social · 07/03/2025
Excited to share our preprint "Using MoTR to probe agreement errors in Russian"! w/ Metehan Oğuz, @wegotlieb.bsky.social, Zuzanna Fuchs Link: osf.io/preprints/ps... 1- We provide moderate evidence that processing of agreement errors is modulated by agreement type (internal vs external agr.)
osf.io
OSF
131