Sign in

Freda Shi

@fredashi.bsky.social
396 followers 185 following 84 posts

Assistant Professor & Canada CIFAR AI Chair, University of Waterloo & Vector Institute | Excited about "grounding" in any form | Feeder of 2 🐈 | 🏸, 🏐, 🏂 | she/her

PostsRepliesMedia
Reposted by Freda Shi
Computational Linguistics Journal @complingjournal.bsky.social · 01/05/2026
Hallucinations pose a substantial challenge to the reliability of LLMs in real-world scenarios. Zhang et al. survey methods of detection, explanation, & mitigation of hallucination, & provide a taxonomy & list of benchmarks for evaluation in this paper: doi.org/10.1162/COLI... @fredashi.bsky.social
031
Freda Shi @fredashi.bsky.social · 01/04/2026
My student and her undergrad advisor submitted their very elegant work on X to "Transactions of X" and got desk-rejected, with the reason: "Your paper is very interesting and a very hardcore contribution to X, but we can't find appropriate reviewers, so we have to desk-reject it." Excuse me???
340
Freda Shi @fredashi.bsky.social · 28/03/2026
Wow, learned about an even more impressive statement today: We are Sorry to inform you that you Didn’t Get the Honor for Voluntary Contribution because the speakers you invited (including yours truly) are Not Famous Enough.
030
Freda Shi @fredashi.bsky.social · 24/03/2026
We are Sorry to inform you that you Didn't Get the Honor for Voluntary Contribution because you onboarded Junior People.
280
Reposted by Freda Shi
Naomi Saphra @nsaphra.bsky.social · 24/03/2026
We are Delighted to Congratulate you on the Honor of the Opportunity of Volunteering
2374
Freda Shi @fredashi.bsky.social · 24/03/2026
Recent takes from this effort: Humans are largely linear animals, and we are so used to all kinds of linear structures, including doing tasks one by one.
100
Reposted by Freda Shi
Abdellah Fourtassi @fourtassi.bsky.social · 08/03/2026
Open PhD/Postdoc position (start: Oct 2026). Topic: AI/LLMs and child language/communicative/cognitive development. The exact project will be shaped with the candidate. Join our team @univ-amu.fr at the intersection of computer and cognitive science (& right next to the Calanques!). Send me your CV!
045
Freda Shi @fredashi.bsky.social · 08/03/2026
I truly feel this AI-assisted development is something completely new. The key difference is the design philosophy: I also like Notion, for example, but I had to adapt myself to their design. Now the assistant bot and myself are doing "bidirectional alignment"---AI coding agents enable this.
170
Freda Shi @fredashi.bsky.social · 08/03/2026
I've been enjoying developing a (safe & fairly reliable) personal assistant bot, partly motivated by some light experience trying OpenClaw. For safety, everything is run & saved locally. I can now manage my todo list, notebook, and calendar simply by speaking to my bot. (1/)
110
Freda Shi @fredashi.bsky.social · 27/01/2026
🐻
120
Freda Shi @fredashi.bsky.social · 24/01/2026
I frankly don't think the ICML policy of having authors rank their papers will work. I have 3-5 submissions (depending on ICLR results) this time to quite different subcommunities. I know my favorite submissions could be quite controversial---people either like or dislike them a lot. (1/2)
110
Reposted by Freda Shi
Abdellah Fourtassi @fourtassi.bsky.social · 20/01/2026
Thrilled to announce the 1st Workshop on Computational Developmental Linguistics (CDL) at ACL 2026 🎉 A new venue at the intersection of development linguistics × modern NLP, spearheaded by @fredashi.bsky.social @marstin.bsky.social, and and outstanding team of colleagues! A thread 🧵
3219
Freda Shi @fredashi.bsky.social · 13/01/2026
AI coding is now the best at implementing things that are (1) either already standardized or not very ambiguous in their natural-language description, and (2) with many details. Yes, I just vibe-coded a helper for some organizational matters, and will do more. 1/
130
Reposted by Freda Shi
Michael Saxon @saxon.me · 30/10/2025
It's #NSF #GRFP application season again so it's time to re-up my GRFP application advice post! Also, check out the cool bsky comment integration I've added to the blog! Engagement with this post will go under the blogpost on my site as comments! saxon.me/blog/2024/gr...
saxon.me
NSF GRFP Application Tips for NLP, AI, CS
Reflections and advice from my successful NSF GRFP proposal in NLP. Why I think my applications worked well, what I wish I did differently, and links to my actual statements and feedback from the GRFP...
183
Freda Shi @fredashi.bsky.social · 31/10/2025
Reviewer timeliness by area (based on a small sample of 14 papers * 4 reviewers each): Multilingual LMs >> VLMs > LLM Reasoning.
130
Reposted by Freda Shi
Artjoms Šeļa @artjomshl.bsky.social · 23/10/2025
We have updated our collection of multilingual poetry corpora PoeTree. With addition of Norwegian, and few metadata fixes it now has a loud label of 1.0.0 release! 🌳 versologie.cz/poetree/vers...
versologie.cz
PoeTree. Poetry corpora in 11 languages
PoeTree is a standardized collection of poetry corpora comprising nearly 335,000 poems in ten languages (Czech, English, French, German, Hungarian, Italian, Norwegian, Portuguese, Russian, Slovenian, ...
34413
Freda Shi @fredashi.bsky.social · 21/10/2025
Should be fun to read!
080
Freda Shi @fredashi.bsky.social · 20/10/2025
I can only go through <40 slides in a one-hour talk... verified multiple times. My job talk has 39 slides (including acknowledgement), and so are my recent talks. Always impressed when folks present 100 slides in the same amount of time.
120
Reposted by Freda Shi
Melanie Mitchell @melaniemitchell.bsky.social · 01/10/2025
In 2016 Hinton predicted that AI would replace all radiologists in five years. Ten years later, why hasn't it happened? This post is a great explainer. www.understandingai.org/p/ai-isnt-re...
understandingai.org
AI isn't replacing radiologists
Radiology combines digital images, clear benchmarks, and repeatable tasks. But demand for human radiologists is at an all-time high.
11348153
Freda Shi @fredashi.bsky.social · 29/09/2025
🚀 ACL ARR is looking for a Co-CTO to join me lead our amazing tech team and drive the future of our workflow. If you’re interested or know someone who might be, let’s connect! RTs & recommendations appreciated.
143
Freda Shi @fredashi.bsky.social · 29/09/2025
Same! I‘ve literally unbidden 0 out of my recommended batch.
030
Freda Shi @fredashi.bsky.social · 05/08/2025
Is there a specific reason that NeurIPS does not show AC identity to reviewers? I'm very curious about who sent the polite reminders, and of course, even more curious about the rude ones.
140
Reposted by Freda Shi
Hokin @hokin.bsky.social · 30/06/2025
#CoreCognition #LLM #multimodal #GrowAI We spent 3 years to curate 1503 classic experiments spanning 12 core concepts in human cognitive development and evaluated on 230 MLLMs with 11 different prompts for 5 times to get over 3.8 millions inference data points. A thread (1/n) - #ICML2025 ✅
1139
Freda Shi @fredashi.bsky.social · 05/05/2025
Just in case this personal latex environmental setup is helpful to a broader crowd ⬇️
040
Reposted by Freda Shi
Ana Marasović @anamarasovic.bsky.social · 03/05/2025
I'm late to the NAACL party, but I've just arrived in Albuquerque! I'll talk about measuring faithfulness of verbalized reasoning on *Sunday* at Repl4NLP at *9:45*.
1333
Reposted by Freda Shi
Najoung Kim @najoung.bsky.social · 04/05/2025
hello NAACL friends I'm giving a keynote today at RepL4NLP at 1:30PM local time, come say hi! I'll mostly be musing about things with light research discussions
Screenshot of a slide that says "what does it take to convince ourselves that a system is exhibiting compositionality?" with a side comment "mostly AI, but humans too!!" for the word system
1171
Freda Shi @fredashi.bsky.social · 29/04/2025
If you are also at NAACL, let's chat!
000
Freda Shi @fredashi.bsky.social · 29/04/2025
On my way to NAACL✈️! If you're also there and interested in grounding, don't miss our tutorial on "Learning Language through Grounding"! Mark your calendar: May 3rd, 14:00-17:30, Ballroom A. Another exciting collaboration with @marstin.bsky.social @kordjamshidi.bsky.social, Jiayuan, and Joyce!
100
Reposted by Freda Shi
Yi (Joshua) Ren @joshuaren.bsky.social · 21/04/2025
📢Curious why your LLM behaves strangely after long SFT or DPO? We offer a fresh perspective—consider doing a "force analysis" on your model’s behavior. Check out our #ICLR2025 Oral paper: Learning Dynamics of LLM Finetuning! (0/12)
16915
Freda Shi @fredashi.bsky.social · 23/04/2025
✈ Just landed in Singapore for #ICLR 2025! DM or email me if you'd like to chat about - Grounded language acquisition and learning - What do vision-language models "know," and what they don't - (Computational) linguistics with/for language models, especially grounded LMs (1/)
130
Freda Shi @fredashi.bsky.social · 28/03/2025
I received a review like this five years ago. It’s probably the right time now to share it with everyone who wrote or got random discouraging reviews from ICML/ACL.
1635
Reposted by Freda Shi
Nouha Dziri @nouhadziri.bsky.social · 27/03/2025
I still can't comprehend how an AC (a professor) accepts the role but then never responds back to emails and never completes their tasks😖! We need urgent 5 emergency reviewers to complete reviews for ACL by the end of today. Area: Ethics, Bias, and Fairness. Please reach out if you can help! Thanks🙏
12411
Freda Shi @fredashi.bsky.social · 17/03/2025
I was assigned to review three, and fortunately got a decent one. This is probably just the submission distribution nowadays.
030
Freda Shi @fredashi.bsky.social · 17/03/2025
🤖 I just made the slides public for this talk. TL; DR: how we computer scientists adapt insights from linguistics to analyze and improve our models. Comments & discussion are welcomed; the recording from Vector is forthcoming. docs.google.com/presentation...
docs.google.com
Vector-240314.pptx
Linguistic Insights Deepen our Understanding of AI Systems The Cases of Reference Frames and Logical Reasoning Freda Shi University of Waterloo, Vector Institute, Canada CIFAR AI Chair fhs@uwaterloo.c...
100
Reposted by Freda Shi
Gautam Kamath @gautamkamath.com · 11/03/2025
There is a postdoctoral position for research on adversarial robustness at UBC, supervised by Mathias Lecuyer (@mathias-lecuyer.bsky.social), Geoff Pleiss, and Nidhi Hegde. Please spread the word! 🇨🇦 docs.google.com/forms/d/e/1F...
docs.google.com
Application for a postdoctoral position on adversarial robustness
Location - Work primarily takes place at UBC, Vancouver, Computer Science department. Position - We are recruiting a postdoctoral researcher for a funded position, under the joint supervision of Math...
1185
Reposted by Freda Shi
Sharon Machlis @smachlis.bsky.social · 09/03/2025
Customizing ggplot for yourself or your organization - slides and #RStats code from @pewresearch.org 's #NICAR25 presentation github.com/pewresearch/...
github.com
GitHub - pewresearch/pewplots-nicar-2025: Materials from NICAR 2025 session "Customizing ggplot for yourself or your organization"
Materials from NICAR 2025 session "Customizing ggplot for yourself or your organization" - pewresearch/pewplots-nicar-2025
0132
Freda Shi @fredashi.bsky.social · 09/03/2025
Another common suggestion I gave to almost all ICML papers I reviewed this time: whenever you write (CS/ML style) theory, please ensure all variables are defined. Although I can complete it with educated guess for most times, it's not always the case.
040
Freda Shi @fredashi.bsky.social · 07/03/2025
It's frustrating to see increasingly many false references in papers I'm reviewing. The authors write "claim about X (Y et al., 2019)" in the paper---as someone who has carefully read Y et al. (2019) and has done follow-up work, I'm confident there's no claim about X at all in it. #AmReviewing
330
Freda Shi @fredashi.bsky.social · 06/03/2025
The experiment I'm most excited about in this work is the disentanglement between language and domain. Even adding a semantically irrelevant sentence in another language to the demonstration---which just simply increases linguistic diversity---helps improve the performance!
030
Freda Shi @fredashi.bsky.social · 06/03/2025
Somewhat surprisingly, high-resource languages with non-Latin alphabets serve as better demonstrations than English. Increasing linguistic diversity in LLMs is not only about linguistic diversity itself---it's also about performance!
1112
Freda Shi @fredashi.bsky.social · 06/03/2025
📜New analysis paper on the surprising effectiveness of multilingual chain of thought. Linguistic diversity in prompts helps solve problems in low-resource languages, even if there's no alphabetical overlap between the demonstrations and the target language.
120
Reposted by Freda Shi
Yoshitomo Matsubara @yoshitomo-matsubara.net · 03/03/2025
[Update & Hiring] Last month, I joined Search and Recommendation Science team at Yahoo! as a Research Scientist🚀 Our team is hiring another Research Scientist🚀🚀 Send me a DM if you have strong research background in deep learning and NLP and want to be considered for the position🙋🙋🙋
493
Reposted by Freda Shi
EMNLP @emnlpmeeting.bsky.social · 27/02/2025
ACL Rolling Review and the EMNLP PCs are seeking input on the current state of reviewing for *CL conferences. We would love to get your feedback on the current process and how it could be improved. To contribute your ideas and opinions, please follow this link! forms.office.com/r/P68uvwXYqfemn
forms.office.com
Microsoft Forms
11113
Freda Shi @fredashi.bsky.social · 27/02/2025
This reminds me of the (rare) failure cases of training/finetuning NNs. Sometimes we got quite poor results without much reason. Restarting with another random seed magically works. Ideas of this work could be the key to understanding what's going on.
1111
Reposted by Freda Shi
Naomi Saphra @nsaphra.bsky.social · 25/02/2025
Ever looked at LLM skill emergence and thought 70B parameters was a magic number? Our new paper shows sudden breakthroughs are samples from bimodal performance distributions across seeds. Observed accuracy jumps abruptly while the underlying accuracy DISTRIBUTION changes slowly!
Distributional Scaling Laws for Emergent Capabilities
Rosie Zhao, Tian Qin, David Alvarez-Melis, Sham Kakade, Naomi Saphra
In this paper, we explore the nature of sudden breakthroughs in language model performance at scale, which stands in contrast to smooth improvements governed by scaling laws. While advocates of "emergence" view abrupt performance gains as capabilities unlocking at specific scales, others have suggested that they are produced by thresholding effects and alleviated by continuous metrics. We propose that breakthroughs are instead driven by continuous changes in the probability distribution of training outcomes, particularly when performance is bimodally distributed across random seeds. In synthetic length generalization tasks, we show that different random seeds can produce either highly linear or emergent scaling trends. We reveal that sharp breakthroughs in metrics are produced by underlying continuous changes in their distribution across seeds. Furthermore, we provide a case study of inverse scaling and show that even as the probability of a successful run declines, the average performance of a successful run continues to increase monotonically. We validate our distributional scaling framework on realistic settings by measuring MMLU performance in LLM populations. These insights emphasize the role of random variation in the effect of scale on LLM capabilities.
36615
Freda Shi @fredashi.bsky.social · 22/02/2025
ARR February Authors: If you keep receiving confusing reminder emails about volunteer qualifications, please ask your nominated volunteer to update the OpenReview profile with either ACL Anthology or DBLP. This is how PCs verify qualification. RT appreciated!
011
Reposted by Freda Shi
Yoav Artzi @yoavartzi.com · 14/02/2025
I am looking for a postdoc. A serious-looking call coming soon, but this is to get it going. Topics include (but not limited to): LLMs (🫢!), multimodal LLMs, interaction+learning, RL, intersection with cogsci, ... see our work to get an idea: yoavartzi.com/pubs Plz RT 🙏
yoavartzi.com
Publications
12413
Reposted by Freda Shi
Kanishka Misra @kanishka.bsky.social · 23/01/2025
Excited that this got accepted at naacl/@naaclmeeting.bsky.social 2025! Massive kudos to Juan Diego and Aaron for being the best co-authors and colleagues one could ask for! 🙏
1383
Freda Shi @fredashi.bsky.social · 27/12/2024
RepL4NLP (collocated with NAACL 2025) is inviting submissions!
192
Reposted by Freda Shi
ACL 2027 @aclmeeting.bsky.social · 16/12/2024
We invite nominations to join the ACL2025 PC as reviewer or area chair(AC). Review process through ARR Feb cycle. Tentative timeline: Review 1-20 Mar 2025, Rebuttal is 26-31 Mar 2025. ACs must be available throughout the Feb cycle. Nominations by 20 Dec 2024: shorturl.at/TaUh9 #NLProc #ACL2025NLP
forms.gle
Volunteer to join ACL 2025 Programme Committee
Use this form to express your interest in joining the ACL 2025 programme committee as a reviewer or area chair (AC). The review period is 1st to 20th of March 2025. ACs need to be available for variou...
01112