Sign in

siddheshp.bsky.social

@siddheshp.bsky.social
51 followers 634 following 11 posts

Grad Student; Into Multilingual NLP

PostsRepliesMedia
siddheshp.bsky.social @siddheshp.bsky.social · 15/09/2026
People use LLMs to answer subjective, culturally situated questions. Most benchmarks ask whether the answer is correct, but they overlook the framing features (eg, insider positioning) 📄 New #EMNLP'26 paper: "Not What, But How: A Communicative Audit of LLM Response Framing"
141
Reposted by @siddheshp.bsky.social
Hanna Wallach @hannawallach.bsky.social · 15/06/2025
Check out the camera-ready version of our ICML position paper ("Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge") to learn more!!! arxiv.org/abs/2502.00561 (6/6)
arxiv.org
Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge
The measurement tasks involved in evaluating generative AI (GenAI) systems lack sufficient scientific rigor, leading to what has been described as "a tangle of sloppy tests [and] apples-to-oranges com...
34411
Reposted by @siddheshp.bsky.social
Yoav Goldberg @yoavgo.bsky.social · 17/12/2024
i mean, people have different goals, and if you cared about some niche aspect of query focused multi doc sum before, it is legit to continue. or you can switch focus and start thinking of HCI. the second became much more possible now, the first maybe hasnt.
141
Reposted by @siddheshp.bsky.social
Edward Grefenstette @egrefen.bsky.social · 19/11/2024
🌶️(?) take: Agents are somehow hot right because people realized that LLM output can be interpreted as a DSL which directs side effects in the world (e.g. tool calls) rather than just returning text in a chat/autocomplete sense. What are the open challenges? A 🧵... [1/11]
916531
Reposted by @siddheshp.bsky.social
Marco @mcognetta.bsky.social · 11/11/2024
#EMNLP has a nice set of tokenization/subword modeling papers this year. It's a good mix of tokenization algorithms, tokenization evaluation, tokenization-free methods, and subword embedding probing. Lmk if I missed some! Here is a list with links + presentation time (in chronological order).
54816
siddheshp.bsky.social @siddheshp.bsky.social · 11/11/2024
We are excited to share our comprehensive survey on cultural awareness in #LLMs! 🗺️ [Was posted on X a few days before] We reviewed 300+ papers across diverse modalities (language, vision-language, etc.) arxiv.org/abs/2411.00860
arxiv.org
Survey of Cultural Awareness in Language Models: Text and Beyond
Large-scale deployment of large language models (LLMs) in various applications, such as chatbots and virtual assistants, requires LLMs to be culturally sensitive to the user to ensure inclusivity. Cul...
120