Sign in

Tim Baumgärtner

@timbmg.bsky.social
99 followers 271 following 11 posts

👨‍💻 NLP PhD Student @ukplab.bsky.social

PostsRepliesMedia
Tim Baumgärtner @timbmg.bsky.social · 01/07/2026
Had a great time presenting our ACL paper "SciCoQA" at the ELLIS / hessian.AI poster session this week! Really enjoyed the conversations and the thoughtful questions from everyone who came by. You can read more about SciCoQA in our #ACL2026 paper: aclanthology.org/2026.acl-lon...
aclanthology.org
SciCoQA: Quality Assurance for Scientific Paper–Code Alignment
Tim Baumgärtner, Iryna Gurevych. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
110
Reposted by Tim Baumgärtner
ACL Rolling Review (ARR) @aclrollingreview.bsky.social · 26/05/2026
📢 ARR-May reviewers can now try REVAS, an experimental review support tool. REVAS gives feedback on review quality criteria and ARR reviewer heuristics, but does not suggest review content or scores. 🔗 aclrollingreview.org/revas-may26 #ARR #EMNLP #ACL #NLProc
aclrollingreview.org
Using the REVAS review assistant tool in the ARR-May cycle
In the March 2026 ARR cycle, we for the first time encouraged reviewers to try the experimental review support tool called REVAS. The instructions were communicated to the cycle reviewers, but we real...
052
Tim Baumgärtner @timbmg.bsky.social · 08/05/2026
"An article about computational science in a scientific publication is not the scholarship itself, it is merely advertising of the scholarship." Claerbout, 1992 Decades later, code is routinely released with papers. But does the code match the ad?
100
Tim Baumgärtner @timbmg.bsky.social · 17/04/2026
REVAS is a peer review assistant I've been working on with colleagues at MBZUAI. The idea is simple: better reviews make better science. If you're reviewing for ARR this cycle, give it a try.
000
Tim Baumgärtner @timbmg.bsky.social · 10/11/2025
💡 TIL @overleaf.com is basically a git repo. In my research workflow, I directly added it as submodule to my code repo. Now I can produce figures and tables, and have them magically uploaded to Overleaf just by pushing the repo. No more renaming, keeping versions straight, and manual uploading 😇
000
Tim Baumgärtner @timbmg.bsky.social · 04/11/2025
💡 TIL, it's super easy to fetch data from Google Sheets into Pandas. Makes it really convenient to annotate some data. Previously, I was always downloading CSVs, losing track of file versions, and loading and merging them sluggishly in Python. 👉 find the code here: gist.github.com/timbmg/6c2d6...
000
Reposted by Tim Baumgärtner
UKP Lab @ukplab.bsky.social · 25/04/2025
🔍 𝗪𝗮𝗻𝘁 𝘁𝗼 𝗲𝘃𝗮𝗹𝘂𝗮𝘁𝗲 𝗺𝗼𝗱𝗲𝗹𝘀 𝗼𝗻 𝘀𝗰𝗶𝗲𝗻𝘁𝗶𝗳𝗶𝗰 𝗤𝗔, 𝗯𝘂𝘁 𝘆𝗼𝘂𝗿 𝗱𝗮𝘁𝗮𝘀𝗲𝘁 𝗹𝗮𝗰𝗸𝘀 𝗿𝗲𝗮𝗹-𝘄𝗼𝗿𝗹𝗱 𝗾𝘂𝗲𝘀𝘁𝗶𝗼𝗻𝘀 𝗮𝘀𝗸𝗲𝗱 𝗯𝘆 𝗲𝘅𝗽𝗲𝗿𝘁𝘀? 🚀 PeerQA is the solution: a dataset with questions from peer reviews and answers from the original authors. (1/🧵) #NLProc
121
Reposted by Tim Baumgärtner
Martin Tutek @mtutek.bsky.social · 21/02/2025
🚨🚨 New preprint 🚨🚨 Ever wonder whether verbalized CoTs correspond to the internal reasoning process of the model? We propose a novel parametric faithfulness approach, which erases information contained in CoT steps from the model parameters to assess CoT faithfulness. arxiv.org/abs/2502.14829
arxiv.org
Measuring Faithfulness of Chains of Thought by Unlearning Reasoning Steps
When prompted to think step-by-step, language models (LMs) produce a chain of thought (CoT), a sequence of reasoning steps that the model supposedly used to produce its prediction. However, despite mu...
24813
Reposted by Tim Baumgärtner
UKP Lab @ukplab.bsky.social · 18/02/2025
𝗙𝗮𝗰𝘁-𝗖𝗵𝗲𝗰𝗸𝗶𝗻𝗴 𝗶𝗻 𝘁𝗵𝗲 𝗔𝗴𝗲 𝗼𝗳 𝗔𝗜 – 𝗔 𝗧𝗮𝗹𝗸 𝗯𝘆 𝗜𝗿𝘆𝗻𝗮 𝗚𝘂𝗿𝗲𝘃𝘆𝗰𝗵 @𝗔𝗜 𝗳𝗼𝗿 𝗚𝗼𝗼𝗱 Misinformation is a new weapon disrupting public debates, scientific discussions, and political decisions. How can we identify and counter misleading content? (1/🧵)
aiforgood.itu.int
Towards real-world fact-checking with large language models
Misinformation poses a growing threat to our society. It has a severe impact on public health by promoting fake cures fear and distrust. Current research
131
Reposted by Tim Baumgärtner
Florent Daudens @fdaudens.bsky.social · 11/02/2025
🤔 An Energy Star for AI? Introducing AI Energy Score: First-ever rating system comparing 166 AI models' energy consumption! From LLaMa to Gemma, get transparent ⭐️1-5 efficiency ratings. Incredible work led by @sashamtl.bsky.social huggingface.co/blog/sasha/a...
0258
Tim Baumgärtner @timbmg.bsky.social · 27/01/2025
Excited to share that our Paper "PeerQA: A Scientific Question Answering Dataset from Peer Reviews" as been accepted to #NAACL2025 Looking forward to presenting it in Albuquerque 🏜️!
050
Reposted by Tim Baumgärtner
Bertram Højer @brtrm.bsky.social · 04/12/2024
What do YOU mean by "intelligence", and does ChatGPT fit your definition? We collected the major criteria used in CogSci and other fields, and designed a survey to find out! Access link: www.survey-xact.dk/collect Code: 4S7V-SN4M-S536 Time: 5-10 mins
bertramhojer.github.io
Perspectives on Intelligence: Community Survey
Research survey exploring how NLP/ML/CogSci researchers define and use the concept of intelligence.
23213