Sign in

Ehud Reiter

@ehudreiter.bsky.social
157 followers 42 following 249 posts

Emeritus (retired) professor at Aberdeen (UK), specialising in Natural Language Generation, evaluation, and AI for Healthcare

PostsRepliesMedia
Ehud Reiter @ehudreiter.bsky.social · 02/10/2026
Going to EMNLP 2026 in Budapest? We are looking for 20-30 EMNLP attendees to help us understand how you handle language barriers during a conference trip. Contact Yujun Wang: y.wang2.25@abdn.ac.uk @weizhaonlp.bsky.social, @markar.bsky.social, Pinzhen Chen, Yujun Wang
022
Reposted by Ehud Reiter
Thomas Dietterich @tdietterich.bsky.social · 01/10/2026
Corrected link: blog.arxiv.org/2026/10/01/u...
blog.arxiv.org
Fair Moderation, Equitable Access, and AI: arXiv’s Updated Rate Limit Policy
arXiv, and the scientific community at large, are facing a watershed moment. Scholarly publishing is currently changing at a rapid pace, and we are seeing a…
072
Ehud Reiter @ehudreiter.bsky.social · 01/10/2026
Next week I am giving an invited talk at a workshop on AI in Rabbinic/Talmud studies, in Israel. Never done anything like this before, will be really interesting to see how AI is used in this space.
020
Ehud Reiter @ehudreiter.bsky.social · 28/09/2026
I see lots of posts "our lab has N papers in VENUE". Congratulations, but what is more useful to me is short 3-5 line descriptions of a paper which will convince me to open it and read abstract. With link to final (not submitted) version
102
Ehud Reiter @ehudreiter.bsky.social · 25/09/2026
Excited that my student Yujun Wang will present (oral) at EMNLP on "Beyond Accuracy: Community Perspectives on Machine Translation". Many of my students have looked at stakeholder requirements/concerns. Need to understand in order to align AI with humans! arxiv.org/abs/2606.09655
arxiv.org
Beyond Accuracy: Community Perspectives on Machine Translation
Despite remarkable progress in machine translation (MT), non-AI communities have raised growing concerns about MT systems, suggesting a noticeable gap between technical advancement and the needs of re...
091
Ehud Reiter @ehudreiter.bsky.social · 23/09/2026
AI is improving productivity in science, but so far this has led to a flood of low-quality spam and large amounts of so-so papers. I dont see increase in the really high-quality research that science needs! Maybe blame perverse incentives in science that priotitise quantity.
142
Ehud Reiter @ehudreiter.bsky.social · 22/09/2026
If AI is risk to humanity, because of human stupidity, not clever AI. In 1983 hallucinating Soviet satellite almost started nuclear war, stopped by sensible human. In 2026, hallucinating AI almost led US military to attack Chinese ship, due to poor human oversight edition.cnn.com/2026/09/18/p...
edition.cnn.com
Exclusive: US military had close call after using AI for false intelligence report, sources say | CNN Politics
The episode shows the risks of using this new, relatively poorly understood technology in the middle of the Iran war
021
Ehud Reiter @ehudreiter.bsky.social · 20/09/2026
love the phrase "paper-shaped objects" "modern AI systems can generate paper-shaped objects even for users who are unfamiliar with the related work and writing conventions for a given area, and who are unable to understand or verify any of a paper’s main claims." blog.iclr.cc/2026/09/02/s...
blog.iclr.cc
Submission policies for ICLR 2027 – ICLR Blog
052
Ehud Reiter @ehudreiter.bsky.social · 17/09/2026
I see submissions from people who are trying to learn how to do research using LLMs. This is dubious, and do not expect TACL (etc) to teach research skills! I suggest you get involved with local research: attend workshop, join seminar series, help in a research project, etc
051
Ehud Reiter @ehudreiter.bsky.social · 17/09/2026
New blog: Localising NLG content in a driving feedback system NLG systems need to localise content for different countries and cultures. Two of my PhD students have written a nice case study paper about this for INLG, in a driving feedback domain. ehudreiter.com/2026/09/17/l...
ehudreiter.com
Localising NLG content in a driving feedback system
NLG systems need to localise content for different countries and cultures. Two of my PhD students developed driving feedback systems in UK and Nigeria, and they have written a nice paper about how …
020
Reposted by Ehud Reiter
Transactions on Machine Learning Research @tmlrorg.bsky.social · 16/09/2026
TMLR has faced a deluge of submissions, necessitating stricter desk rejection policies due to limited reviewer capacity Co-EiC Nihar Shah reached out to authors of 10 papers slated for desk reject. Could they answer questions about their *own* submission? medium.com/@TmlrOrg/ask...
medium.com
Asking Authors About Their Own Papers
By Nihar B. Shah
116665
Ehud Reiter @ehudreiter.bsky.social · 15/09/2026
Congratulations to my student Mengxuan Sun for passing her PhD viva with minor corrections! Mengxuan's thesis was on Designing and Evaluating Natural Language Processing Tools for Patient-Centred Communication in Cancer Care
000
Ehud Reiter @ehudreiter.bsky.social · 14/09/2026
Years ago you only needed papers when applying for job as post-doc or faculty. Now secondary-school students are using LLMs to write papers, to help them get into a good UG university... Inevitable result is wave of AI slop, which (as below) "authors" may not even understand.
122
Ehud Reiter @ehudreiter.bsky.social · 08/09/2026
New blog: Desk rejection is an unfortunate necessity Editors can desk reject papers without sending them to reviewers, eg because of wrong format or weak content. I dont like doing this, but flood of “AI slop” papers makes this an unfortunate necessity. ehudreiter.com/2026/09/08/d...
ehudreiter.com
Desk rejection is an unfortunate necessity
Editors in venues can desk reject papers without sending them to reviewers, for reasons such as wrong format, LLM-written, topic out of scope, and weak content. I dont like doing this, but the floo…
022
Ehud Reiter @ehudreiter.bsky.social · 04/09/2026
Im doing format desk reject for TACL. Figs must be readable, and appendices are limited to 5 pages for reproducability, and 3 for additional figs/tables (NOT free text). Good scientific writing is concise, we dont want shrunken figs or bloated appendices!
141
Ehud Reiter @ehudreiter.bsky.social · 01/09/2026
Position change: Professor -> Emeritus Professor. I am now officially retired!
151
Ehud Reiter @ehudreiter.bsky.social · 29/08/2026
Interesting op-ed by Ciaran Martin in Economist (paywall) about threat of AI hackers. Points out that if a cybersecurity company allowed hacking software to escape a testing environment, it would probably face lawsuits and prosecution. Maybe different rules apply to OpenAI and Anthropic?
120
Ehud Reiter @ehudreiter.bsky.social · 26/08/2026
My email e.reiter@abdn.ac.uk may temporarily stop working after 31 Aug (when I retire and become emeritus), if you need to reach me you should CC s04er6@abdn.ac.uk
010
Ehud Reiter @ehudreiter.bsky.social · 25/08/2026
New blog (personal): Cycling in Northern Ireland This is a personal blog, about a recent bike trip which was mostly to Northern Ireland. ehudreiter.com/2026/08/25/c...
ehudreiter.com
Cycling in Northern Ireland
This is a personal blog, about a recent bike trip which was mostly to Northern Ireland.
000
Ehud Reiter @ehudreiter.bsky.social · 24/08/2026
Let me know if you are interested in joining the reviewer team at TACL (Transactions of ACL) journal. This is "old fashioned" reviewing - one paper at a time, sent to knowledgeable reviewers. Reviewers should have a PhD, and at least 1K citations and 10 first-author papers.
152
Ehud Reiter @ehudreiter.bsky.social · 07/08/2026
In AI Healthcare, skyrocketing benchmakes do not lead to real-world impact. Real world is messy, BM do not predict impact, poor fit to what to what users want, and lack of ambition in changing system. I suspect this applies to many other uses of AI! ehudreiter.com/2026/08/05/a...
ehudreiter.com
AI for Healthcare: Great benchmark scores but minimal impact
The paradox of AI in healthcare is that while benchmarks show soaring (and sometimes superhuman) performance, in the real world AI is not actually helping people much. This is partially because rea…
010
Ehud Reiter @ehudreiter.bsky.social · 05/08/2026
New blog: AI for Healthcare: Great benchmarks but minimal impact AI in healthcare benchmarks show soaring perf, but in the real world AI is not actually helping people much. Reasons include messy real world, weak eval, deployment, and lack of utility ehudreiter.com/2026/08/05/a...
ehudreiter.com
AI for Healthcare: Great benchmarks but minimal impact
The paradox of AI in healthcare is that while benchmarks show soaring (and sometimes superhuman) performance, in the real world AI is not actually helping people much. This is partially because rea…
032
Ehud Reiter @ehudreiter.bsky.social · 31/07/2026
Somebody emailed me to ask about being a reviewer at TACL, and unfortunately this ended up in my spam folder and got deleted. Apologies! If you emailed me about this and did not get a response, please email me again
001
Ehud Reiter @ehudreiter.bsky.social · 28/07/2026
Congrats to my student Mengxuan Sun for submitting her PhD thesis! Mengxuan is joint CS/Medicine and is looking at using NLP to help cancer patients.
042
Ehud Reiter @ehudreiter.bsky.social · 24/07/2026
Reading about OpenAI AI attacking Huggingface reminded me a bit of the Internet worm released by Morris in 1988 (shows my age). Details completely different, but in both cases an "experiment" got out of control, did damage, and drew media attention en.wikipedia.org/wiki/Morris_...
en.wikipedia.org
Morris worm - Wikipedia
030
Ehud Reiter @ehudreiter.bsky.social · 23/07/2026
Really interesting paper on evaluating usability of speech translation, by asking recipients to answer questions about key information (in scenarios such as communicating health problems) arxiv.org/abs/2606.06177
arxiv.org
Ouvia: A User-centered Framework for Measuring Usability of Speech Translation in Real-World Communication Scenarios
Speech translation (ST) is increasingly adopted in user applications, yet its evaluation largely focuses on decontextualized testbeds and holistic quality, rather than end users' communication needs. ...
020
Ehud Reiter @ehudreiter.bsky.social · 23/07/2026
TACL is looking for more co-editors, apply by 31 July if interested www.aclweb.org/portal/conte... Also TACL is always looking for people to join its standing reviewer pool. In general, TACL reviewers should have PhD, 10 first-author papers, and 1K citations
aclweb.org
Call for nominations for TACL Co-Editors-in-Chief, 2027–2029 | ACL Member Portal
001
Ehud Reiter @ehudreiter.bsky.social · 22/07/2026
Somewhat scary paper showing dubious Kaggle medical data sets being used in both research papers and real-world clinical practice. Blind use of dubious data is not just a problem in NLP and AI... doi.org/10.1186/s129...
doi.org
Evidence of unreliable data and poor data provenance in clinical prediction model research and clinical practice - BMC Medicine
Background Clinical prediction models are often created using large routinely collected datasets. It is essential that prediction models are developed with appropriate data and methods and transparently reported to ensure that decisions are based on reliable predictions. Kaggle is a popular competition and data repository website where users learn and apply analysis skills on a range of datasets. Methods We identified two large, publicly available Kaggle datasets, on stroke and diabetes, that lack clear data provenance, but are widely used in clinical prediction models in peer reviewed publications. We used exploratory analyses to examine the quality of data and reporting of information using nine items from the TRIPOD+AI statement checklist. Results Data provenance assessment using nine TRIPOD+AI items revealed major deficiencies, with minimal details for either dataset including no information on when, where, why or how the data were collected. The authenticity of both datasets could not be verified and have no reliable provenance of authenticity and should not be used for informing research or practice. From these two datasets, we found 125 clinical prediction model studies. Three prediction models had evidence of use in clinical practice, one model was cited in a medical device patent, and the models were cited in 86 review articles. Conclusions We recommend that journals and data repositories mandate data provenance reporting to safeguard published research. Prediction models based solely on inauthentic or unreliable datasets should never be used to directly inform decisions on patient care.
021
Ehud Reiter @ehudreiter.bsky.social · 20/07/2026
New blog: Memories of being an NL researcher in 1990 I “reminisce ” about being an NL researcher in 1990, when I got my PhD. Community was much smaller than 2026, but in many ways it was nicer, including less pressure and a more open research culture. ehudreiter.com/2026/07/20/m...
ehudreiter.com
Memories of being an NL researcher in 1990
Since I am about to retire, I decided to “reminisce ” about what it was like to be an NL researcher in 1990, when I got my PhD. The community was much smaller than 2026, but in many way…
010
Ehud Reiter @ehudreiter.bsky.social · 17/07/2026
AI safety news UK: AISI eliminates its societal resilience team (which covered things like suicide risk from bots) China: new government regulations aim to reduce emotional dependence on AI Looks like China takes emotional risks far more seriously than UK (or US)...
041
Ehud Reiter @ehudreiter.bsky.social · 16/07/2026
As reviewer, I pointed out a fundamental misconception in paper. Authors reponded that they had seen many published ACL (etc) papers with same problem, its also embedded in benchmarks. If they cannot trust what they see published in ACL, how can they build on other peoples work? I dont have answer..
071
Reposted by Ehud Reiter
Nils Feldhus @nfel.bsky.social · 13/07/2026
📢 Call for Papers: YNLG 2026 The Young Researchers in Natural Language Generation workshop is a 2d in-person event part of INLG 2026 @inlg.bsky.social in Utrecht 🇳🇱, with poster sessions, keynote talks, roundtable discussions, and a one-day hackathon. Due: August 10, 2026 ynlg-workshop.github.io
0109
Ehud Reiter @ehudreiter.bsky.social · 09/07/2026
New blog: What is the purpose of ACL conferences? What is the main pupose of ACL conferencess: meeting people, enhancing CVs, identifying good papers, or providing a home for exciting science? The best reviewing system depends on the goal of our conf ehudreiter.com/2026/07/09/w...
ehudreiter.com
What is the purpose of ACL conferences?
The reviewing system for ACL conferences is struggling. In order to fix it, we should be clear about what the main pupose of the conferences is: meeting people, enhancing CVs, identifying good pape…
0114
Reposted by Ehud Reiter
Emiel van Miltenburg @evanmiltenburg.bsky.social · 08/07/2026
Our book review of @ehudreiter.bsky.social 's book on Natural Language Generation is finally out: direct.mit.edu/coli/article... #NLProc
direct.mit.edu
Natural Language Generation by Ehud Reiter
063
Ehud Reiter @ehudreiter.bsky.social · 08/07/2026
Congratulations to Zeerak and Dirk for their well-deserved Test of Time award on hate speech detection! Its also striking to see a Test of Time award go to a paper presented at a Student Research Workshop. Maybe SRW are more open to crazy new ideas than xACL main conferences?
170
Ehud Reiter @ehudreiter.bsky.social · 06/07/2026
I'm giving a talk to the local branch of the Royal Statistical Society next week. Looking forward to it! Statisticians have been building models for a long time, we can learn from them. They also take data issues very seriously unlike most machine learning people I know.
000
Reposted by Ehud Reiter
Techmeme @techmeme.com · 05/07/2026
AWS says Mechanical Turk will no longer accept new customers and that it is placing the crowdsourcing service in maintenance, signaling its future retirement (Simon Sharwood/The Register) Main Link | Techmeme Permalink
02916
Ehud Reiter @ehudreiter.bsky.social · 02/07/2026
TACL is looking for more Editors (EiC) www.aclweb.org/portal/conte...
aclweb.org
Call for nominations for TACL Co-Editors-in-Chief, 2027–2029 | ACL Member Portal
011
Ehud Reiter @ehudreiter.bsky.social · 01/07/2026
Are there any papers which won both Best Paper award when presented and a Test of Time award later? I cannot find any in xACL. If so, suggests BP awards do not signify long-term research impact
010
Ehud Reiter @ehudreiter.bsky.social · 29/06/2026
I am becoming an editor (EiC) at TACL (Transactions of ACL) journal. Not what I expected to be doing in retirement, but I am a strong believer in journals as the best way to present research, and I want to help TACL become an even better venue for publishing NLP research.
0111
Reposted by Ehud Reiter
Emiel van Miltenburg @evanmiltenburg.bsky.social · 26/06/2026
INLG (@inlg.bsky.social) submissions are now open! Please submit all of your work on Natural Language Generation. Submit systems, demos, experiments, position papers, squibs, linguistic analyses, as long as it is NLG-related. For more information, see: 2026.inlgmeeting.org #NLProc
2026.inlgmeeting.org
INLG2026
The 19th International Natural Language Generation Conference is scheduled to be held in Utrecht, the Netherlands from October 17 to 21, 2026.
164
Ehud Reiter @ehudreiter.bsky.social · 26/06/2026
Really interesting survey about peer review www.cs.cmu.edu/~nihars/prep... (extended version of CACM paper). Lots of interesting and indeed worrying observations about peer review
cs.cmu.edu
010
Ehud Reiter @ehudreiter.bsky.social · 26/06/2026
New blog: Future of NLG evaluation In a recent position paper, I argued that NLG evaluation in the future needs to be become more rigorous. It also needs to move beyond benchmarks, and focus more on impact, qualitative, and safety evaluation. ehudreiter.com/2026/06/26/f...
ehudreiter.com
Future of NLG evaluation
In a recent position paper, I argued that NLG evaluation in the future needs to be become more rigorous. It also needs to move beyond benchmarks, and focus more on impact, qualitative, and safety e…
020
Ehud Reiter @ehudreiter.bsky.social · 25/06/2026
My student @iniakpothompson.bsky.social developed a safe driving app for Nigeria. Evaluation was promising, and he is currently trying to convert his research project into a reality which saves lives in Nigeria. If anyone wants to learn more about this work, see iniakpothompson.com/research-pro...
iniakpothompson.com
Safe Drive Africa — Culturally Attuned AI for Road Safety in Nigeria | Iniakpokeikiye Peter Thompson
Safe Drive Africa: a completed PhD research programme combining smartphone telematics, ML-based alcohol detection, and culturally attuned AI feedback to reduce unsafe driving in Nigeria. 19.4% reducti...
011
Ehud Reiter @ehudreiter.bsky.social · 24/06/2026
Really enjoyed helping my colleague Jakub Zbrzezny from our Divinity dept look at how well LLMs can translate biblical materials in a local Arabic dialect. Not surprisingly, LLMs good at translating out of dialect, but struggle to translate into dialect. aclanthology.org/2026.retroev...
aclanthology.org
The Arabic Bible as an Evaluation Tool: The Case Study of the Khalīlī Arabic Dialect
Jakub Zbrzeżny, Ehud Reiter, Wei Zhao. Proceedings of the 1st Symposium on Natural Language Generation Evaluations. 2026.
021
Ehud Reiter @ehudreiter.bsky.social · 22/06/2026
Radical suggestion. Why not *lower* prestige of xACL, for example by including Findings in main conf. Then people chasing N papers in "top" venues will submit elsewhere, making xACL more manageable. Let Neurips deal with AI slop...
121
Reposted by Ehud Reiter
Saad Mahamood @saad.me.uk · 17/06/2026
I am pleased to announce that the proceedings for RetroEval have now been published on ACL Anthology: aclanthology.org/volumes/2026...
aclanthology.org
Proceedings of the 1st Symposium on Natural Language Generation Evaluations - ACL Anthology
011
Ehud Reiter @ehudreiter.bsky.social · 17/06/2026
Reviewed a paper which extensively used a dataset I helped to create, but showed zero awareness of data issues and the domain. I guess people just want data to throw into LLMs, dont care about data issues even when these are carefully explained in dataset paper.
192
Ehud Reiter @ehudreiter.bsky.social · 15/06/2026
[AI leads to] an expansion of individual scientists’ impact but a contraction in collective science’s reach, as AI-augmented work moves collectively towards areas richest in data... AI tools seem to automate established fields rather than explore new ones www.nature.com/articles/s41...
nature.com
Artificial intelligence tools expand scientists’ impact but contract science’s focus - Nature
Artificial intelligence boosts individual scientists’ output, citations and career progression, but collectively narrows research diversity and reduces collaboration, concentrating work in data-rich a...
041
Reposted by Ehud Reiter
Saad Mahamood @saad.me.uk · 09/06/2026
We had a fantastic time last week discussing the current challenges in NLG evaluation and celebrating the career of @ehudreiter.bsky.social. Pictures and a few video clips are now available: retroeval.github.io/_pages/media I would like to thank @uniofaberdeen.bsky.social for their support.
042