Sign in

Institute of Formal and Applied Linguistics

@ufal.mff.cuni.cz
651 followers 71 following 218 posts

Computational linguistics • Natural language processing • Formal linguistics • Machine translation | at Faculty of Mathematics and Physics, Charles University

PostsRepliesMedia
Institute of Formal and Applied Linguistics @ufal.mff.cuni.cz · 05/10/2026
From data highways to oral history: ÚFAL researchers presented at the CLARIN Conference in Barcelona. Ondřej Košarko & Pavel Straňák compared EOSC-CZ & CLARIN infrastructure architectures, while Christopher Brückner showcased MalachNER, turning oral history into searchable data. #NLP #CLARIN2026
011
Reposted by Institute of Formal and Applied Linguistics
Jindřich Libovický @jlibovicky.bsky.social · 21/09/2026
I really enjoyed giving a talk at the Machine Translation Marathon 2026 about our research on tokenization evaluation. Thank you for having me! For those interested, the slides from my presentation are available here: docs.google.com/presentation...
072
Institute of Formal and Applied Linguistics @ufal.mff.cuni.cz · 02/09/2026
Congrats, Adnan! 🎉 If you're at #EMNLP2026 in Budapest 🇭🇺, don't miss this one.
020
Reposted by Institute of Formal and Applied Linguistics
Jana Straková @janastrakova.bsky.social · 13/07/2026
We're releasing a new multilingual model for NameTag 3: nametag3-multilingual-260521. It achieves state-of-the-art performance on 33 test datasets across 23 languages. Try the demo: lindat.mff.cuni.cz/services/nam... Documentation: ufal.mff.cuni.cz/nametag/3
lindat.mff.cuni.cz
071
Institute of Formal and Applied Linguistics @ufal.mff.cuni.cz · 09/07/2026
Nalin Kumar & @tuetschek.bsky.social presented Modular Monolingual Adaptation using Pretrained Language Models aclanthology.org/2026.acl-ind... with their tricks for saving parameters and improving performance when fine-tuning pretrained LMs for low-resource languages.
032
Institute of Formal and Applied Linguistics @ufal.mff.cuni.cz · 08/07/2026
What are Universal Dependencies? 🌍 UD is a framework for annotating grammar consistently across languages, now covering 100+ languages. It's used across multilingual NLP: parsing, linguistic typology studies, and model interpretability studies. Explore: universaldependencies.org
universaldependencies.org
Universal Dependencies
001
Institute of Formal and Applied Linguistics @ufal.mff.cuni.cz · 08/07/2026
🎉 Congratulations to our colleague Dan Zeman, who together with Marie-Catherine de Marneffe, @chrmanning.bsky.social & Joakim Nivre has won the ACL 2026 Computational Linguistics High Impact Paper Award for "Universal Dependencies"! 👏 aclanthology.org/2021.cl-2.11
1101
Institute of Formal and Applied Linguistics @ufal.mff.cuni.cz · 08/07/2026
Reasoning Gets Harder for LLMs Inside A Dialogue aclanthology.org/2026.acl-lon... by @ivankartac.bsky.social, M. Lango & @tuetschek.bsky.social Same reasoning tasks, isolated vs. inside a dialogue: 9 LLMs 📉 consistently do worse once conversation, roles & tool-use requirements enter the picture.
022
Institute of Formal and Applied Linguistics @ufal.mff.cuni.cz · 07/07/2026
Semantic-pragmatic Annotations in the Prague Dependency Treebank aclanthology.org/2026.finding... @mariemikulova.bsky.social , E. Hajicova, J. Mirovsky, A. Nedoluzhko, M. Novak, P. Synkova, J. Stepanek, B. Stepankova, @hajicjan.bsky.social A 3M+ token corpus, fully annotated & freely available.
020
Institute of Formal and Applied Linguistics @ufal.mff.cuni.cz · 07/07/2026
Thesis Proposal by Kristýna Onderková: Intentional Inference for Insight Generation aclanthology.org/2026.acl-srw... Why do LLM insights feel shallow? Often, it is unstated assumptions that fill gaps in the task. This thesis makes models surface them & push toward deeper, more trustworthy reasoning.
021
Institute of Formal and Applied Linguistics @ufal.mff.cuni.cz · 07/07/2026
Thesis proposal by Honza Bronec (supervised by @jindrahelcl.bsky.social) 👉 Targeted and Unified Cross-Lingual Unlearning from Multilingual Language Models aclanthology.org/2026.acl-srw... Aims to unify cross-lingual unlearning benchmarks and edit only the layers that store the knowledge.
010
Institute of Formal and Applied Linguistics @ufal.mff.cuni.cz · 07/07/2026
LLM as a Meta-Judge: Synthetic Data for NLP Evaluation Metric Validation aclanthology.org/2026.acl-srw... by Lukáš Eigler, @jlibovicky.bsky.social & David Hurych Rankings from synthetic LLM-generated data almost perfectly match real human judgments. 🤖⚖️ at Student Research Workshop
010
Institute of Formal and Applied Linguistics @ufal.mff.cuni.cz · 07/07/2026
One-step Nonautoregressive Natural Language Generation with Shortcut Flow Matching Models by Jędrzej Warczyński, Mateusz Lango & @tuetschek.bsky.social aclanthology.org/2026.acl-sho... One-step generation that actually works: BLEU more than doubles vs classic flow matching.
032
Institute of Formal and Applied Linguistics @ufal.mff.cuni.cz · 06/07/2026
Dušan Variš presents joint work w/ @abyste.bsky.social & @jlibovicky.bsky.social: TokCollate: A Comprehensive Tool for Tokenizer Evaluation and Visualization across Languages aclanthology.org/2026.acl-dem... A dashboard for your tokenizer 👀 See which languages you left behind. Open-source, MIT.
030
Institute of Formal and Applied Linguistics @ufal.mff.cuni.cz · 05/07/2026
#ACL2026 continues with the main conference (main + findings + demos + industry + SRW), and @ufal.mff.cuni.cz folks will present 8️⃣ papers. Stop by and check with our colleagues. All times in PDT.
882
Institute of Formal and Applied Linguistics @ufal.mff.cuni.cz · 05/07/2026
UFAL-CUNI at SemEval-2026 Task 11: An Efficient Modular Neuro-symbolic Method for Syllogistic Reasoning by @ivankartac.bsky.social, Honza Bronec, Kristýna Onderková, @zdenekkasner.cz, Mateusz Lango & @tuetschek.bsky.social aclanthology.org/2026.semeval... 💪 4B model + prover beats zero-shot LLMs
051
Institute of Formal and Applied Linguistics @ufal.mff.cuni.cz · 05/07/2026
Lucie Poláková presents: Challenges in Machine Translation of Interactive Multimodal Exercises aclanthology.org/2026.bea-1.1... MT isn't as solved as it seems: terminological, multimodal, XML-encoded exercises bring new challenges, incl. LLMs 'fixing' counterfactuals.
020
Institute of Formal and Applied Linguistics @ufal.mff.cuni.cz · 04/07/2026
A Pocket Offline Model for Simultaneous Speech Translation as CUNI Submission to IWSLT2026 aclanthology.org/2026.iwslt-1.22 AlignAtt4LLM: Fast AlignAtt for Decoder-Only LLMs @ IWSLT2026 Simultaneous Speech Translation Task aclanthology.org/2026.iwslt-1.32 presented simultaneously by Dominik Macháček
030
Institute of Formal and Applied Linguistics @ufal.mff.cuni.cz · 04/07/2026
@straka-milan.bsky.social winning the Unconstrained track in this shared task with his system CorPipe aclanthology.org/2026.codi-1.27 CorPipe 26 tops the unconstrained track by 9.5 points and the LLM track by 2.8, with ablations on model size, empty node prediction & zero-shot transfer.
000
Institute of Formal and Applied Linguistics @ufal.mff.cuni.cz · 04/07/2026
Michal Novák presented "Findings of the Fifth Shared Task on Multilingual Coreference Resolution: Expanding Datasets for Long-Range Entities" on the CODI-CRAC workshop, the CRAC26 shared task overview. CorefUD 1.4: 27 datasets, 19 languages. #NLProc aclanthology.org/2026.codi-1....
001
Institute of Formal and Applied Linguistics @ufal.mff.cuni.cz · 02/07/2026
Congrats! Stop by to chat with @ivankartac.bsky.social at #ACL2026 in San Diego 🇺🇸 He's presenting 2 papers: A modular neuro-symbolic system pairing small LLMs with a theorem prover for syllogistic reasoning, and BOULDER, a new benchmark showing LLM reasoning drops sharply inside multi-turn dialogue.
030
Institute of Formal and Applied Linguistics @ufal.mff.cuni.cz · 01/07/2026
#ACL2026 starts with two days of tutorials and workshops. @ufal.mff.cuni.cz folks will be there and present their work, here's our schedule, so don't miss our presentations 👇
561
Institute of Formal and Applied Linguistics @ufal.mff.cuni.cz · 01/07/2026
If you are at #ACL2026, don't miss @mariemikulova.bsky.social et al. presenting a semantic-pragmatic annotation layer for the 3M+ token Prague Dependency Treebank aclanthology.org/2026.finding...
aclanthology.org
020
Reposted by Institute of Formal and Applied Linguistics
Open Euro LLM @openeurollm.bsky.social · 15/06/2026
Input, more input 🤖⚡ Just like Jonny 5 in Short Circuit, our baby model is reading every single token from its pretraining dataset. So far: 10 trillion tokens, 36 languages + code & math as their own "languages" 📚🌍💻 We’re tracking progress & sharing it openly 👇 (1/2)
static.klipy.com
Ally Sheedy and Johnny 5 in Short Circuit
ALT: Ally Sheedy and Johnny 5 in Short Circuit
191
Reposted by Institute of Formal and Applied Linguistics
Tokenization Workshop (TokShop) @COLM2026 @tokshop.bsky.social · 23/06/2026
📣 TokShop 2026 deadline extended! 🗓️ New submission deadline: Friday, June 26, 2026 (AoE) Research papers (up to 9 pages) and extended abstracts (up to 2 pages) are welcome. Submit: openreview.net/group?id=col... More info: tokenization-workshop.github.io
034
Reposted by Institute of Formal and Applied Linguistics
Jindřich Libovický @jlibovicky.bsky.social · 09/06/2026
Lukáš Eigler defended his thesis (co-supervised with David Hurych, @valeoai.bsky.social) 🎉 Congrats! #NLP metric validation needs 🐌💰 human judgment data. Our fix: generate synthetic data for metric validation. ✅ Tested on MT, QA, summarization. To appear #ACL2026 SRW: arxiv.org/abs/2603.09403
arxiv.org
LLM as a Meta-Judge: Synthetic Data for NLP Evaluation Metric Validation
Validating evaluation metrics for NLG typically relies on expensive and time-consuming human annotations, which predominantly exist only for English datasets. We propose LLM as a Meta-Judge, a scalabl...
091
Institute of Formal and Applied Linguistics @ufal.mff.cuni.cz · 23/06/2026
One of a dozen master's theses defended at our institute this season. Thesis season is always a reminder of the talented students passing through here. 🎓
000
Reposted by Institute of Formal and Applied Linguistics
Tokenization Workshop (TokShop) @COLM2026 @tokshop.bsky.social · 22/05/2026
Announcing First Call for Papers: Second Tokenization Workshop 🔡 📣 ▶️ Non-archival submissions of two types: Research papers (up to 9 pages) ▶️ Extended abstracts (up to 2 pages) Submission deadline June 23, 2026 (AoE) Acceptance notification on July 24, 2026 (AoE) tokenization-workshop.github.io
1138
Reposted by Institute of Formal and Applied Linguistics
Marie Mikulová @mariemikulova.bsky.social · 20/05/2026
Great discussions, inspiring talks, and lots of interest around the 🌲Prague Dependency Treebank🌲 at #LREC2026. Hope we helped give PDT some well-deserved visibility there! @ufal.mff.cuni.cz ➡️ lrec.elra.info/lrec2026-mai... ➡️ lrec.elra.info/lrec2026-mai... ➡️ lrec.elra.info/lrec2026-mai...
041
Institute of Formal and Applied Linguistics @ufal.mff.cuni.cz · 20/05/2026
Automatic Suggestions Help Extending Eventive Ontology: A Case Study on SynSemClass by @janastrakova.bsky.social, Eva Fučíková, Zdenka Urešová & @hajicjan.bsky.social lrec.elra.info/lrec2026-mai... Auto-suggestions improve semantic ontology annotation agreement in SynSemClass case study.
071
Institute of Formal and Applied Linguistics @ufal.mff.cuni.cz · 20/05/2026
SEEM-CZ: Annotation and Classification of Epistemic Markers in Czech by Bára Štpánková, Michal Novák, Tomáš Musil, Lucie Poláková lrec.elra.info/lrec2026-mai... Czech epistemic markers dataset (~4,000 uses) with annotations and XLM-RoBERTa classifiers for NLP.
120
Institute of Formal and Applied Linguistics @ufal.mff.cuni.cz · 20/05/2026
DReUD: Discourse Relations in Universal Dependencies by Jiří Mírovský and Pavlína Synková lrec.elra.info/lrec2026-mai... UD-based shallow discourse relation annotation scheme + DReUD parser for Czech 🇨🇿 & English 🇬🇧
110
Institute of Formal and Applied Linguistics @ufal.mff.cuni.cz · 20/05/2026
CzechDocs: A Multiway Parallel Dataset of Formatted Documents for Minority Languages in Czechia 🇨🇿 by Pepa Jon and Ondřej Bojar 📝 lrec.elra.info/lrec2026-mai... 📂 github.com/cepin19/Czec...
110
Institute of Formal and Applied Linguistics @ufal.mff.cuni.cz · 20/05/2026
HotelCheckSpan: A Benchmark Dataset for LLM Faithfulness huggingface.co/datasets/pat... by @patuchen.bsky.social, @tuetschek.bsky.social and @saad.me.uk 🏨 Hotel summary faithfulness benchmark with span-level error labels. lrec.elra.info/lrec2026-mai...
151
Institute of Formal and Applied Linguistics @ufal.mff.cuni.cz · 20/05/2026
#LREC2026 is over. It had huge @ufal.mff.cuni.cz presence 👯🙆‍♂️🧑‍🤝‍🧑🙋🧓🧒 of 2️⃣6️⃣ people and 1️⃣4️⃣ main conference papers. If you missed our presentations, you can still checkout our #NLProc and #CL papers 👇
120
Institute of Formal and Applied Linguistics @ufal.mff.cuni.cz · 15/05/2026
Final day of the main conference with 4️⃣ @ufal.mff.cuni.cz contributions! If you're around, check our posters! 👀
020
Reposted by Institute of Formal and Applied Linguistics
Tokenization Workshop (TokShop) @COLM2026 @tokshop.bsky.social · 14/05/2026
TokShop will be at #COLM2026! 🗓️ October 9th, 2026 📍 San Francisco, USA More details and a call for papers coming soon.
0118
Institute of Formal and Applied Linguistics @ufal.mff.cuni.cz · 14/05/2026
On Thursday, we continue with 4️⃣ more paper presentations 🤩
130
Reposted by Institute of Formal and Applied Linguistics
Open Euro LLM @openeurollm.bsky.social · 13/05/2026
All ready to share information about #OpenEuroLLM with the #LREC2026 crowd. Let's talk data, infra, evals and open multilingual LLM models together! Come to booth #5 at the poster area 1, Elyxir Building. #multingualLLMs #openLLMs #diverseLLMs #safeLLMs
063
Institute of Formal and Applied Linguistics @ufal.mff.cuni.cz · 12/05/2026
#LREC2026 🌴🇪🇸 continues with day 1 of the main conference. @ufal.mff.cuni.cz folks will present 5 papers: From annotating medieval manuscripts to auditing LLM hallucinations: Prague NLP is everywhere 🚀
163
Institute of Formal and Applied Linguistics @ufal.mff.cuni.cz · 12/05/2026
Proud to have this work at #LREC2026! 🎉 Fingers crossed and good luck with the poster on Thursday! 🤞
020
Institute of Formal and Applied Linguistics @ufal.mff.cuni.cz · 12/05/2026
Šárka Zikanová et al. Introducing corpora Hlava Cor and Hlava AD Humans disagree on text meaning just as much as models do — and on the same hard cases 🤝😬 ufal.mff.cuni.cz/hvar/hlava-cor ufal.mff.cuni.cz/hvar/hlava-ad
020
Institute of Formal and Applied Linguistics @ufal.mff.cuni.cz · 12/05/2026
Meaning Annotation Experience. Tribute to Petr Sgall (Jan Štěpánek et al.) Spatial semantics: where philosophy meets spreadsheets. Multiple annotators, 3M Czech tokens, and one humbling conclusion — meaning is messy and labels lie. Thanks Petr Sgall for the theory, sorry for the disagreements. 📍
121
Institute of Formal and Applied Linguistics @ufal.mff.cuni.cz · 12/05/2026
Panel dicussion The Role of Symbolic Representations in the Era of LLMs @jepusto.bsky.social, @bonverbial.bsky.social, Louise McNally, Susan Windisch and our @hajicjan.bsky.social
110
Institute of Formal and Applied Linguistics @ufal.mff.cuni.cz · 12/05/2026
First Shared Task on UMR Parsing (Overview by Dan Zeman and Jan Štěpánek et al.) ufal.mff.cuni.cz/umr-parsing First shared task on UMR parsing: 7 languages from 4 families. Turns out cross-linguistic semantic graphs are hard — especially when you can't even see the training data. 📊
110
Institute of Formal and Applied Linguistics @ufal.mff.cuni.cz · 12/05/2026
Minoo Nassajian et al. Named Entity Recognition for Persian Literary Text Even the best NER tools get lost in the desert when faced with literary text — turns out training only on news doesn't prepare you for The Little Prince. 🌹🐍 huggingface.co/mansoorhamid...
110
Institute of Formal and Applied Linguistics @ufal.mff.cuni.cz · 12/05/2026
Hana Hledíková presenting Deverbal nouns in UMR: link event/result/agent nouns to their verb, skip location nouns. Simple rule, big consistency win.
110
Institute of Formal and Applied Linguistics @ufal.mff.cuni.cz · 11/05/2026
Congrats to our student Aleš Manuel Papáček to his very first paper made of bachelor thesis! 👉 Never-before-computationally-studied problem: predicting the etymological origin of individual morphs in Czech words (native vs. borrowed, and from which language). lrec-conf.org/proceedings/...
130
Institute of Formal and Applied Linguistics @ufal.mff.cuni.cz · 11/05/2026
The second day of workshops goes on with two more presentations...
010
Institute of Formal and Applied Linguistics @ufal.mff.cuni.cz · 11/05/2026
🔸 🔶 Petr Sgall was a pioneer who saw the future of NLP long before the age of LLMs 🔶 🔸 As we mark the 100th anniversary of his birth, Eva Hajičová and Jarmila Panevová have shared a brilliant tribute to his life and work. 🔗 Read the full tribute here: www.matfyz.cz/clanky/ste-v...
matfyz.cz
Sté výročí průkopníka počítačové lingvistiky prof. Petra Sgalla
V květnu tohoto roku si připomínáme 100 let od narození emeritního profesora Univerzity Karlovy a dlouholetého pracovníka Matematicko-fyzikální fakulty UK PhDr. Petra Sgalla, Dr. h. c. mult. (27. květ...
072