Sign in

Marzena Karpinska

@markar.bsky.social
4K followers 978 following 121 posts

#nlp researcher interested in evaluation including: multilingual models, long-form input/output, processing/generation of creative texts previous: postdoc @ umass_nlp phd from utokyo marzenakrp.github.io

PostsRepliesMedia
Marzena Karpinska @markar.bsky.social · 06/10/2026
We analyze some trends in AI ideas and Human ideas to understand what the detector is picking up on. We see that: - AI overstates claims beyond what is provided - Makes statements that 'appear' specific until you try to actually apply them.
1121
Marzena Karpinska @markar.bsky.social · 06/10/2026
IdeaLens can also detect AI ideas in languages other than English despite never seeing them during training. This also works in low-resource languages such as Marathi or Khmer. The model flags 94- 96% of AI-ideated documents & only 0.6-0.8% of AI-written documents.
120
Marzena Karpinska @markar.bsky.social · 06/10/2026
IdeaLens can also successfully detect human writing based on AI ideas.
110
Marzena Karpinska @markar.bsky.social · 06/10/2026
IdeaLens can detect varying levels of human input that went into the writing. For documents generated from a very detailed human outline, IdeaLens flags only 6.8% as AI, compared to ProseLens's 86.5% and Pangram's 92.4%.
110
Marzena Karpinska @markar.bsky.social · 06/10/2026
On conventional AI detection benchmarks, where prose and ideas come from the same source, IdeaLens correctly flags 91.1% as AI at 0.6% FPR, outperforming all open-source detectors. ProseLens does even better, flagging 98.9% at 0.7% FPR.
110
Marzena Karpinska @markar.bsky.social · 06/10/2026
Traditional detectors are very accurate when #prose and #ideas come from the same source. But #IdeaLens is the only detector that can detect AI ideas hidden behind human prose.
120
Marzena Karpinska @markar.bsky.social · 06/10/2026
We represent documents as outlines of their ideas and train #IdeaLens to fit Pangram 3 labels after paraphrasing the outlines to remove stylistic cues from the source. As a control, we train #ProseLens directly on the raw documents.
110
Marzena Karpinska @markar.bsky.social · 06/10/2026
Can you tell whether the #ideas in a piece of writing came from a human or an AI? As AI writing becomes common, policies on AI use focus on who came up with the ideas in the text rather than who wrote it. AI detectors answer the latter, not the former. We release #IdeaLens, a detector of AI ideas.
First page of a research paper: IDEALENS: DETECTING AI IDEAS IN LONG-FORM WRITING
24114
Marzena Karpinska @markar.bsky.social · 04/09/2026
My first PhD student (Weidong Zhang, yet to come to blsky) brought us new lab's gadgets and sweets 😍😍😍😍 I'm so happy! Thank you 🥰
020
Marzena Karpinska @markar.bsky.social · 21/07/2026
this checkbox at #arr really seems like a get-out-of-jail-free card; if we really need to allow for lang edits, why not require the reviewers to submit their pre-GPTed draft along with the edited one? 😭😭😭
1131
Marzena Karpinska @markar.bsky.social · 07/07/2026
5 years later MTurk is on its way out. A lot has improved in open-ended gen eval, but it still suffers from underreporting, lack of statistical analysis, and sometimes sloppy design. Perhaps authors should always do their own tasks to understand the implications of how they designed evals.
030
Marzena Karpinska @markar.bsky.social · 07/07/2026
A great lesson learned from #ACL2026 test of time award winner: sometimes it just takes time for people to appreciate your work
081
Marzena Karpinska @markar.bsky.social · 07/07/2026
Great message from @barbaraplank.bsky.social unifying CL & NLP
031
Marzena Karpinska @markar.bsky.social · 06/07/2026
Only 17% of May #ARR authors are qualified for reviewing, and almost 40% of papers have no qualified people?!?!
2304
Marzena Karpinska @markar.bsky.social · 06/07/2026
Spotted by @jennarussell.bsky.social at #ACL2026 A sober reminder to check the output when using Claude et al to generate your presentation. Context, the user likely prompted for 8 min long presentation since that was the limit. The model, just added this silly "8 minutes" to all slides.
010
Marzena Karpinska @markar.bsky.social · 26/06/2026
A lot of people focus on AI-written fiction, but what about AI literary translation? 📚 We find that AI translation can be readable.. ‼️BUT it also flattens characters' voices 🧙and is less immersive 🫣 than published human translations. Below is my favorite quote from a reader 📖 lait.cs.sfu.ca
0216
Marzena Karpinska @markar.bsky.social · 23/05/2026
One more thing @tuhinchakr.bsky.social 's post reminded me of... people tend to rationalize and see things not there. We saw it already in GPT-2 stories - we *expect* things to *mean* something, so we tend to see things that are not there... (link to this old paper: aclanthology.org/2021.emnlp-m...)
032
Marzena Karpinska @markar.bsky.social · 16/01/2026
Now is probably a good time to share that I left my job at @microsoft.com (will forever miss this team) and moved to Vancouver, Canada, where I'm starting my lab as an assistant professor at the gorgeous @sfu.ca 🏔️ I'm looking to hire 1-2 students starting in Fall 2026. Details in 🧵
1121
Marzena Karpinska @markar.bsky.social · 18/10/2025
I'm not sure why people lost the ability to do related work properly but if you absolutely need to use AI at least proofread it? (And they most likely edited with ai) www.pangram.com/history/01bf...
050
Marzena Karpinska @markar.bsky.social · 07/10/2025
Come to talk with us today about the evaluation of long form multilingual generation at the second poster session #COLM2025 📍4:30–6:30 PM / Room 710 – Poster #8
062
Marzena Karpinska @markar.bsky.social · 06/10/2025
Off to #COLM fake Fuji looks really good today. 本物は下からしか見たことがないが、今日は少なくとも偽物が上から見えて嬉しい。
060
Marzena Karpinska @markar.bsky.social · 06/10/2025
I feel like it was worth waking up early
040
Marzena Karpinska @markar.bsky.social · 20/08/2025
Happy to see this work accepted to #EMNLP2025! 🎉🎉🎉
0121
Marzena Karpinska @markar.bsky.social · 08/08/2025
GPT-5 lands first place on NoCha, our long-context book understanding benchmark. That said, this is a tiny improvement (~1%) over o1-preview, which was released almost one year ago. Have long-context models hit a wall? Accuracy of human readers is >97%... Long way to go!
Screenshot of benchmark with gpt-5 on top with 68.46% accuracy.
1186
Marzena Karpinska @markar.bsky.social · 07/04/2025
We have updated #nocha (long-context benchmark measuring how well models process book-length narratives) with #Llama4 Scout. Sadly, the performance was below the random level, much lower than the reported model's performance on a retrieval task (needle in the haystack). novelchallenge.github.io
161
Marzena Karpinska @markar.bsky.social · 02/04/2025
We have updated #nocha, a leaderboard for reasoning over long-context narratives 📖, with some new models including #Gemini 2.5 Pro which shows massive improvements over the previous version! Congrats to #Gemini team 🪄 🧙 Check 🔗 novelchallenge.github.io for details :)
Leaderboard showing performance of language models on claim verification task over book-length input. o1-preview is the best model with 67.36% accuracy followed by Gemini 2.5 Pro with 64.17% accuracy.
0114
Marzena Karpinska @markar.bsky.social · 21/02/2025
An absolutely awesome lineup of language pairs for the 20th iteration of WMT 🍾🎉
060
Marzena Karpinska @markar.bsky.social · 28/01/2025
No matter how we tried to modify #LLM generated text (paraphrasing, humanization), people who frequently use LLMs for writing are consistently good at detecting model-generated text, though they change cues they rely on! Congrats @jennarussell.bsky.social on first paper!
192
Marzena Karpinska @markar.bsky.social · 29/12/2024
We've added #o1 and #Llama 3.3 70B to the #Nocha leaderboard for long-context narrative reasoning! Surprisingly, o1 performs worse than o1-preview, and Llama 3.3 70B matches proprietary models like gpt4o-mini & gemini-Flash. Check out our website for more results! More in 🧵
Screenshot of the nocha leaderboard with o1-preview model performing the best at 67.36%
1354
Marzena Karpinska @markar.bsky.social · 11/11/2024
I will be present our paper on LMs performance on long-context reasoning task at #EMNLP2024 (Tue 16:00-17:30; riverfront hall) Come and chat with us! 🧚🦋
2203
Marzena Karpinska @markar.bsky.social · 11/11/2024
I really wanted to run NEW #nocha benchmark claims on #o1 but it won't behave 😠 - 6k reasoning tokens is often not enough to get an ans and more means being able to process only short books - OpenAI adds sth to the prompt: ~8k extra tokens-> less room for book+reason+generation!
Image showing prompt token count as per the tokenizer (tiktoken) which is 117,609 tokens, and as per what openai API claims it to be, which is 125,385 tokens. There is about 7000 extra tokens added coming from who knows where.
163