Sign in

Thennal D K

@thennal.bsky.social
41 followers 397 following 31 posts

nlp researcher, fan of fungi any/all, english/മലയാളം/日本語

PostsRepliesMedia
Reposted by Thennal D K
Maria Antoniak @mariaa.bsky.social · 03/10/2026
Maybe worth saying aloud that I do worry about safety risks (hacking, bio, an “out of control” agent). I just also see those risks as deeply entangled with money, politics, personalities, regulation, such that I doubt we can “solve safety” if we don’t first address foundational issues.
3688
Reposted by Thennal D K
Mike Rugnetta @mikerugnetta.com · 30/09/2026
“lately, she’s been filming AI how-to content on topics like how to get an AI assistant to buy you groceries so you can make Goop Kitchen’s Brentwood Chinese chicken salad at home”
6362
Reposted by Thennal D K
Colin @colin-fraser.net · 16/09/2026
LLMs won’t wipe out humanity because they just don’t have that dog in them.
825228
Reposted by Thennal D K
Colin @colin-fraser.net · 14/09/2026
It rules that in 2026 you can steal $600,000 just by asking politely
69711
Reposted by Thennal D K
Colin @colin-fraser.net · 08/09/2026
many of you are spending far too much energy on finite time blowups when you should be striving for a finite time glow up
14611
Reposted by Thennal D K
leon @leyawn.bsky.social · 29/08/2026
post by leyawn

jester, roast this peasant child for saying i have no clothes on. more vulgar, jester. use the forbidden words
95150032772
Reposted by Thennal D K
Colin @colin-fraser.net · 27/08/2026
I also continue to believe that my mental model is the best one, which does not conceptualize the “agents” as discrete entities, but rather as the protagonists in a story that the LLM has been trained to write about a cool little guy who goes on hacker adventures.
434462
Reposted by Thennal D K
Eryk Salvaggio @eryk.bsky.social · 05/08/2026
Court rules “Mambo No. 6” must be written by a human
0379
Reposted by Thennal D K
Alane Suhr @suhr.bsky.social · 03/08/2026
Wittgenstein would love this game en.wikipedia.org/wiki/Mad_Libs
en.wikipedia.org
Mad Libs - Wikipedia
021
Reposted by Thennal D K
Eryk Salvaggio @eryk.bsky.social · 01/08/2026
It’s not becoming a little guy but it is become a weird language calculator and we’d be better off if we dropped the “little dude’s arriving soon” shit & think about wtf a language calculator is instead
3284
Thennal D K @thennal.bsky.social · 16/05/2026
every time I see a violin plot in a paper I take compounding psychic damage
000
Reposted by Thennal D K
Colin @colin-fraser.net · 15/05/2026
I still always think about the early AI paper (circa 2023?) paper where they get people to guess whether a poem is by a famous poet or ChatGPT and the people perform worse than chance and when you dig a bit deeper into the data you find it’s because they simply hate the real poems
711915
Reposted by Thennal D K
Eryk Salvaggio @eryk.bsky.social · 13/05/2026
Moving beyond training data explanations doesn’t cede ground, it’s actually a less mystically charged combo of code & ranking algos, while data-at-scale has a culty “accumulation of information creates thinking” motivation that is low-key reinforced when credited with increasingly complex language
091
Reposted by Thennal D K
eternalist @eternalism-when.bsky.social · 09/05/2026
4583123
Reposted by Thennal D K
Hans Hatzel @hanshatzel.bsky.social · 10/04/2026
Joint translation and label projection does not hurt translation performance! Our paper "Just Use XML: Revisiting Joint Translation and Label Projection" (accepted to ACL Findings) rejects the idea that label projection for span annotations should be performed as a secondary step after translation.
arxiv.org
Just Use XML: Revisiting Joint Translation and Label Projection
Label projection is an effective technique for cross-lingual transfer, extending span-annotated datasets from a high-resource language to low-resource ones. Most approaches perform label projection as a separate step after machine translation, and prior work that combines the two reports degraded translation quality. We re-evaluate this claim with LabelPigeon, a novel framework that jointly performs translation and label projection via XML tags. We design a direct evaluation scheme for label projection, and find that LabelPigeon outperforms baselines and actively improves translation quality in 11 languages. We further assess translation quality across 203 languages and varying annotation complexity, finding consistent improvement attributed to additional fine-tuning. Finally, across 27 languages and three downstream tasks, we report substantial gains in cross-lingual transfer over comparable work, up to +39.9 F1 on NER. Overall, our results demonstrate that XML-tagged label projection provides effective and efficient label transfer without compromising translation quality.
231
Thennal D K @thennal.bsky.social · 04/03/2026
dead internet theory but it's just agents emailing eachother
000
Reposted by Thennal D K
Carl T. Bergstrom @carlbergstrom.com · 11/11/2025
What do you do after you’re done jumping the shark? Whatever it is, Nature Careers is all in.
Nature
CAREER COLUMN
11 November 2025
I have Einstein, Bohr and Feynman in my pocket
Grappling with difficulties in your career? Try asking an Al-powered advisory panel of experts, suggests Carsten Lund Pedersen.The make-up of my advisory board often changes, and has included an eclectic mix.
Besides Bohr, Feynman and Einstein, I've also tapped microbiologist Alexander Fleming, poet Piet Hein and anti-apartheid activist Nelson Mandela. Sometimes I include experts from my specific disciplines of Al and marketing; other times, l 'invite' artist Pablo Picasso or architect Bjarke Ingels for a completely different perspective. But whatever the board's composition, I typically retain a core group of three seminal scientists.
4132751
Thennal D K @thennal.bsky.social · 08/04/2025
Stop using Word Error Rate! The lovely @sthirkal.bsky.social and I made a poster for my recent NAACL 2025 Findings paper, highlighting the issues in multilingual ASR evaluation and proposing viable alternatives.
A poster describing the content of our paper.
000
Reposted by Thennal D K
Arianna Bisazza @arianna-bis.bsky.social · 08/04/2025
Modern LLMs "speak" hundreds of languages... but do they really? Multilinguality claims are often based on downstream tasks like QA & MT, while *formal* linguistic competence remains hard to gauge in lots of languages Meet MultiBLiMP! (joint work w/ @jumelet.bsky.social & @weissweiler.bsky.social)
2216
Reposted by Thennal D K
Shana Gadarian @sgadarian.bsky.social · 08/04/2025
International students will stop coming to American universities if their visas are going to be at risk. This will make our intellectual community poorer and also make tuition more expensive for domestic students.
7590163
Reposted by Thennal D K
The Onion @theonion.com · 05/02/2025
FBI Uncovers Al-Qaeda Plot To Just Sit Back And Enjoy Collapse Of United States
theonion.com
FBI Uncovers Al-Qaeda Plot To Just Sit Back And Enjoy Collapse Of United States
WASHINGTON—Putting the nation on alert against what it has described as a “highly credible terrorist threat,” the FBI announced today that it has uncovered a plot by members of al-Qaeda to sit back an...
10646837415825
Reposted by Thennal D K
Alane Suhr @suhr.bsky.social · 01/02/2025
Most of my colleagues are shocked when I bring up these comments to them. Being in academia doesn't mean we are protected, or even simply ignored by this. We and our institutions are being actively attacked by policies like funding freezes, and I don't see why it wouldn't get worse than that.
0222
Reposted by Thennal D K
Peli Grietzer @peligrietzer.bsky.social · 01/02/2025
We need a name for the 'paradox' where frontier LLMs are useful in domain x if your expertise in domain x go much deeper than theirs and harmful otherwise
9654
Reposted by Thennal D K
Jonathan Peelle @jpeelle.bsky.social · 31/01/2025
My university (Northeastern) got rid of their page on Diversity, Equity, and Inclusion, and redirected the URL to a page on "belonging". It's not clear to me how a private university having a page on their values violates the law? web.archive.org/web/20250124...
web.archive.org
Office of Belonging – Belonging at Northeastern
86911
Reposted by Thennal D K
Abhilasha Ravichander @lasha.bsky.social · 31/01/2025
We are launching HALoGEN💡, a way to systematically study *when* and *why* LLMs still hallucinate. New work w/ Shrusti Ghela*, David Wadden, and Yejin Choi 💫 📝 Paper: arxiv.org/abs/2501.08292 🚀 Code/Data: github.com/AbhilashaRav... 🌐 Website: halogen-hallucinations.github.io 🧵 [1/n]
3348
Reposted by Thennal D K
George Takei @georgetakei.bsky.social · 31/01/2025
This is what the government did with 120K+ Japanese Americans in 1942. I know. I was there in those camps.
34539889731613
Thennal D K @thennal.bsky.social · 23/01/2025
Happy to share that this paper has now been accepted at NAACL 2025!
000
Reposted by Thennal D K
Jeremy Howard @howard.fm · 19/12/2024
I'll get straight to the point. We trained 2 new models. Like BERT, but modern. ModernBERT. Not some hypey GenAI thing, but a proper workhorse model, for retrieval, classification, etc. Real practical stuff. It's much faster, more accurate, longer context, and more useful. 🧵
19620147
Reposted by Thennal D K
Colin @colin-fraser.net · 19/12/2024
Here's why "alignment research" when it comes to LLMs is a big mess, as I see it. Claude is not a real guy. Claude is a character in the stories that an LLM has been programmed to write. Just to give it a distinct name, let's call the LLM "the Shoggoth".
929479
Reposted by Thennal D K
Sam Power @spmontecarlo.bsky.social · 27/12/2024
I saw recently that this had been cited in some lecture notes, so I decided to try to preserve it for posterity (since the original tweet had been deleted by me in the interim).
1717
Reposted by Thennal D K
Adam Rogers @jetjocko.bsky.social · 26/12/2024
At risk of repeating myself: The Luddites weren’t against technology. They were against getting put out of work by a technology that did a version of their job faster but worse, in the service of increasing profits for their bosses.
8461882038
Reposted by Thennal D K
Jameel Jaffer @jameeljaffer.bsky.social · 13/12/2024
Is artificial intelligence supercharging the problem of political misinformation? In a word: no. New from @randomwalker.bsky.social & @sayash.bsky.social at @knightcolumbia.org. knightcolumbia.org/blog/we-look...
knightcolumbia.org
We Looked at 78 Election Deepfakes. Political Misinformation Is Not an AI Problem.
4248
Reposted by Thennal D K
Irene Chen @irenetrampoline.bsky.social · 20/12/2024
How do disparities in healthcare access affect ML models? 💰📉🧐 We found that low access to care -> worse EHR data quality -> worse ML performance in a dataset of 134k patients. Work with Anna Zink (on the faculty job market rn!) + Hongzhou Luan, presented at #ML4H2024
Bar chart of different barriers to healthcare
14014
Reposted by Thennal D K
Ted Underwood @tedunderwood.com · 25/12/2024
I’m a Bourdieusian, when I give a movie 7/10 it means it goes here
The image is titled “Figure 2: French literary field in the second half of the 19th century.” It appears to be a diagram representing various forms of literature in relation to their audience and degree of consecration.

Key features of the diagram include:
	•	Axes:
	•	The vertical axis measures autonomy (from poor, no economic profit at the bottom, to autonomy, no audience at the top).
	•	The horizontal axis ranges from “Left” (unknown, poor, young) to “Right” (institutionalized consecration, bourgeois audience).
	•	Quadrants:
	•	Top-left: High autonomy with intellectual audience (“Art for Art’s Sake”), including Symbolists (Mallarmé) and Decadents (Verlaine).
	•	Top-right: High degree of consecration with bourgeois audience, featuring psychological novels, “society novels,” and Parnassians.
	•	Bottom-left: Low consecration, no audience, labeled “Bohemia,” including little reviews and younger, unknown works.
	•	Bottom-right: Low consecration and a mass audience, including popular novels, journalism, and industrial art like vaudeville.
	•	Connections:
	•	Arrows link various literary movements or works, such as the influence of naturalist novels (Zola) on the broader literary field.

Annotations like “10” and “7” are added on the diagram but their purpose is unclear from the image alone. If you need help interpreting it further, feel free to provide more context!
4638
Reposted by Thennal D K
Hypervisible @hypervisible.blacksky.app · 25/12/2024
I’m dreaming of a slop Xmas…
404media.co
Merry Slopmas!
AI-generated Christmas classics that dwell in the uncanny valley are giving listeners the creeps.
911534
Reposted by Thennal D K
The Washington Post @washingtonpost.com · 23/12/2024
As data centers cause America to run out of power, Arizona is emblematic of the sacrifices ordinary people face to fuel the boom.
washingtonpost.com
In the shadows of Arizona’s data center boom, thousands live without power
As data centers drain America’s power grids, a fierce battle is being waged for electricity. On Navajo Nation land, many are on the losing end.
21388141
Thennal D K @thennal.bsky.social · 16/12/2024
honey you didnt close the fridge door again im adjusting my priors on the divorce
000
Thennal D K @thennal.bsky.social · 06/12/2024
extremely funny sequence of events that's emblematic of issues in the field in general. and ofc new work will ignore both these datasets and just use the original MMLU anyway
000
Reposted by Thennal D K
lukelukeluke @lukelukeluke.bsky.social · 02/12/2024
Here are some nice mushrooms
A group of rich brown Ischnoderma resinosum mushrooms is topped with a pile of snow where they grow off a log, their cream undersides flowing downward
33101948
Reposted by Thennal D K
Laura @lauraruis.bsky.social · 20/11/2024
How do LLMs learn to reason from data? Are they ~retrieving the answers from parametric knowledge🦜? In our new preprint, we look at the pretraining data and find evidence against this: Procedural knowledge in pretraining drives LLM reasoning ⚙️🔢 🧵⬇️
36850139
Thennal D K @thennal.bsky.social · 01/12/2024
I believe the natural language processing as a field has severe issue with multilingual evaluation. This paper focuses specifically on automatic speech recognition evaluation, and advocates for an a more inclusive, more accurate evaluation scheme than the one currently employed by e.g. OpenAI.
100
Thennal D K @thennal.bsky.social · 01/12/2024
some of my recent work in #nlp: Large Language Models Are Overparameterized Text Encoders arxiv.org/abs/2410.14578
arxiv.org
Large Language Models Are Overparameterized Text Encoders
Large language models (LLMs) demonstrate strong performance as text embedding models when finetuned with supervised contrastive training. However, their large size balloons inference time and memory r...
280
Reposted by Thennal D K
Linda Bellissima @lindabellissima.bsky.social · 30/11/2024
Good Saturday morning all #photography #Nature #fungus #landscape #mycelium
32164577
Reposted by Thennal D K
lukelukeluke @lukelukeluke.bsky.social · 30/11/2024
Here is a nice mushroom
A single Mycena mushroom has its cap illuminated by autumn sunlight where it grows on a dark log against yellow forest undergrowth
794527752345