Sign in

Rémy Decoupes

@remydec.bsky.social
102 followers 113 following 22 posts

GeoAI (NLP & LLM) and GIS for disaster response and crisis management at UMR TETIS, INRAE remydecoupes.pages-forge.inrae.fr/w…

PostsRepliesMedia
Reposted by Rémy Decoupes
André Loez & podcast Paroles d’histoire @andreloez.bsky.social · 23/09/2026
Unpopular opinion Histoire de France : quelqu'un qui serait capable de réciter dans l'ordre tous les rois / présidents depuis les origines, mais sans connaître l'enquête linguistique de 1863 ou les réalités qu'elle illustre, ne connaît et ne comprend en fait rien à la trajectoire historique du pays
carte tirée du livre d'E Weber sur l'enquête linguistique de 1863 montrant un grand nombre de départements non francophones (sud, ouest, est)
39533196
Reposted by Rémy Decoupes
Cartes @cartes.app · 16/09/2026
🤖 Internet est le nouveau far-west. Et ça explique pourquoi notre service est dégradé depuis plusieurs semaines. Auparavant le contrat était simple, un éditeur autorisait un bot d'indexation à consulter son site car ça lui rapportait de l'audience.
7182140
Reposted by Rémy Decoupes
BlackboxNLP @blackboxnlp.bsky.social · 09/09/2026
🎉 Thrilled to announce the keynote speakers for BlackboxNLP 2026 @emnlpmeeting, with three incredible perspectives on interpretability: 🔍 Ivan Titov 🔍 Sheridan Feucht @sfeucht.bsky.social 🔍 Michael Hahn @m-hahn.bsky.social Join us on October 29th! 🇭🇺
0104
Reposted by Rémy Decoupes
Ai2 @ai2.bsky.social · 26/08/2026
A Thai research team adapted our Dolma data-curation toolkit to build Mangosteen, a 47B-token corpus for Thai LLMs. They used Dolma to filter widely used web datasets into a smaller corpus that improved Thai LLM performance despite using less data. 🧵 buff.ly/fvLSIru
1161
Rémy Decoupes @remydec.bsky.social · 25/08/2026
Training data attribution: capability provenance 🚀
020
Reposted by Rémy Decoupes
Philippe Reka visionscarto.net @reka.visionscarto.net · 06/08/2026
Un jour, un article... un mois au cœur de la plateforme visionscarto.net Jour 6 : L’évolution du français aux États-Unis www.visionscarto.net/francais-eta... par Albert Valdman « dans tas d’pierre-là (dans le tas de pierres) p’tsit berger-là (le petit berger) m’as casser ta tête avec bâton-là »
001
Rémy Decoupes @remydec.bsky.social · 03/08/2026
Hallucination des LLMs et géographie. Au delà du fait que l'emplacement des pays est aléatoire, la visualisation en elle même n'a pas de cohérence. www.lemonde.fr/afrique/arti...
lemonde.fr
Quand le gouvernement américain publie une carte fantaisiste de l’Afrique
Lors d’une conférence internationale sur le sida, au Brésil, fin juillet, le département d’Etat a présenté une carte de l’Afrique où la Côte d’Ivoire s’est retrouvée dans le sud du continent et le Nig...
000
Rémy Decoupes @remydec.bsky.social · 28/07/2026
😭
000
Reposted by Rémy Decoupes
Dr. Serge Zaka @sergezaka.bsky.social · 26/06/2026
+1000% de mortalité chez les poules, +200% chez les porcs, +45% chez les bovins. Plusieurs millions d'animaux d'élevage sont morts sous l'effet de la canicule. 1/13
411204992
Reposted by Rémy Decoupes
Yonatan Belinkov @boknilev.bsky.social · 17/06/2026
Are you wondering if LLM interpretability results generalize, reproduce, etc.? Check out the reproducibility challenge and submit your work reproducing papers in this area: bsky.app/profile/blac...
1156
Reposted by Rémy Decoupes
Ai2 @ai2.bsky.social · 11/06/2026
LLMs are no longer created w/ human data alone. They rely on other models to generate & filter data, evaluate outputs, & guide dev work. So what is a modern LLM built on? Olmo 3 → 89 model + 183 dataset dependencies; Nemotron 3 → 273 + 560 We made ModSleuth to trace this. 🧵
15511
Reposted by Rémy Decoupes
rec @i-am-rec.bsky.social · 02/06/2026
🧮 LREC 2026 proceedings are out, and I just counted the annotation tool citations. 💡 INCEpTION came out on top with 39 citations and 28 of those actually using it! ❤️ Massive thanks to the LREC community for choosing INCEpTION! #lrec2026 #inceptiontap #opensource #textannotation #opensource
linkedin.com
INCEpTION @ LREC 2026
As you might know, I am the maintainer of the INCEpTION open source tool for the semantic and linguistic annotation of text documents. And the LREC conference is a good way to gauge how popular INCEpT...
172
Reposted by Rémy Decoupes
Gabriele Sarti @gsarti.com · 13/05/2026
Excellent survey on causal interpretability by @amuuueller.bsky.social and many BauLab members, don't miss it!
1133
Reposted by Rémy Decoupes
Julien Gossa @juliengossa.cpesr.fr · 05/05/2026
« plus une page est utile à Copilot, plus elle est citée. Plus elle est citée, moins elle est cliquée. L'IA cannibalise ses propres sources. [L'utilisateur] lit la réponse résumée par Copilot et n'a plus besoin de cliquer sur le lien source. »
0125
Reposted by Rémy Decoupes
Ai2 @ai2.bsky.social · 23/04/2026
OlmoEarth Studio now lets you compute & export custom embedding vectors from our OlmoEarth foundation models. Choose your area, time range, encoder, resolution, & imagery sources—Studio delivers a GeoTIFF you can use however you like. 🧵
1266
Rémy Decoupes @remydec.bsky.social · 24/03/2026
🤩
020
Reposted by Rémy Decoupes
Ilker Kesen @ilkerkesen.bsky.social · 17/03/2026
📢I'm organizing a BoF session at #EACL2026 called Tokenization & Beyond, aiming to gather researchers exploring tokenization and alternatives such as byte-level and pixel-based approaches. Sign up using the form if you're interested! #NLProc @eaclmeeting.bsky.social
1119
Reposted by Rémy Decoupes
Patrick Marques @pmarques35.bsky.social · 11/03/2026
[1/8] @obs-des-inegalites.bsky.social Louis Maurin examine comment mesurer les inégalités dans l’espace. Il montre que la géo des inégalités dépend fortement de l’échelle choisie et des méthodes d’observation. Les cartes doivent donc être interprétées avec prudence. #geography #spatialinequality
inegalites.fr
Inégalités entre territoires : de quoi parle-t-on ?
Mesurer les inégalités dans un espace géographique n’est pas aussi évident qu’il y parait. Quelle échelle choisir ? Quels sont les pièges à éviter ? L’analyse de Louis Maurin.
1209
Reposted by Rémy Decoupes
Damien Van Achter (davanac) @davanac.bsky.social · 02/03/2026
L'État français déploie son propre serveur MCP pour connecter les IA aux données publiques. Code source ouvert, contrôle souverain : une infrastructure qui évite la capture par les géants tech tout en respectant les standards ouverts.
da.van.ac
Data.gouv lance son protocole ouvert contre la capture des données publiques
L'État français expérimente MCP, un protocole ouvert pour connecter les IA aux données publiques, avec code source accessible et focus sur la souveraineté.
1159
Reposted by Rémy Decoupes
Alexander Doria @dorialexander.bsky.social · 24/02/2026
New amazing interpretability/SAE work on Baguettotron! Almost surprised how much the high entropy section are actually connected to analytical features: lyramakesmusic.github.io/bread-slicer/
1615
Reposted by Rémy Decoupes
Daniel van Strien @danielvanstrien.bsky.social · 20/02/2026
Llama.cpp joins Hugging Face github.com/ggml-org/lla...
llama.cpp logo + Hugging Face logo
2547
Reposted by Rémy Decoupes
Unité INRAE MaIAGE @inrae-maiage.bsky.social · 20/02/2026
MaIAGE recrute ! Un poste permanent d'IR en TAL est ouvert au concours. Descriptif du poste, contacts et modalités de candidature: jobs.inrae.fr/concours/con...
025
Reposted by Rémy Decoupes
AI Forensics @aiforensics.org · 12/02/2026
🧵 AI bias isn’t just a chatbot problem. It’s increasingly built into our phones 📱 Our new audit finds the core on-device model powering Apple Intelligence is far from neutral.
171
Reposted by Rémy Decoupes
Processor of Natural Languages @processorofnl.bsky.social · 03/02/2026
It's insightful and true because it's in ICLR format
021
Reposted by Rémy Decoupes
Lea Tardieu @leatardieu.bsky.social · 27/01/2026
🗺️ We are offering a position as an INRAE research fellow in spatial economics or quantitative geography focusing on global health approaches in our laboratory - UMR TETIS (Territories, Environment, Remote Sensing, and Spatial Information). 📍 All information can be found here: #OneHealth #Spatial
jobs.inrae.fr
Junior research scientist in quantitative geography or spatial economics applied to regional "One Health" approach
CR26-ACT-1 - You will work at the TETIS research unit in Montpellier, at the Maison de la Télédétection, which also houses the Espace-Dev research unit and companies specializing in remote sensing and spatial information. The lab is a highly multidisciplinary unit developing applied research in the field of spatial information. It offers a welcoming environment for researchers in the social sciences (geography, economics) and encourages the development of innovative approaches by promoting interaction with researchers from other disciplines (remote sensing, data science, spatial modeling). You will join the “Diagnosis and Anticipation through Spatial Modeling and Analysis of Territories” research group and participate in the unit's scientific activities focused on the major challenges of "Preserving biodiversity and strengthening the health of socio-ecosystems“ and ”Characterizing inequalities and promoting socio-spatial justice."You will also benefit from a large network of collaborators in the Montpellier area (e.g., i-site MUSE Exposum, MSH Sud, MEEDIN, Health Ecology networks), the INRAE network (Territoires de Santés seminar, ModStatSAP network) in France, and international networks (PEER Network, Ecosystem Services Partnership, IPBES).Addressing the challenges related to health independently of those related to biodiversity, food, and climate change can lead to ineffective (and even costly) policies. Your research will contribute to the development of quantitative and spatial methodologies that operationalize ‘One health’ and nexus approaches and enable the evaluation of land use policies with a global health perspective. This may involve the concept of « one health territories » (« territoires de santés »), referring to spaces of varying sizes where collective actions are deployed to manage risks in a comprehensive manner. You will study the types of spatial organizations (e.g. landscape structure) that promote or hinder global health, developing methods in the short term to map multiple risks (pathogenesis) and then, in the medium term, analyses of what constitutes global health (salutogenesis). From a geographical perspective, the research focus will be on developing quantitative and spatial analysis methodologies (statistics, modeling) based on multiple data sets (epidemiological, biophysical, and socioeconomic). From an economic perspective, you will evaluate human health-focused development policies versus integrated approaches, addressing issues of urbanization, climate change, and pressure on biodiversity.You will join the TETIS lab (Montpellier) organized into research groups developing innovative approaches to spatial information to support regions facing major social and environmental challenges. This structure is designed to promote cross-disciplinary perspectives and interactions. You will be responsible for conducting research on the operationalization of nexus and One health approaches by developing methodologies and analytical tools that are useful for land use planning and health stakeholders in addressing for instance the following questions: How do spatial effects (agglomeration, displacement) affect global health in a given territory? Which populations are most impacted by land use planning policies? Which land use planning policies (dis)advantage global health?These questions will be developed through case studies involving health and environmental stakeholders (e.g., DGS, DGALN, ARS, etc.). In the short term, your research will focus on pathogenesis approaches, for which data (environmental and epidemiological data from the unit's projects—e.g., MOOD, BEYOND, THEIA Spatial Data Center) are available, in order to assess and map risks to animal health, human health, or plant health in France. In the medium term, your research will focus on salutogenesis approaches, examining forms of land use planning that promote health (e.g., prevention of zoonoses through biodiversity).To conduct your research, you will have privileged access to the spatial modeling tools produced by the unit (e.g., Ocelet, ArboCarto, Urban InVEST, BiodivMapR, FORDEAD, etc.) and will be able to participate in French collaborative projects (e.g., Vulnerability of Populations and Ecosystem Services - VULPES, Agralife - Sustainable Livestock Farming, Vegetation-Ecosystem Services and Arboviruses) and international projects (e.g., Horizon Europe MOSAIC). Within INRAE, the position is linked to the BIOSEFAIR and XRISQUES metaprograms. Finally, you will have access to financial resources in accordance with the unit's operating rules.
023
Reposted by Rémy Decoupes
Antonin Poché @antoninpoche.bsky.social · 20/01/2026
🔥I am super excited for the official release of an open-source library we've been working on for about a year! 🪄interpreto is an interpretability toolbox for HF language models🤗. In both generation and classification! Why do you need it, and for what? 1/8 (links at the end)
1229
Reposted by Rémy Decoupes
Epoch AI @epochai.bsky.social · 13/01/2026
We loved this quick visual rundown of the Frontier Datacenters hub. Thanks to Rowan Cheung for featuring our project! www.youtube.com/shorts/szAW...
youtube.com
This map reveals some of the hidden facts behind AI data centers 👀 #trendingshorts #ai #research
Epoch AI, a nonprofit research group, is using satellite imagery and public records to track the rapid expansion of AI datacenters across the United States.B...
171
Reposted by Rémy Decoupes
AgroParisTech @agroparistech.fr · 13/01/2026
38 participants ont joué à un escape game pour sensibiliser aux bonnes pratiques de gestion et publication des données scientifiques. Une initiative de l'UMR TETIS et l'UMR Espace Dev avec le soutien de l’Appel à Projets Science Ouverte AgroParisTech. 📰 Lire l'article : tinyurl.com/33rhpvfw
031
Reposted by Rémy Decoupes
Me AI @me-ai.bsky.social · 10/01/2026
Arabic has 422 million speakers, yet most AI systems treat it as an afterthought. The Technology Institute of the UAE just released Falcon-H1-Arabic, the first Arabic language model built on hybrid Mamba-Transformer architecture. This isn't another scaled-up model with better Arabic.. (1/7)
ArXiv page 1
142
Rémy Decoupes @remydec.bsky.social · 07/01/2026
Really interesting AI Forensics report about how we search for information using AI chatbots. It raises yet another layer of concern. aiforensics.org/work/governi...
aiforensics.org
From 'Googling' to 'Asking ChatGPT': Governing AI Search
We provide an overview of the most relevant changes and shifts between traditional search engines and AI search functionalities. This report also offers a framework for situating AI search in the curr...
000
Reposted by Rémy Decoupes
EarthArXiv Bot @eartharxivbot.bsky.social · 03/12/2025
Longitudinal assessment of research in GIScience domain shows a positive impact of reproducible research practices doi.org/10.31223/X5RJ3W
001
Reposted by Rémy Decoupes
Lê Nguyên Hoang @science4all.org · 28/11/2025
Yesterday, an #OpenReview vulnerability led to the leak of reviewer identities of all the major academic AI conferences, including the ongoing #ICLR2026 conferences. #ICLRLeaks This is both a huge disaster, and an opportunity to tackle the serious flaws of AI research. eu.36kr.com/en/p/3572028...
eu.36kr.com
Academic Circle in Uproar: ICLR Reviewers Reveal Identities, Low Scores Given by Friends
True Open Review: Unveiling Transparent and Authentic Evaluations
1358
Reposted by Rémy Decoupes
Sebastian Raschka (rasbt) @rasbt.bsky.social · 03/12/2025
This interesting week started with DeepSeek V3.2! I just wrote up a technical tour of the predecessors and components that led up to this: 🔗 magazine.sebastianraschka.com/p/technical-... - Multi-Head Latent Attention - RLVR - Sparse Attention - Self-Verification - GRPO Updates
magazine.sebastianraschka.com
A Technical Tour of the DeepSeek Models from V3 to V3.2
Understanding How DeepSeek's Flagship Open-Weight Models Evolved
1367
Rémy Decoupes @remydec.bsky.social · 01/12/2025
BERT-like models are still the most downloaded models on Hugging Face (45%), compared with decoder-only models (9%). www.reddit.com/r/LlamaFarm/...
reddit.com
From the LlamaFarm community on Reddit: "We're in an LLM bubble, not an AI bubble" - Here's what's actually getting downloaded on HuggingFace and how you can start to really use AI.
Explore this post and more from the LlamaFarm community
000
Reposted by Rémy Decoupes
Wikipedia @wikipedia.org · 21/11/2025
25 years of humanity at its best. #Wikipedia25 Donate now ➡️ donate.wikipedia25.org
7371116
Reposted by Rémy Decoupes
ARTE @artefr.bsky.social · 21/11/2025
Les géants américains et chinois se disputent la suprématie dans le secteur de l'IA et cette révolution technologique soulève évidemment des enjeux stratégiques et éthiques 🤖
youtube.com
Intelligence artificielle : une compétition mondiale | Le dessous des cartes - ARTE
Les géants américains et chinois se disputent la suprématie dans le secteur de l'intelligence artificielle, avec des investissements colossaux. L'IA générative, qui produit textes, images et musiques, est au coeur de cette bataille. Cette révolution technologique soulève des enjeux stratégiques et é
02712
Reposted by Rémy Decoupes
n8rob.bsky.social @n8rob.bsky.social · 12/11/2025
Interested in developing LLMs that work for dialectal Arabic? Introducing the AMIYA shared task: Arabic Modeling In Your Accent, just accepted to VarDial 2026. Please consider submitting and joining us in Morocco if you do! sites.google.com/view/vardial...
sites.google.com
VarDial 2026 - Shared Tasks
AMIYA (عامية) Shared Task: Arabic Modeling In Your Accent The AMIYA shared task will offer a chance for researchers to demonstrate innovations and improvements in language modeling of dialectal Arabic...
174
Reposted by Rémy Decoupes
Platial Analysis Lab @platialanalysis.bsky.social · 11/11/2025
The program for the 6th Spatial Data Science Symposium is now online. Registration will open soon. The symposium is Dec 4 & 5. Plan to check out the Thematic Sessions, Paper Sessions, Emerging Researchers Panel, Keynotes, and Interview. sdss2025.spatial-data-science.net/program.html
086
Reposted by Rémy Decoupes
EMNLP @emnlpmeeting.bsky.social · 07/11/2025
🌉 #EMNLP2026 will be October 24-29th in Budapest! 🌉 Thanks all for a great conference, and see you at the next one!
An image of a conference presentation slide showing that EMNLP 2026 will be held October 24-29th in Budapest, with an audience below
1214
Reposted by Rémy Decoupes
Sebastian Raschka (rasbt) @rasbt.bsky.social · 04/11/2025
My new field guide to alternatives to standard LLMs: Gated DeltaNet hybrids (Qwen3-Next, Kimi Linear), text diffusion, code world models, and small reasoning transformers. 🔗 magazine.sebastianraschka.com/p/beyond-sta...
05315
Reposted by Rémy Decoupes
404 Media @404media.co · 03/11/2025
Because of an onslaught of AI-generated research, specifically in the computer science (CS) section, arXiv is going to limit which papers can be published. @mjgault.bsky.social has more:
404media.co
arXiv Changes Rules After Getting Spammed With AI-Generated 'Research' Papers
Cornell University’s academic paper repository will no longer accept Computer Science papers still under review.
27723
Reposted by Rémy Decoupes
Julien Chaumond @julien-c.hf.co · 02/11/2025
Training LLMs end to end is hard. But way more people should, and will, be doing it in the future. The @hf.co Research team is excited to share their new e-book that covers the full pipeline: · pre-training, · post-training, · infra. 200+ pages of what worked and what didn’t. ⤵️
415227
Reposted by Rémy Decoupes
Alexander Doria @dorialexander.bsky.social · 28/10/2025
So we're hiring.
32010
Rémy Decoupes @remydec.bsky.social · 28/10/2025
Our latest paper has just been published. We explore geographical biases in language models and their implications (with a focus on crisis monitoring). @interdonatos.bsky.social , @matroche.bsky.social , M.Teisseire & S.Valentin. Code: github.com/tetis-nlp/ge... link.springer.com/article/10.1...
link.springer.com
Evaluation of geographical distortions in language models - Machine Learning
Geographic bias in language models (LMs) is an underexplored dimension of model fairness, despite growing attention being given to other social biases. We investigate whether LMs provide equally accur...
011
Rémy Decoupes @remydec.bsky.social · 28/10/2025
It's a very nice article in which the author compares the changes in the neural network architecture among the most famous 2025 LLMs: The Big LLM Architecture Comparison magazine.sebastianraschka.com/p/the-big-ll...
magazine.sebastianraschka.com
The Big LLM Architecture Comparison
From DeepSeek-V3 to Kimi K2: A Look At Modern LLM Architecture Design
010
Reposted by Rémy Decoupes
Tom Aarsen @tomaarsen.com · 22/10/2025
🤗 Sentence Transformers is joining @hf.co! 🤗 This formalizes the existing maintenance structure, as I've personally led the project for the past two years on behalf of Hugging Face. I'm super excited about the transfer! Details in 🧵
2316
Reposted by Rémy Decoupes
Ludovic Moncla @ludovicmoncla.bsky.social · 16/10/2025
Un projet à l’intersection de l’histoire, de la géomatique, des sciences du langage et de l’IA. Les slides de ma présentation : docs.google.com/presentation... Modèles et données : huggingface.co/GEODE Démonstration : huggingface.co/spaces/GEODE...
docs.google.com
2025 RnMSH Talk
L’usage de l’IA pour une étude interdisciplinaire de la géographie dans l’Encyclopédie de Diderot et d’Alembert Ludovic Moncla ludovic.moncla@insa-lyon.fr Pratiques d’intelligence artificielle appliqu...
001
Reposted by Rémy Decoupes
EOSC SIESTA @eosc-siesta.eu · 06/10/2025
🎥 Tools SIESTA | Anjana 🧩 A Python library that makes data anonymization simple & powerful with techniques like generalization, suppression & microaggregation. 👉 Watch: youtu.be/yw3tOd8WuIU #EOSCSIESTA
youtu.be
Anonymizing Sensitive Tabular Data: A Practical Guide with Anjana
YouTube video by EOSC SIESTA
036
Reposted by Rémy Decoupes
Alexander Doria @dorialexander.bsky.social · 27/09/2025
And new paper out: Pleias 1.0: the First Family of Language Models Trained on Fully Open Data How we train an open everything model on a new pretraining environment with releasable data (Common Corpus) with an open source framework (Nanotron from HuggingFace). www.sciencedirect.com/science/arti...
817951