Paul Bose @pbose.bsky.social · 27/07/2025At the time I was using both inline completions and the chat function (at least the inline chat, not the agent capabilities, which didn't exist yet). 000
Paul Bose @pbose.bsky.social · 02/07/2025Interesting, thanks for sharing. I didn't know about this. Tbh I mostly did this exercise to test out the gender and region prediction package. But thought the results might be interesting. Nice to know that there is a resource directly from RePEc. 000
Paul Bose @pbose.bsky.social · 02/07/2025Check out my blog post for a detailed explanation of my methodology and findings: www.paulbose.com/thisandthat/... 100
Paul Bose @pbose.bsky.social · 02/07/2025I used the "nametrace" package (github.com/parobo/namet...) to predict gender and region of origin based on author names rather than checking each author's actual gender or background. Therefore, take the results with a grain of salt.github.comGitHub - parobo/nametrace: A python package to predict demographic information from names.A python package to predict demographic information from names. - parobo/nametrace 100
Paul Bose @pbose.bsky.social · 02/07/2025The rise in female representation in the top 5% appears to be a global phenomenon, with increases observed in both North America/Europe and the "rest of the world." The trend is slightly stronger in North America and Europe, but these regions also had more ground to cover in terms of catching up. 100
Paul Bose @pbose.bsky.social · 02/07/2025The growth for regions other than US/Canada/Europe is primarily driven by scholars from Eastern Asia, Southern Asia and South America. Unfortunately, Africa remains extremely underrepresented. 100
Paul Bose @pbose.bsky.social · 02/07/2025While these numbers are super low, there's a small of positive change, ... a slow one. - The share of women has edged up from roughly 9% to 12% over the past 12 years. Progress, but the pace is too slow! - Representation from "the rest of the world" has increased from 16% to 25% in 2025. 100
Paul Bose @pbose.bsky.social · 02/07/2025How is economics doing in terms of representation of women and researchers with a region of origin outside North America or Europe? I looked at the top 5% authors on RePEc in the last 12 years. - Only 12% are women - 25% a region of origin other than US/Canada/Europe #EconSky #AcademicSky #PoliSky 161
Paul Bose @pbose.bsky.social · 26/06/2025Need to predict people's gender or region of origin from their name for your research? Check out my python package "nametrace" which provides a simple modern API to do just that. 110
Paul Bose @pbose.bsky.social · 25/06/2025It was great to be in Clermont Ferrand to present my work with @econom.bsky.social on local social media activity after refugee arrival. Thanks so much to the organizers and the other participants. It was a super interesting workshop! 041
Reposted by Paul BoseZohal Hessami @zohalhessami.bsky.social · 05/06/2025🚨3 days left to apply!media.tenor.coma man is sitting at a desk in front of a computer and smiling .ALT: a man is sitting at a desk in front of a computer and smiling . 021
Paul Bose @pbose.bsky.social · 27/05/2025Our results suggest that refugee influx caused a sharp but short lived spike in salience of refugees. People who remained active tweeters on the topic started to show more opposition of refugees after a while however. We combine the analysis with extremely local voting data and find similar results. 000
Paul Bose @pbose.bsky.social · 27/05/2025Olivier Marie is presenting our paper (joint with Renske Stans) on social media salience of refugees and election effects in Linnaeus today. I am super excited that this paper ready to be presented more widely! We study local twitter discourse around the timing of refugee arrival during 2015/2016. 140
Paul Bose @pbose.bsky.social · 20/05/2025As an example I show how to analyze the sentiment of the last 100 posts of the reddit CEOs spez and kn0thing in mere seconds. 000
Paul Bose @pbose.bsky.social · 20/05/2025Are you using LLMs for your research and want to classify millions of text? This can be a very slow and expensive process. But it doesn't have to be. In my blogpost I explain how to use multiple GPUs and vLLM to analyse thousands of texts with an LLM super fast! #econsky #polisky lnkd.in/di9AbciU 143
Paul Bose @pbose.bsky.social · 07/05/2025Want to learn how to finetune a large language model for your specific needs? I wrote a short post on how to train LLAMA for a classification task: www.paulbose.com/thisandthat/... #econsky #poliskypaulbose.com Finetuning an LLM to do classification | Paul Bose LLMs are ubiquitous at this point. ChatGPT for example is constantly used to generate labels for unstructured data such as text. However using ChatGPT might be too expensive or not perform very well f... 101
Paul Bose @pbose.bsky.social · 26/02/2025BONUS: use ollama to interact with the models from python and receive structured responses. This is super helpful for classification or structured outputs for e.g. Text summary. 000
Paul Bose @pbose.bsky.social · 26/02/2025Want to learn how to run your own version of Deepseek R1 or Meta's LLAMA model on a remote high-performance computing server? I wrote a brief blog post explaining how you can install and run the models using ollama. www.paulbose.com/thisandthat/... #econsky #polisky 162
Paul Bose @pbose.bsky.social · 06/02/2025FYi, in case you might care, here is an explanation how the models moderate internally (i.e. without using the API, but using system prompts). www.lesswrong.com/posts/jGuXSZ...lesswrong.comRefusal in LLMs is mediated by a single direction — LessWrongThis work was produced as part of Neel Nanda's stream in the ML Alignment & Theory Scholars Program - Winter 2023-24 Cohort, with co-supervision from… 011
Reposted by Paul BoseFrancesco Sobbrio @fsobbrio.bsky.social · 29/01/2025Last chance to submit a paper (Deadline Jan 31). 3rd CEPR Workshop on Media, Technology, Politics, and Society cepr.org/events/3rd-b...cepr.org3rd BOCCONI - CEPR Workshop on Media, Technology, Politics, and SocietySearch the site 043
Paul Bose @pbose.bsky.social · 28/01/2025I seems the abliterated versions are not censored but the the base model seems to be: bsky.app/profile/paul... 020
Paul Bose @pbose.bsky.social · 28/01/2025Interesting point, your screenshots do somewhat indicate hard coded instructions in the model, but it could still be only on the API version. Would like to see this for the local model. 110
Paul Bose @pbose.bsky.social · 28/01/2025Yes, I would assume that Deepseek generates the unmoderated answers internally, but has some instructions to not show them in the official API, but I wouldn't be surprised if the local version does show the hidden responses. (Let me know if you get to run it locally, I wanted to look into this too) 120
Paul Bose @pbose.bsky.social · 28/01/2025This is using the official Deepseek AI right? Or are these kind of responses hard coded into the open source model as well (i.e. if I run it locally, will it give me different responses)? 210
Paul Bose @pbose.bsky.social · 13/10/2024I mostly code in python and at some point decided to use parquet whenever possible. For me a very big advantage is that datatypes are also stored and that you can query the files similar to how you would an sql database. I am not sure if nowadays you can read them into Stata though. 010
Paul Bose @pbose.bsky.social · 10/10/2024Yes that's true. I have to admit, personally I frequently get to the point were I have to "review" code that I would not have written like this myself and might not understand. But at the same time it enables me to use packages that I didn'd ever use before without much preparation. 010
Paul Bose @pbose.bsky.social · 10/10/2024Interesting, I just saw someone post this article a couple of days ago, that suggest there is no productivity gain for devs using Copilot. My guess would be that there are substantial difference in productivity gains by experience and skill level? www.cio.com/article/3540...cio.comDevs gaining little (if anything) from AI coding assistantsCode analysis firm sees no major benefits from AI dev tool when measuring key programming metrics, though others report incremental gains from coding copilots with emphasis on code review. 110
Paul Bose @pbose.bsky.social · 04/09/2024For official documents (also in Italian), I have had very good experience with deepl.com, for not super long docs its free and the translation is generally better than Google translate I feel. 010
Paul Bose @pbose.bsky.social · 04/09/2024Amazing tip. Didn't know about this one yet. Super helpful :) 110
Paul Bose @pbose.bsky.social · 03/09/2024 Also helpful: Write a comment with what you want Copilot to do (** codeblock that does XYZ), often the inline code prediction will have useful code suggestions. 010
Paul Bose @pbose.bsky.social · 03/09/2024I have to admit that I haven't tried GPT4 much for Stata yet. I would guess that it outperforms Copilot for more complex prompts. The advantage of Copilot is definitely the inline capability, which speeds up coding without needing to prompt anything. 100
Paul Bose @pbose.bsky.social · 03/09/2024I am mainly using Copilot with VSCode for python and find it super helpful. For me the main advantage over ChatGPT is the inline prediction. With Stata it is definitely less helpful (I think because there is just less Stata code on github), but the inline predictions are still pretty good I think. 210
Paul Bose @pbose.bsky.social · 14/03/2024The full paper is available on arXiv (arxiv.org/abs/2403.05700) and is forthcoming at @LrecColing 2024. It was truly a pleasure to work with my great co-authors on this project. Thank you! 000
Paul Bose @pbose.bsky.social · 14/03/2024We train an XLM classifier that combines information on users' tweets, and self-written bio. The classifier is trained on Italian user, but also performs well on a German test set. 100
Paul Bose @pbose.bsky.social · 14/03/2024Can we learn something about gender and age from a person's tweets? Yes! In a new project with Lorenzo Luro, Mahyar Habibi, Dirk Hovy and Carlo Schwarz, we provide a labeled dataset of 20k Twitter users. We show that demographic classifiers using tweets outperform traditional tools. 152
Paul Bose @pbose.bsky.social · 15/11/2023I think Clément de Chaisemartin recently posted a new version of did_multiplegt which performs about 100 times faster than the old version. I think it was called did_multiplegt_dyn and should be on SSC. 120
Paul Bose @pbose.bsky.social · 30/10/2023I don't know how many department seminars are still online, but there are two young researcher seminar series that are hosted online and open: @ayew.bsky.social: www.monash.edu/business/imp... and PhD-EVS: sites.google.com/view/phd-evs... 130