Sign in

Jan Kabatek 💙💛

@jankabatek.com
1.6K followers 613 following 894 posts

Associate Professor at the Melbourne Institute, University of Melbourne. www.jankabatek.com

PostsRepliesMedia
Jan Kabatek 💙💛 @jankabatek.com · 05/08/2026
On a related note, searches for "VPN":
030
Jan Kabatek 💙💛 @jankabatek.com · 05/08/2026
This is a striking chart. It shows the consequences of mandatory age-verification for accessing, ehm, select websites in Australia. The search term in question is the name of one of the most visited sites on the internet. I'll let you guess which one.
120
Jan Kabatek 💙💛 @jankabatek.com · 04/06/2026
SEHO2026 is delivering both great talks and stunning views ❤️🇮🇹
020
Jan Kabatek 💙💛 @jankabatek.com · 14/05/2026
Day 2 of the LEAP summit is kicking into great with a tour-de-force keynote by Mark Hatzenbuehler (Harvard) on the topic of structural stigma among LGBTQ+ populations
020
Jan Kabatek 💙💛 @jankabatek.com · 13/05/2026
Day 1 has been chock-full of incredible talks, including an inspired keynote by Donn Feir on the status of LGBTQ+ research in economics, and a wonderful community panel chaired by Joe Ball, the Victorian Commissioner for LGBTIQA+ Communities
Prof. Donn Feir asking a questionDr. Karina Saxby introducing the community panelCommissioner Joe Ball chairing the community panel
010
Jan Kabatek 💙💛 @jankabatek.com · 13/05/2026
The inaugural LEAP Summit for LGBTQ+ Research and Researchers in Asia-Pacific is in full swing! 🥳🌈 @leap-econ.bsky.social @ksaxby.bsky.social @unimelb.edu.au
283
Jan Kabatek 💙💛 @jankabatek.com · 24/03/2026
And all of this started because Stanisław Ulam didn't feel like calculating the odds of winning a game of Solitaire...
050
Jan Kabatek 💙💛 @jankabatek.com · 20/01/2026
If you'd like to embed a Bluesky feed on your website, check out this neat and simple widget: github.com/Vincenius/bs...
060
Jan Kabatek 💙💛 @jankabatek.com · 02/10/2025
A really neat article on fertility: ourworldindata.org/total-fertil... Charts out the fertility dynamics and explains the complicated relationships between different statistical measures, such as TFR and CCFR.
080
Jan Kabatek 💙💛 @jankabatek.com · 08/09/2025
This indicator is extremely useful. One can easily track the latent drivers of poverty, and evaluate which ones are becoming more important over time 📈
001
Jan Kabatek 💙💛 @jankabatek.com · 18/08/2025
And that's a wrap of the 2025 Labour Econometrics Workshop 😊 We've seen great presentations, inspiring keynotes, and all the flavours of Melbourne weather 🌈 Thank you everyone for making this year extra special! ❤️
060
Jan Kabatek 💙💛 @jankabatek.com · 05/08/2025
Furthermore, we have experienced a 15% spike in divorce rates in 2021. This was also likely fueled by the lockdowns. The spike came in 2021 (not in 2020), because Australia mandates 12-month separation period prior to divorce.
100
Jan Kabatek 💙💛 @jankabatek.com · 05/08/2025
During Covid, Australian marriage rate went down by 35%. This reflected: 1) many couples with marriage intentions waiting out the lockdowns 2) decline in relationship formation during lockdowns
100
Jan Kabatek 💙💛 @jankabatek.com · 01/07/2025
FWIW, Joop et al's work is accredited in the author's reanalysis link.springer.com/article/10.1...
110
Jan Kabatek 💙💛 @jankabatek.com · 15/04/2025
Hmmm.... #Nature
291
Jan Kabatek 💙💛 @jankabatek.com · 09/04/2025
. 𝒓𝒆́𝒈 𝒚 𝒔𝒖𝒓 𝒙, 𝒓𝒐𝒃𝒖𝒔𝒕𝒆
613915
Jan Kabatek 💙💛 @jankabatek.com · 16/01/2025
à propos of nothing, here's a little #Stata script that sets up a standardized folder structure for new research projects (inspired by @asjadnaqvi.bsky.social) 💾 LINK: jankabatek.com/stata/folder...
35913
Jan Kabatek 💙💛 @jankabatek.com · 04/12/2024
I did not go up to 20M, but this should give you an idea... (normal reg obviously runs out of memory) I have not looked into the guts of these methods, but I am fairly confident that they use the same FWL matrix transformations, which do not get much more intensive as the number of FEs grows
011
Jan Kabatek 💙💛 @jankabatek.com · 03/12/2024
The GTOOLS picture was rendered weird... here is the right one:
000
Jan Kabatek 💙💛 @jankabatek.com · 03/12/2024
The GTOOLS picture was rendered weird... here is the right one:
000
Jan Kabatek 💙💛 @jankabatek.com · 03/12/2024
✔️ So consider using my command 𝐞𝐱𝐩𝐚𝐧𝐝𝐫𝐚𝐧𝐤, which is 20-30x faster and free of redundancies: . net install expandrank, from("https://jankabatek.com/stata/expandrank/") replace . sysuse cancer, clear . drop _* . expandrank studytime, name(time_disc) fast
110
Jan Kabatek 💙💛 @jankabatek.com · 03/12/2024
Note that the relative performance depends *𝐜𝐫𝐢𝐭𝐢𝐜𝐚𝐥𝐥𝐲* on the size of your dataset (see the pictures below). 👉 𝐫𝐞𝐠𝐡𝐝𝐟𝐞𝐣𝐥 performs particularly well in big datasets (10M+), but badly in small ones. (FWIW, I don't know why the factorized regressions are subject to that structural break 📉)
200
Jan Kabatek 💙💛 @jankabatek.com · 03/12/2024
👇 The chart below illustrates the relative performance of the aforementioned regression approaches in a scenario with 1M observations and a variable number of fixed effects. 🖥️ My hardware specs: i7-8700 (8 cores), 16GB RAM, no GPU processing, OS Win11
Relative performance of four regression approaches when estimating a HDFE model in a dataset with 1 million observations. Standard regression performs the best up to approx. 200 fixed effects, whereas the areg and reghdfejl perform the best in scenarios with more fixed effects.
131
Jan Kabatek 💙💛 @jankabatek.com · 03/12/2024
1 ) 𝐆𝐓𝐎𝐎𝐋𝐒 for the win! First and foremost, embrace the 𝐆𝐓𝐎𝐎𝐋𝐒 suite by @mcaceresb.bsky.social . ssc install gtools Mauricio created many blazing-fast alternatives to Stata's native commands. If you're not using them, you're missing out! Also, he is on the market 💪
210
Jan Kabatek 💙💛 @jankabatek.com · 03/12/2024
Here's another batch of tips for dealing with massive datasets in #Stata! Today's theme is 𝑺𝑷𝑬𝑬𝑫 🚀
1256
Jan Kabatek 💙💛 @jankabatek.com · 29/11/2024
The reference was to this figure... though now I realise that "p" used here probably means something else than p-values...?
110
Jan Kabatek 💙💛 @jankabatek.com · 26/11/2024
Celebrating our successful ARC Discovery bid 🥳 Huge thank you goes to Vincent Mancini for being the driving force behind this exciting project!
380
Jan Kabatek 💙💛 @jankabatek.com · 20/11/2024
Conditional colouring (introduced only in Stata 18 ❗) can be used to distinguish significant and insignificant coefficient estimates in 𝗰𝗼𝗲𝗳𝗽𝗹𝗼𝘁 👍 . sysuse auto, clear . reg trunk i.rep78 . gen COL=0 in 1/5 . replace COL=1 in 2 . replace COL=1 in 5 . coefplot, recast(bar) colorvar(COL) citop
170
Jan Kabatek 💙💛 @jankabatek.com · 15/11/2024
Or just show the full picture... 🤷‍♂️
100
Jan Kabatek 💙💛 @jankabatek.com · 15/11/2024
Add the vertical axis line with the -//- symbol (denoting cropped scaling) for extra style points... 👍
100
Jan Kabatek 💙💛 @jankabatek.com · 15/11/2024
To appreciate this, let's crop the lower range of the y-axis a bit more. The grid no longer splits the chart into equal segments, and the position of the lowest gridline cues the reader to the fact that the x-axis is unlikely to coincide with zero. The visual becomes immediately less deceptive ✔️
100
Jan Kabatek 💙💛 @jankabatek.com · 15/11/2024
Amusing as it is, I don't like this chart. Here's why 👇
230
Jan Kabatek 💙💛 @jankabatek.com · 30/10/2024
Extremely happy to learn that my Honours student Max Yong has secured the John Monash Scholarship to pursue his postgraduate degree at Harvard Kennedy School 🥳 johnmonash.com/contact/find...
160
Jan Kabatek 💙💛 @jankabatek.com · 22/10/2024
Fixed it for you, Elsevier
061
Jan Kabatek 💙💛 @jankabatek.com · 21/10/2024
I am very excited to share with you a heavily-updated edition of my tips for dealing with massive datasets in #Stata 🥳 Today's theme is 𝐌𝐄𝐌𝐎𝐑𝐘. Other themes will follow over the coming weeks 🙂 [a thread]
How to handle large datasets in Stata. 
Episode 1: Memory
23612
Jan Kabatek 💙💛 @jankabatek.com · 11/10/2024
Mini #Stata tip that came in handy today: Use graphing option 𝘅𝗹𝗮𝗯𝗲𝗹( ,𝗳𝗼𝗿𝗺𝗮𝘁(%𝟬𝟮.𝟬𝗳)) to add leading zeros to single-digit calendar year labels
150
Jan Kabatek 💙💛 @jankabatek.com · 03/10/2024
Each country tailors its rules to its specific context... Here are the official do's and don'ts of administrative data access in Australia 🦘
020
Jan Kabatek 💙💛 @jankabatek.com · 02/10/2024
I am a big fan of this #Stata script I cobbled some time ago: . ssc install expandrank It expands on the 'expand' script by adding the next logical step: create a rank variable. 𝗕𝗨𝗧: the rank is calculated without sorting the data, which renders the code super fast w/ big datasets (~30x faster).
191
Jan Kabatek 💙💛 @jankabatek.com · 20/09/2024
A #Stata bookmark to myself. 𝗛𝗼𝘄 𝘁𝗼 𝗳𝗼𝗿𝗰𝗲 𝗮 𝗹𝗲𝗴𝗲𝗻𝗱 𝗶𝗻𝘁𝗼 𝘁𝗵𝗲 𝗽𝗹𝗼𝘁 𝗿𝗲𝗴𝗶𝗼𝗻: . sysuse auto, clear . twoway lfitci price weight, legend(ring(0) pos(5) bmargin(medium)) ring(0) is the essential option pos(1-12) determines the (clock) position of the legend bmargin() sets the offset from the border
172
Jan Kabatek 💙💛 @jankabatek.com · 16/09/2024
As you would expect, the problem disappears if we define x as a 𝗱𝗼𝘂𝗯𝗹𝗲, so that both sides of the if-condition belong to the same data type:
100
Jan Kabatek 💙💛 @jankabatek.com · 16/09/2024
The calculations include evaluations of if-conditions, which is why x [stored as 𝗳𝗹𝗼𝗮𝘁 0.69999999] is lower than 0.7 [interpreted as 𝗱𝗼𝘂𝗯𝗹𝗲 0.6999999999999...]
100
Jan Kabatek 💙💛 @jankabatek.com · 13/09/2024
Another example for good measure:
020
Jan Kabatek 💙💛 @jankabatek.com · 13/09/2024
This #Stata discovery will haunt me forever... 👇 FLOATING-POINT PRECISION ISSUES 👇
5174
Jan Kabatek 💙💛 @jankabatek.com · 11/09/2024
Reminded myself of this neat Stata trick: 𝗰𝗵𝗮𝗿 command allows you to adjust the width of specific columns in the Data Editor: sysuse bplong char patient[_de_col_width_] 20 char sex[_de_col_width_] 7 browse It's perfect for dealing with ultra-long string variables in the data editor.
143
Jan Kabatek 💙💛 @jankabatek.com · 15/11/2023
And here's a multi-coloured code: sysuse auto, clear reg tr i.rep78 gen COL = _n coefplot, recast(bar) citop colorvar(COL) colorcuts(0(0.25)5)
000
Jan Kabatek 💙💛 @jankabatek.com · 15/11/2023
Here's an example that shows how to use conditional colouring in Stata 18 to distinguish significant and insignificant coefficient estimates: sysuse auto, clear reg tr i.rep78 gen COL=0 replace COL=1 in 2 replace COL=1 in 5 coefplot, recast(bar) colorvar(COL) citop
140