Sign in

alex hayes

@alexpghayes.com
4.7K followers 2K following 484 posts

assistant prof @ oregon state statistics. networks, causal inference, contagion, measurement error, #rstats. he/him www.alexpghayes.com

PostsRepliesMedia
alex hayes @alexpghayes.com · 13/08/2026
very cool investigation into consequences of the friendship paradox
130
Reposted by alex hayes
Jessica Hullman @jessicahullman.bsky.social · 11/08/2026
Nice post from @eytan.adar.prof on the cargo cult science of "AI-native" PhD students, who come in super productive because they perform all major parts of research with genAI. No one needs all these quick papers, and they can really hurt students down the road. eytanadar.medium.com/ai-native-ph...
eytanadar.medium.com
AI-Native PhD Students
If you’re a student and just want the tl;dr, feel free to jump to the end.
24312
alex hayes @alexpghayes.com · 08/08/2026
I think the thing is you have to genuinely design the class as a service class and that's could be unfamiliar/unpopular/less of an ad for your major
020
alex hayes @alexpghayes.com · 08/08/2026
I think there are a number of reasonable curricula for a singular stat class (Calling Bullshit and friends)! Could develop around Spiegelhalters Art of Statistics very reasonably, etc, etc
140
Reposted by alex hayes
Adam L @adam-lg.bsky.social · 08/08/2026
My argument would be more: we shouldn't be teaching people's one stats class as "here's a cookbook of tests to get a p-value"
3161
alex hayes @alexpghayes.com · 08/08/2026
I am somewhat surprised by the lack of, let's say, public numeracy courses in gen ed requirements Seems like there is lots of intro to numerical thinking and less on numerical consumption
172
alex hayes @alexpghayes.com · 07/08/2026
Brutal
Screencap from linked report claiming 94 percent of DI athletics programs cost more than they bring in
010
alex hayes @alexpghayes.com · 06/08/2026
seeing lots of use in LLM land at the moment as a way to reduce model evaluation costs -- only evaluate on the hardest/most informative questions arxiv.org/pdf/2402.14992
arxiv.org
100
Reposted by alex hayes
Nick Strayer @nstrayer.bsky.social · 31/07/2026
Check out the blog post Wasim and Cindy Tong wrote announcing the GA launch: opensource.posit.co/blog/2026-07... Give it a try and let us know how you get on!
opensource.posit.co
Positron's Jupyter Notebook Editor Is Now Generally Available
The Positron Jupyter Notebook editor is now generally available, bringing a first-class notebook experience with built-in environment management, interactive data exploration, streamlined version cont...
0144
alex hayes @alexpghayes.com · 30/07/2026
Yeah it's better. Render time is actually reasonable for large documents
030
alex hayes @alexpghayes.com · 29/07/2026
I'll be at New Researchers Conference 2026 this weekend chatting about some recent work on estimating peer effects in noisy networks! arxiv.org/abs/2605.03204
0103
alex hayes @alexpghayes.com · 29/06/2026
Some LaTeX resources I'd appreciate: - A modern guide to best LaTeX practices - A concise guide to mathematical writing - Typst <-> LaTeX conversion gotchas - Claude skills for all of the above
060
alex hayes @alexpghayes.com · 12/06/2026
@grimalkina.bsky.social i love your learning-opportunities claude skill -- do you know of anything similar for web chat interfaces? the web interface is much nicer than clis when discussing math since it renders latex i'm curious about user prompts both for myself and also to suggest to students!
160
alex hayes @alexpghayes.com · 22/05/2026
my assumption is that multilevel models need more structure than just a random intercept to beat glmnet since there a lambda such that ridge and random intercepts are equal
110
alex hayes @alexpghayes.com · 22/05/2026
@avehtari.bsky.social curious if any of these have come up in your work
010
alex hayes @alexpghayes.com · 22/05/2026
Returning to this! Do you know of any empirical applications of this estimator or sample code?
100
alex hayes @alexpghayes.com · 22/05/2026
i'm looking for interesting teaching examples where (1) random effects/multilevel models predict better than glmnet (2) glmnet predicts better than random effects #rstats #statistics
3115
alex hayes @alexpghayes.com · 15/05/2026
I appreciated reading the original comment. I see www.sciencedirect.com/science/arti... suggests bootstrapping for estimating risk of predictive models; do you have more extensive advice about model selection for clinical models somewhere?
sciencedirect.com
Internal validation of predictive models: Efficiency of some procedures for logistic regression analysis
The performance of a predictive model is overestimated when simply determined on the sample of subjects that was used to construct the model. Several …
000
Reposted by alex hayes
Thomas Dietterich @tdietterich.bsky.social · 13/05/2026
We are implementing a similar policy at @arxiv.bsky.social. If there is incontrovertible evidence of LLM slop in a paper, this means the authors did not take the time to read the LLM output and we can't trust anything else in the paper. Penalty is 1 year ban from arXiv followed by...
15412139
alex hayes @alexpghayes.com · 13/05/2026
I'm curious what the other signatures are!
100
alex hayes @alexpghayes.com · 10/05/2026
i wrote up some notes! what's a good email?
110
alex hayes @alexpghayes.com · 06/05/2026
i'm quite excited about this work and how it allows us to handle noise in networks with spectral tools we still have some time before we submit to a journal, so questions, comments and feedback are very welcome!
010
alex hayes @alexpghayes.com · 06/05/2026
our theory covers weighted networks observed with additive noise in simulations we show our approach also works for: - networks with missing edges - ego-centric network data - aggregated relational data this is because we know how to estimate the spectrum well under these noise processes
110
alex hayes @alexpghayes.com · 06/05/2026
this leads to a general approach for handling noise in low-rank networks 1. choose a spectral estimator appropriate to your noise process 2. plug in your low-rank estimate of E[A] for A in a typical two-stage least squares estimator
110
alex hayes @alexpghayes.com · 06/05/2026
so far this is intellectually interesting. but it turns out it is also immensely useful the latent contagion estimator only requires an estimate of the spectrum of E[A] to construct crucially, we can often get a good estimate of E[A] from a noisy observation of A
110
alex hayes @alexpghayes.com · 06/05/2026
why are peer and latent contagion two-stage least squares estimators equivalent in low-rank networks? the network A concentrates quickly around E[A]. estimates of E[A] also concentrate quickly around E[A] in sufficiently dense networks the noise in A or the estimate around E[A] gets averaged away
110
alex hayes @alexpghayes.com · 06/05/2026
this is not a generic result; it requires low-rank structure in your network! think: random dot product graphs the converse is also true: a peer contagion estimator can also be used to estimate coefficients under a latent contagion data generating process
110
alex hayes @alexpghayes.com · 06/05/2026
if you use latent contagion as working model to derive an estimator, and then use that estimator on data from a peer contagion, everything just works the incorrectly specified latent contagion estimator is asymptotically indistinguishable from the correct peer contagion estimator!
110
alex hayes @alexpghayes.com · 06/05/2026
2. the second, and potentially much larger problem, is that contagion in a latent space might not be causally meaningful contagion in a latent space might not correspond to a "real" contagion! this is where things get really cool
110
alex hayes @alexpghayes.com · 06/05/2026
there's two problems with modeling contagion in a latent space 1. these models can have vanishing signal-to-noise ratios. this derailed the project for 18 months while we figured out what was going on, and lead to a whole other manuscript explaining the problem bsky.app/profile/alex...
110
alex hayes @alexpghayes.com · 06/05/2026
what can you do if the network is noisy? one idea is to use a latent space network model and model peer effects as happening in a latent space then, the presence or absence of individual edges matters less, and instead you need to be able to estimate the latent structure of the network
Screencap of paper introduction, in particular the formula for latent and peer contagion network autoregressive models
120
alex hayes @alexpghayes.com · 06/05/2026
typically, peer effects estimators assume that you have a perfectly accurate observation of a social network but most people who have worked with network data know that measuring networks is challenging! you often only see a noisy variant of the network
120
alex hayes @alexpghayes.com · 06/05/2026
new print! keith levin and i came with a new approach to estimating peer effects in noisy, low-rank networks i'm very excited about our idea because the approach works with a variety of different noise processes arxiv.org/abs/2605.03204
Screenshot of the first page of the pre-print, showing the title, abstract and first paragraph of the text.
1102
Reposted by alex hayes
Noah Greifer @noahgreifer.bsky.social · 01/05/2026
This is the clearest and most accessible introduction to DML I've ever read: arxiv.org/abs/2504.08324 Congrats to the authors @aahrens.bsky.social, @markeschaffer.bsky.social, et al on this amazing paper and great accompanying R package! I'm now DML-pilled.
arxiv.org
An Introduction to Double/Debiased Machine Learning
This paper provides an introduction to Double/Debiased Machine Learning (DML). DML is a general approach to performing inference about a target parameter in the presence of nuisance functions: objects...
24813
alex hayes @alexpghayes.com · 20/04/2026
thanks so much!
000
alex hayes @alexpghayes.com · 17/04/2026
Was having some bluesky issues, sorry for the multiple responses there!
010
alex hayes @alexpghayes.com · 16/04/2026
Done! Thanks for all your work on Quarto github.com/quarto-dev/q...
020
alex hayes @alexpghayes.com · 16/04/2026
This is far outside my wheelhouse but I'll see if I can get an agent to make some reproducible examples
110
alex hayes @alexpghayes.com · 16/04/2026
Not sure how helpful this is, but here's the commit of Claude/Gemini fixes for axe-core violations on my blog github.com/alexpghayes/...
110
alex hayes @alexpghayes.com · 15/04/2026
i'd love to hear about it!
100
alex hayes @alexpghayes.com · 15/04/2026
currently i'm just copy-pasting the axe-core warnings i get into claude and hoping it knows how to resolve things better than i do it seems like the things claude is doing are fixes patching html / css directly rather than changes to the content of my qmd files
000
alex hayes @alexpghayes.com · 15/04/2026
gotcha. one thing that is hard to tell is how much accessibility failures are coming from choices i'm making vs quarto side things is it helpful to you if i report accessibility warnings i get about my website?
210
alex hayes @alexpghayes.com · 15/04/2026
what's the recommended workflow for checking the accessibility of a #quarto website? currently i'm setting format: html: axe: output: document in _quarto.yml and checking every page visually and it's not particularly efficient #rstats
430
alex hayes @alexpghayes.com · 03/04/2026
I just gave Claude my latex cv and it basically one-shotted recreating it in Typst
100
Reposted by alex hayes
Hadley Wickham @hadley.nz · 20/03/2026
We have summer internships y'all! Come work at Posit on the PyData, tidymodels, shiny, or Connect teams: grnh.se/tigz810a3us. You will have an awesome time, learn a ton, and help advance our open source and pro tools 🧰 #rstats #pydata
Kermit the frog screaming with excitement
27245
alex hayes @alexpghayes.com · 19/03/2026
A lot of Imbens' and Athey's recent work is on estimators for panel data that generalize diff-in-diff, especially in the context of 2+ time periods In general though that literature moves faster than I can keep up with
120
alex hayes @alexpghayes.com · 19/03/2026
See also www.tandfonline.com/doi/full/10.... Worth noting that changes score = diff-in-diff, and there's a huge literature on diff-in-diff
tandfonline.com
ANCOVA versus Change Score for the Analysis of Two-Wave Data
There is a longstanding debate on whether the analysis of covariance (ANCOVA) or the change score approach is more appropriate when analyzing non-experimental longitudinal data. In this article, we...
591
alex hayes @alexpghayes.com · 16/03/2026
"slides here" link on paulgp.com/2018/04/30/b... is dead for me
paulgp.com
Beamer Tips
111
alex hayes @alexpghayes.com · 14/03/2026
"If you give the same writing advice three times, make a Claude skill about it" is probably a reasonable new heuristic... ...and especially when working with folks who need lots of feedback
021
alex hayes @alexpghayes.com · 14/03/2026
Claude skills for reviewing manuscripts that I would appreciate: - stylistic advice for math writing (punctuating eqs, etc) - academic tropes (gaps, see quote) and linguistic oddities (gerund form) to avoid - structure of a good abstract/introduction - etc
270