Sign in

statmills.bsky.social

@statmills.bsky.social
37 followers 150 following 11 posts

Data Science, CFB, Abundance Agenda www.statmills.com

PostsRepliesMedia
statmills.bsky.social @statmills.bsky.social · 09/04/2026
In case anyone missed it from earlier this week my latest blog post extends Gradient Boosting to fit just about any model you want.
000
statmills.bsky.social @statmills.bsky.social · 07/04/2026
The result is smooth curves that can learn high dimensional interaction effects that you can fit at scale!
000
statmills.bsky.social @statmills.bsky.social · 07/04/2026
Then fitting our Gradient Boosting Spline model is as simple as calling {DecisionTreeRegressor} in a for loop to build up our coefficient predictions.
100
statmills.bsky.social @statmills.bsky.social · 07/04/2026
{JAX} makes this super easy because all we need to do as practitioners is to define our loss function as a function of our parameters and let autodiff handle calculating the partial derivatives
100
statmills.bsky.social @statmills.bsky.social · 07/04/2026
Gradient Boosting Machines will generate predictions for each observation by iteratively learning to predict the gradient of the loss function at each iteration. In this blog post I show you can use the same methodology to fit entire parameter sets
100
statmills.bsky.social @statmills.bsky.social · 07/04/2026
I have a new blog post out today that I'm really excited about. I walk through how you can use Gradient Boosting to fit entire vectors of parameters for each observation, not just a single prediction. statmills.com/2026-04-06-g... #pydata #rstats
statmills.com
Gradient Boosting Parameters
In this post I’ll walk through how you can use the same principles behind Gradient Boosting Machines (GBM) to predict parameters of models in the same way traditional GBMs predict 1-dimensional target...
142
statmills.bsky.social @statmills.bsky.social · 05/02/2025
There is an additional layer of smoothing you can do with low-rank smoothers, which not only smooths the data but speeds up the processing because we use less parameters!
000
statmills.bsky.social @statmills.bsky.social · 05/02/2025
The smoothing you can do with location data is really interesting! Similar to penalizing neighboring coefficients in a GAM, you can penalize the difference between neighboring census tracts
100
statmills.bsky.social @statmills.bsky.social · 05/02/2025
This allows us to look at some cool maps like where first time home buyers are buying, and where the most expensive homes are
100
statmills.bsky.social @statmills.bsky.social · 05/02/2025
My new blog post explores some open housing data with some interesting (to me) maps. I walked through how to smooth location data using the {mgcv} package with Markov Random Fields for those interested in learning more! statmills.com/2025-02-04-f...
statmills.com
A First Look at Some Atlanta Housing Data
I recently found out that the Federal Housing Authority publishes a ton of granular housing data and wanted to start exploring the data for Atlanta, where I live. I’m not sure there will be anything r...
110
statmills.bsky.social @statmills.bsky.social · 05/02/2025
I hate how it's never the actual data analysis that trips me up when switching between R and Python, it's the silly base stuff like `int` and `len` that takes me multiple tries to switch over 🤬
030