Sign in

hhp9x.bsky.social

@hhp9x.bsky.social
22 followers 418 following 0 posts
PostsRepliesMedia
Reposted by @hhp9x.bsky.social
Andrew Heiss @andrew.heiss.phd · 12/03/2025
New preprint! A general overview of stats in public policy research with this (oversimplified but still helpful) separation of methods into description, explanation, and prediction #policysky HTML/PDF: stats.andrewheiss.com/snoopy-spring/ SocArXiv: doi.org/10.31235/osf...
This essay provides an overview of statistical methods in public policy, focused primarily on the United States. I trace the historical development of quantitative approaches in policy research, from early ad hoc applications through the 19th and early 20th centuries, to the full institutionalization of statistical analysis in federal, state, local, and nonprofit agencies by the late 20th century. I then outline three core methodological approaches to policy-centered statistical research across social science disciplines: description, explanation, and prediction, framing each in terms of the focus of the analysis. In descriptive work, researchers explore what exists and examine any variable of interest to understand their different distributions and relationships. In explanatory work, researchers ask why does it exist and how can it be influenced. The focus of the analysis is on explanatory variables (X) to either (1) accurately estimate their relationship with an outcome variable (Y), or (2) causally attribute the effect of specific explanatory variables on outcomes. In predictive work, researchers as what will happen next and focus on the outcome variable (Y) and on generating accurate forecasts, classifications, and predictions from new data. For each approach, I examine key techniques, their applications in policy contexts, and important methodological considerations. I then consider critical perspectives on quantitative policy analysis framed around issues related to a three-part “data imperative” where governments are driven to count, gather, and learn from data. Each of these imperatives entail substantial issues related to privacy, accountability, democratic participation, and epistemic inequalities—issues at odds with public sector values of transparency and openness. I conclude by identifying some emerging trends in public sector-focused data science, inclusive ethical guidelines, open research practices, and future directions for the field.	Description	Explanation	Prediction
General question	What exists?	Why does it exist? How can it be influenced?	What will happen next?
Focus of analysis	Focus is on any variable—understanding different variables and their distributions and relationships	Focus is on X —understanding the relationship between X and Y, often with an emphasis on causality	Focus is on Y —forecasting or estimating the value of Y based on X, often without concern for causal mechanisms
Names for variable of interest	—		Explanatory variable
	Independent variable
	Predictor variable
	Covariate		Outcome variable
	Dependent variable
	Response variable
Goal of analysis	Summarize and explore data to identify patterns, trends, and relationships	Estimation: Test hypotheses or theories and make inferences about the relationship between one or more X variables and Y
 
Causal attribution: A special form of estimating—make inferences about the causal relationship between a single X of interest and Y through credible causal assumptions and identification strategies	Generate accurate predictions; maximize the amount of explainable variation in Y while minimizing prediction error
Evaluation criteria	—	Confidence/credible intervals, coefficient significance, effect sizes, and theoretical consistency	Metrics like root mean square error (RMSE) and R^2; out-of-sample performance
Typical approaches	Univariate summary statistics like the mean, median, variance, and standard deviation; multivariate summary statistics like correlations and cross-tabulations	t-tests, proportion tests, multivariate regression models; for causal attribution, careful identification through experiments, quasi-experiments, and other methods with observational data	Multivariate regression models; more complex black-box approaches like machine learning and ensemble modelsTable of contents
Introduction
Brief history of statistics in public policy
Core methodological approaches
Description
Explanation
Prediction
The pitfalls of counting, gathering, and learning from public data
Future directions
References
415128
Reposted by @hhp9x.bsky.social
Andreas Krause @arkrause.bsky.social · 17/02/2025
We've released our lecture notes for the course Probabilistic AI at ETH Zurich, covering uncertainty in ML and its importance for sequential decision making. Thanks a lot to @jonhue.bsky.social for his amazing effort and to everyone who contributed! We hope this resource is useful to you!
16410
Reposted by @hhp9x.bsky.social
Martin Huber @causalhuber.bsky.social · 14/02/2025
Just submitted the final proofs of my new book, Impact Evaluation in Firms and Organizations, set to be published by @mitpress.bsky.social this summer! Compared to my first book, #CausalAnalysis, it’s shorter, less technical, and focused on business applications. Here’s the table of contents!
1253
Reposted by @hhp9x.bsky.social
Aaron Sojourner @aaronsojourner.org · 11/02/2025
Former Commissioner of the U.S. Bureau of Labor Statistics on here! Welcome, @ericagroshen.bsky.social! If you care about workers, employers, jobs, and the federal statistical system, an essential follow.
47126
Reposted by @hhp9x.bsky.social
Charlie Gao @shikokuchuo.net · 10/02/2025
For everyone curious to read a bit more on {mirai}, my note on mirai v2.0 is now out on rweekly. Easier distributed computing, cancellation of remote tasks, and the biggest announcement of all: mirai now powers parallel map in purrr #tidyverse. #rstats shikokuchuo.net/posts/25-mir...
shikokuchuo.net
shikokuchuo{net}: mirai v2.0
Continuous Innovation
0177
Reposted by @hhp9x.bsky.social
Charlie Gao @shikokuchuo.net · 06/02/2025
Parallelization just landed in the dev version of purrr: purrr.tidyverse.org/dev/referenc... Really pleased that the mirai framework makes this possible. Huge credit to the tidyverse maintainers @hadley.nz @lionelhenry.bsky.social and @davisvaughan.bsky.social ! #rstats #tidyverse
purrr.tidyverse.org
Parallelization in purrr — parallelization
purrr's map functions have a .parallel argument to parallelize a map using the mirai package. This allows you to run computations in parallel using more cores on your machine, or distributed over the ...
69632