Sign in

Jan Vanhove

@janhove.bsky.social
248 followers 86 following 130 posts

Linguist, statistician, and saxophonist of sorts. Senior lecturer with the Department of Multilingualism, Fribourg (CH). I offer support and training in quantitative research methods. 🏠 janhove.github.io 🎷 beatmoustache.ch

PostsRepliesMedia
Reposted by Jan Vanhove
Charlie Stross @cstross.bsky.social · 10h
Oh my: LibreOffice just made "no AI" one of their official features: blog.documentfoundation.org/blog/2026/09...
blog.documentfoundation.org
Yes, no AI is now a feature - TDF Community Blog
One of the comments on the announcement of LibreOffice 26.8 was a simple question: “So, no AI is now a feature?” LibreOffice does not reject artificial intelligence out of hand, but an office suite us...
3532251213
Reposted by Jan Vanhove
Shravan Vasishth @shravanvasishth.bsky.social · 01/10/2026
Applications are open for the 2027 summer school on statistical methods for linguistics and psychology: All details are here: smlp.science
smlp.science
The Eleventh Summer School on Statistical Methods for Linguistics and Psychology
01718
Jan Vanhove @janhove.bsky.social · 29/09/2026
I'll blog around the thesis some other time to break down what the abstract says into something that I could have parsed a year ago.
Abstract: The sliced Wasserstein distance is a fairly recently developed metric between probability distributions. In this Master’s thesis, I investigate how it can be used in conjunction with Gaussian
process regression to solve regression problems in which the inputs are represented by prob-
ability distributions. After spelling out the theoretical foundations (Wasserstein and sliced
Wasserstein distances, positive-definite kernels, and Gaussian process regression), I present a
toy example that demonstrates that a workflow incorporating sliced Wasserstein distances can
capitalise on distributional differences in the inputs that remain invisible to methods relying
solely on marginal distributions and bivariate correlations. I then analyse how global affine
transformations of the input distributions affect sliced Wasserstein distances and the kernels
derived from them. In particular, I show that different whitening procedures produce the
same sliced Wasserstein distances. Asymptotic analyses reveal that extremely anisotropic
transformations can lead to degenerate kernel matrices, but that hyperparameter tuning can
offset this degeneracy. Simulations and a real-data example corroborate these findings and
further suggest that kernels based on Wasserstein distances along the cardinal axes can match or
even outperform sliced Wasserstein distances, when the relevant differences among the input
distributions are approximately aligned with the coordinate axes
000
Jan Vanhove @janhove.bsky.social · 29/09/2026
Master's degree in statistics and data science: completed. If anyone's interested in my thesis on "sliced Wasserstein distances and their application to regression problems", you can find it here: github.com/janhove/msct....
152
Jan Vanhove @janhove.bsky.social · 28/09/2026
Six years after enrolling as a B.Sc. student in maths and CS, I'll receive my master's degree in stats and data science tomorrow. I'll blog about my M.Sc. thesis when I get round to it, but in the mean time, here's my B.Sc. thesis: janhove.github.io/CDCL_Vanhove....
Title of B.Sc. thesis: SAT solving using conflict-driven clause learning
and its application to classical planning.Abstract: The Boolean satisfiability (SAT) problem asks if, given a propositional formula, there is some
way to assign truth values to its variables that render the formula true. SAT is NP-complete, so
a general efficient algorithm for solving it may not exist. However, SAT problems encountered
in industrial applications tend to exhibit some structure that make them efficiently solvable in
practice. This Bachelor’s thesis presents a popular approach to solving such SAT problems that is
based on so-called conflict-driven clause learning (CDCL). It also discusses a handful of tweaks
commonly implemented, in some form or other, in current CDCL-based SAT solvers. Finally, it
explains how planning problems can be solved using SAT solvers and illustrates this process
using two types of examples.An illustration of how the algorithm works.
120
Jan Vanhove @janhove.bsky.social · 18/09/2026
New blog post: "Cluster analysis: A skeptic’s guide". In which I ask a few questions that social scientists wishing to identify hidden classes in their data ought to address. janhove.github.io/posts/2026-0...
A two-dimensional scatterplot showing a three-cluster solution.

Caption: "Figure 2: A visualisation of a Latent Profile Analysis fit. Such visualisations may help readers appreciate that the clusters aren’t nicely separated and that the researchers’ notion of clusters may not correspond to their own."Table of contents:

Refresher: What is cluster analysis?
What are the clusters for?
What clusters, exactly?
Does the pipeline work?
What does the solution look like?
When running follow-up analyses, how is the uncertainty in the cluster assignments taken into account?
Was it worth it?
Conclusion
References
416146
Reposted by Jan Vanhove
Jan Vanhove @janhove.bsky.social · 14/09/2026
I wrote a thing: "On the justification of sampling-based statistics in randomised experiments." It's geared towards non-statisticians in the quantitative social sciences, and I'd love to get some feedback! janhove.github.io/justificatio...
When popular statistical tools such as the $t$-test are taught, the narrative is one in which the data are sampled randomly from some population. Real studies employing random sampling are, however, much rarer than the ubiquity of these tools would suggest, and students may wonder what justifies their use in the absence of random sampling or even of a clear notion of what the population might be. Reassuringly, some of these tools can be understood in an altogether
more common setting: experiments in which units are randomly assigned to conditions. This perspective clarifies why familiar procedures such as the $t$-test tend to perform reasonably well in randomised experiments, but also underscores a difference in the scope of inference afforded by random sampling versus random assignment. Furthermore, this perspective foregrounds flexible randomisation-based methods, whose validity follows from the study design and which, as I will suggest, represent a natural starting point for teaching statistical inference to non-mathematicians.
21512
Jan Vanhove @janhove.bsky.social · 14/09/2026
I wrote a thing: "On the justification of sampling-based statistics in randomised experiments." It's geared towards non-statisticians in the quantitative social sciences, and I'd love to get some feedback! janhove.github.io/justificatio...
When popular statistical tools such as the $t$-test are taught, the narrative is one in which the data are sampled randomly from some population. Real studies employing random sampling are, however, much rarer than the ubiquity of these tools would suggest, and students may wonder what justifies their use in the absence of random sampling or even of a clear notion of what the population might be. Reassuringly, some of these tools can be understood in an altogether
more common setting: experiments in which units are randomly assigned to conditions. This perspective clarifies why familiar procedures such as the $t$-test tend to perform reasonably well in randomised experiments, but also underscores a difference in the scope of inference afforded by random sampling versus random assignment. Furthermore, this perspective foregrounds flexible randomisation-based methods, whose validity follows from the study design and which, as I will suggest, represent a natural starting point for teaching statistical inference to non-mathematicians.
21512
Jan Vanhove @janhove.bsky.social · 08/09/2026
TIL dat "warm en koud blazen" Belgisch-Nederlands is. Waarvoor dank, @jeroenreygaert.bsky.social. Wel oppassen met de letterlijke vertaling naar het Duits.
000
Jan Vanhove @janhove.bsky.social · 01/09/2026
Conjecture: If a paper's title emphasises the statistical tool used, chances are that the study's just an excuse for the authors to use a fancy tool.
100
Jan Vanhove @janhove.bsky.social · 27/08/2026
New blog post: On superficially similar research questions. janhove.github.io/posts/2026-0...
janhove.github.io
On superficially similar research questions – Jan Vanhove
2132
Jan Vanhove @janhove.bsky.social · 22/08/2026
Optimist: The cup is half full. Pessimist: The cup is half empty. Topologist: The cup is fumpty.
000
Jan Vanhove @janhove.bsky.social · 17/08/2026
The draw_boxplots() function in this post is similar in spirit to the plot_r() function from the 10-year-old post "What data patterns can lie behind a correlation coefficient?". Both allow you to create random Anscombe-like lineups that highlight the need for plotting your data. #rstats
Sixteen scatterplots with wildly different patterns that all show 50 observations of two variables correlated at r = 0.5.
0102
Jan Vanhove @janhove.bsky.social · 17/08/2026
New blog post: Different distributions, same boxplot: janhove.github.io/posts/2026-0... In which I share functions for generating markedly different distributions yielding identical boxplots. Overlaying the individual data points whenever possible can prevent many misinterpretations.
Six quite different distributions that are depicted as identical boxplots.
030
Jan Vanhove @janhove.bsky.social · 12/08/2026
When anova() doesn't do ANOVA: A short post on what R's anova() function actually does when you use it to compare mixed-effects models, why you should take its output with a pinch of salt, and how you can obtain more accurate results. janhove.github.io/posts/2026-0...
The black curve shows the distribution function of the 𝜒² distribution with 8 degrees of freedom, which is the asymptotic distribution of the likelihood ratio test statistic in our example. The red curve shows the estimated distribution function of the test statistic under the null distribution. The difference is quite substantial. The test statistic that was actually observed is highlighted by the dashed vertical line.
000
Jan Vanhove @janhove.bsky.social · 11/08/2026
Ah well. Time to rewrite my lecture notes on data and code sharing.
010
Jan Vanhove @janhove.bsky.social · 05/08/2026
"A random variable P is called an exact p-value if Prob(P ≤ α | Hₒ) ≤ α for each α ∈ (0,1)." Most general definition. Doesn't explain how you can construct such P, why some constructions are more useful than others, what approximate p-values are, or why anyone might care, in fairness.
020
Jan Vanhove @janhove.bsky.social · 26/07/2026
Michel Wuyts
000
Jan Vanhove @janhove.bsky.social · 13/07/2026
A new blog post in which I discuss the strong and weak null hypotheses in randomised experiment and how they can be tested: janhove.github.io/posts/2026-0... These tests don't rely on sampling assumptions but on more easily verifiable assignment assumptions.
janhove.github.io
Testing strong and weak null hypotheses in randomised experiments – Jan Vanhove
051
Reposted by Jan Vanhove
👻 Jack Djinn Wail 🧞 @jacktindale.bsky.social · 07/07/2026
STOP INTERVIEWING RANDOM MEMBERS OF THE PUBLIC WHO ARE HANGING AROUND THE TOWN CENTRE IN THE MIDDLE OF THE AFTERNOON YOU MAY AS WELL ASK A CAT FOR HOW REFLECTIVE OF PUBLIC OPINION IT IS www.bbc.co.uk/news/article...
bbc.co.uk
What do Farage's Clacton constituents think of his resignation?
People living in Clacton share contrasting feelings of frustration and support towards Nigel Farage.
2841175
Jan Vanhove @janhove.bsky.social · 03/06/2026
"When interaction effects are hard to explain, try nested effects." janhove.github.io/posts/2026-0...
010
Jan Vanhove @janhove.bsky.social · 29/05/2026
Sometimes, you just know that the line must go up. New blog post: A quick introduction to isotonic regression. janhove.github.io/posts/2026-0...
010
Jan Vanhove @janhove.bsky.social · 17/05/2026
I quite often give students a few arbitrarily selected excerpts in which the article they've read is cited and ask them to compare what was done in the article vs what they would've thought was done in the article based only on the reference. Makes for interesting discussions.
181
Jan Vanhove @janhove.bsky.social · 24/02/2026
Just out: "Does multilingualism really protect against accelerated ageing? Critical comments on Amoruso et al. (2025)" www.tandfonline.com/doi/full/10....
tandfonline.com
Does multilingualism really protect against accelerated ageing? Critical comments on Amoruso et al. (2025)
In a much-publicised large-scale study, [Amoruso, Lucia, Hernan Hernandez, Hernando Santamaria-Garcia, Sebastian Moguilner, Agustina Legaz, Pavel Prado, Jhosmary Cuadros, et al. 2025. “Multilingual...
000
Jan Vanhove @janhove.bsky.social · 06/01/2026
My criticism of Amoruso et al.'s "Multilingualism protects against accelerated aging" study in preprint form: osf.io/45xwm/files/...
111
Reposted by Jan Vanhove
David Leavitt @davidleavitt.bsky.social · 04/01/2026
The FIFA Peace prize used to mean something
521158176
Jan Vanhove @janhove.bsky.social · 15/12/2025
A recent study purports to have found that multilingualism protects against accelerated ageing. I've taken a closer look at it, and it doesn't look good. New blog post: "Does multilingualism really protect against accelerated ageing? Some critical comments" janhove.github.io/posts/2025-1...
Positive trend between monolingualism and accelerated ageing.Negative trend between per capita GDP and accelerated ageing.
56427
Jan Vanhove @janhove.bsky.social · 09/11/2025
New blog post: The population model and the randomisation model of statistical inference janhove.github.io/posts/2025-1...
Figure 5: The distribution of the mean difference under the sharp null hypothesis for 
 for the Brasília replication of the gambler’s fallacy.
000
Jan Vanhove @janhove.bsky.social · 10/10/2025
New blog post: Clarifying research questions by sketching possible outcomes janhove.github.io/posts/2025-1...
092
Jan Vanhove @janhove.bsky.social · 10/10/2025
High-level whining: Replace your osf.io view-only links once your paper's accepted.
000
Jan Vanhove @janhove.bsky.social · 08/10/2025
For the record, I've got dibs on "The Reverend Bass and His Flat Priors".
010
Reposted by Jan Vanhove
Jo Wolff @jowolff.bsky.social · 05/10/2025
I wouldn’t normally endorse AI prompts but these are indeed essential for all academics.
610336
Jan Vanhove @janhove.bsky.social · 13/09/2025
"It is well known that horses don't require dentistry (Rohlfing 2025)."
032
Jan Vanhove @janhove.bsky.social · 12/09/2025
Revised and translated to English: The course materials for 'Introduction to quantitative data analysis'. Available from github.com/janhove/Anal....
050
Jan Vanhove @janhove.bsky.social · 12/09/2025
I've updated the lecture notes for the class on Quantitative Methodology - available from github.com/janhove/Quan....
010
Reposted by Jan Vanhove
Nick Brown @steamtraen.eu · 14/11/2024
French wordplay jokes are the best...
041
Reposted by Jan Vanhove
Shravan Vasishth @shravanvasishth.bsky.social · 19/10/2024
Applications are open for the ninth summer school on statistical methods for linguistics and psychology (SMLP 2025), to be held August 25-29, 2025 in Potsdam, Germany: vasishth.github.io/smlp2025/
vasishth.github.io
The Ninth Summer School on Statistical Methods for Linguistics and Psychology
023
Jan Vanhove @janhove.bsky.social · 10/09/2024
New blog post: Exact significance tests for 2 × 2 tables janhove.github.io/posts/2024-0... #stats #rstats
janhove.github.io
Jan Vanhove :: Blog - Exact significance tests for 2 × 2 tables
010
Jan Vanhove @janhove.bsky.social · 04/09/2024
Updated (German-language) teaching materials: Einführung in die quantitative Analyse: janhove.github.io/resources.ht...
janhove.github.io
Jan Vanhove :: Blog - Teaching resources
010
Reposted by Jan Vanhove
Julia M. Rohrer @dingdingpeng.the100.ci · 27/08/2024
New blog post! If you're a substantive researcher, you mainly do statistics to answer substantive questions. So let's stop putting the statistical model first. www.the100.ci/2024/08/27/l...
the100.ci
Let’s do statistics the other way around
Summer in Berlin – the perfect time and place to explore the city, take a walk in the Görli, go skinny dipping in the Spree, attend an overcrowded, overheated conference symposium on cross-lagged pane...
1214663
Jan Vanhove @janhove.bsky.social · 30/07/2024
My Bachelor's thesis (2024): "SAT solving using conflict-driven clause learning and its application to classical planning" janhove.github.io/CDCL_Vanhove...
janhove.github.io
000
Jan Vanhove @janhove.bsky.social · 26/07/2024
Updated teaching materials: Introduction to the general linear model. janhove.github.io/resources.ht...
janhove.github.io
Jan Vanhove :: Blog - Teaching resources
010
Jan Vanhove @janhove.bsky.social · 26/07/2024
New primer: Visualising data in R. janhove.github.io/datasets_gra...
janhove.github.io
Visualising data in R: A primer
000
Jan Vanhove @janhove.bsky.social · 26/07/2024
New primer: Working with datasets in R. janhove.github.io/datasets_gra...
janhove.github.io
Working with datasets in R: A primer
000
Jan Vanhove @janhove.bsky.social · 18/03/2024
For some reason, applied linguists just love formulating three research questions in each paper, regardless of whether they actually have one, three, or five. :)
020