Sign in

Kohei Watanabe

@koheiw.bsky.social
64 followers 16 following 5 posts

I analyze textual data for a living and fun using R packages that I develop. Visit blog.koheiw.net for more details.

PostsRepliesMedia
Kohei Watanabe @koheiw.bsky.social · 06/09/2026
Text analysis is gradually moving away from bag-of-words analysis, so I created a new package for Gaussian mixture topic models blog.koheiw.net?p=2437. I did not not do much C++ programing thanks to the Armadillo package. GMTM is extremely fast thanks to multi-threading. #text-as-data #quanteda
010
Kohei Watanabe @koheiw.bsky.social · 24/05/2025
I released the wordvector package v0.5.0. It is rapidly getting better and different from the original Word2vec package. Please read "Align word vectors of multiple Word2vec models" about the new function blog.koheiw.net?p=2299 #rstats #quanteda
blog.koheiw.net
Align word vectors of multiple Word2vec models
I have been developing a new R package called wordvector since last year. I started it as a fork of the Word2vec package but made several important changes to make it fully compatible with quanteda…
020
Reposted by Kohei Watanabe
Johannes B. Gruber @jbgruber.bsky.social · 19/12/2024
Nice paper showing just *how* irreprodroducible research with proprietary generative LLMs is. Luckily there are open source alternatives (and they are very easy to use too!)
1229
Kohei Watanabe @koheiw.bsky.social · 13/12/2024
A few days ago, I received an email from a researcher asking if text analysis is becoming irrelevant because of AI... blog.koheiw.net?p=2254 #text-as-data #quanteda
010
Kohei Watanabe @koheiw.bsky.social · 23/11/2024
If you think the number of topics, k, is the only important parameter for topic models, you need to read this post and the research paper. blog.koheiw.net?p=2233 I created a new model to optimize the Dirichlet priors to analyze imbalanced corpus more accurately. #rstats #quanteda
blog.koheiw.net
A new topic model for analysis imbalanced corpus
I have been developing and testing a new topic model called model Distributed Asymmetric Allocation (DAA) because latent Dirichlet allocation (LDA) takes a long time to fit to a large corpus but do…
0122