Reposted by Christopher AkikiMarco @mcognetta.bsky.social · 30/09/2026🚨 [Token][ization] Paper Alert 🚨 Tokenization is a wildly understudied area of language modeling despite it having effects across all of NLP. Over the past ~8 months, 32 (!) tokenizer researchers put together the most comprehensive survey of the field. Check it out! 112333
Reposted by Christopher AkikiLichess @lichess.org · 20/07/2026Were you watching the World Cup or playing chess? Well, we have the stats for it... (1/2) 1357
Christopher Akiki @cakiki.bsky.social · 17/07/2026t-SNE on SigLIP embeddings of 1,360 @oreilly.bsky.social book cover animals. The penguin cluster is my favorite. 020
Christopher Akiki @cakiki.bsky.social · 14/07/2026This version has 1360 @oreilly.bsky.social book covers, sorted by color. 020
Christopher Akiki @cakiki.bsky.social · 11/07/2026Image quilt of all @oreilly.bsky.social animals. 133
Reposted by Christopher AkikiLichess @lichess.org · 26/06/2026We recently crossed 6 million chess puzzles in our open database. Like the platform itself, Lichess data is free and open source. Go and build something cool with it, Available wherever you get your datasets! 5457
Christopher Akiki @cakiki.bsky.social · 22/06/2026work in progress: 1 million synthetic personas from NVIDIA's Nemotron-Personas-USA dataset. 020
Christopher Akiki @cakiki.bsky.social · 04/06/2026Been trying to reproduce datashader functionality using only Apache Arrow and Acero as a learning exercise. This is a 1-billion point Clifford attractor rendered in ~8s from a 9GB parquet file. (No JIT) 010
Reposted by Christopher AkikiMarco @mcognetta.bsky.social · 11/02/2026I'm looking for 5-10 #chess players to test out a tool I'm building. Preferably who play on @lichess.org and are 1200+ in rapid or blitz. And if you coach chess at all, I'd be extra grateful to have you test it! NOTE: it is _NOT_ an "LLM chess coach" tool, I promise! 🙏 264
Reposted by Christopher AkikiLichess @lichess.org · 07/02/2026Any chess position with 8 pieces on the board and at least one pair of opposed pawns has been solved! Lichess can now tell you definitively if it's a win, loss or draw with no engine required. Read about the massive technological undertaking to accomplish this partial 8 piece tablebase on our blog: 24010
Christopher Akiki @cakiki.bsky.social · 20/01/2026UMAP connectivity plots of 3,627 chess openings from the @lichess.org datasets (huggingface.co/datasets/Lic...) 161
Christopher Akiki @cakiki.bsky.social · 18/01/2026Which colormap do you think looks the nicest? I'm leaning toward plasma. 300
Christopher Akiki @cakiki.bsky.social · 16/01/2026Scatterplot of 4 million computer science authors, laid out according to co-authorship connections. Large blob in the bottom left are all single authors; removing them lets the plot breathe more somehow. The source of the data is the @dblp.org bibliography. 142
Reposted by Christopher AkikiShayne Longpre @shaynelongpre.bsky.social · 26/11/2025Who is winning the open AI race? Our new study Economies of Open Intelligence maps @hf.co 851k models' downloads 2020→2025. 1) Power rebalance: US tech ↓; China + community ↑ 2) Models size & efficient ↑ (MoE, quant, multimodal) 3) Intermediary layers ↑ (adapters/quantizers) 4) Transparency ↓ /🧵 293
Reposted by Christopher AkikiLichess @lichess.org · 11/11/2025Researchers at Google DeepMind used our free puzzle database and reinforcement learning to train a model to generate creative chess puzzles. ➡️ Read more on this by Tom Zahavy from the DeepMind discovery team: lichess.org/@/tomas135/b...lichess.orgAI-Generated Chess PuzzlesA new research by the Discovery team at @GoogleDeepMind using RL and generative models to discover creative chess puzzles 0123
Christopher Akiki @cakiki.bsky.social · 04/11/2025Three different ways to represent colo(u)r. Work in progress, inspired by an old post by Kat Zhang / The Poet Engineer. 151
Christopher Akiki @cakiki.bsky.social · 31/10/2025I made this annotated scatter plot of 1 million FineWeb-Edu documents for @sashamtl.bsky.social's new TED talk. 141
Reposted by Christopher AkikiXiaoyi @cleefouti.bsky.social · 28/10/2025When the fish left the river: 215046
Christopher Akiki @cakiki.bsky.social · 27/10/2025Also really love how organic the plot looks with "inferno" (left) and "viridis" (right). 041
Christopher Akiki @cakiki.bsky.social · 26/10/2025Thanks to @jamesabednar.bsky.social I realized I had used the wrong background color for the colormap I had chosen. This is another version of the plot (different embeddings) with the corrected background. 110
Christopher Akiki @cakiki.bsky.social · 28/09/2025526.9 million player deaths in 24.7 million levels of Super Mario Maker 2. Data by @tgr.bsky.social 050
Christopher Akiki @cakiki.bsky.social · 11/07/2025Really cool new embeddings exploration tool by @domoritz.de and colleagues from Apple. Can't wait to build with this. Also includes a streamlit component and a Jupyter widget. 120
Christopher Akiki @cakiki.bsky.social · 28/02/2025Woah! EA just open sourced "Command and Conquer: Red Alert" and a bunch of other CnC games! github.com/electronicar... 121
Reposted by Christopher AkikiLichess @lichess.org · 02/02/2025Lichess is now on @kaggle.com! Use our puzzles, openings, and engine evaluation datasets directly in your kaggle notebooks: www.kaggle.com/organizations/lichess ♟️ 1524
Christopher Akiki @cakiki.bsky.social · 08/12/2024The folks at Foursquare released a @hf.co dataset of 104.5 million places of interest and here's all of them plotted using datashader 2173
Christopher Akiki @cakiki.bsky.social · 06/12/2024I recently used the @lichess.org puzzles dataset to experiment with chess position embeddings and visualize 4.5M starting positions. (hf.co/datasets/Lic...) 2284
Reposted by Christopher AkikiLichess @lichess.org · 06/12/2024The Lichess database of games, puzzles, and engine evaluations is now on @hf.co - huggingface.co/Lichess. Billions of chess data points to download, query, and stream and we're excited to see what you'll build with it! ♟️ 🤗 39423
Christopher Akiki @cakiki.bsky.social · 13/02/2024Early experiment visualizing of Cohere For AI's newly-released Aya dataset. Multilingual corpora are always so fun to play with. 020
Christopher Akiki @cakiki.bsky.social · 26/09/2023835 languages. 3.5 million bible verses. Work in progress. 140
Christopher Akiki @cakiki.bsky.social · 26/09/2023UMAP connectivity graphs—with edgehammer bundling—are always something to gaze at. 030
Christopher Akiki @cakiki.bsky.social · 25/09/2023Revisiting John Williamson's prime factors plot with a few differences in implementation. I am using UMAP and Datashader to visualize the first million integers. Not quite there yet. 030
Christopher Akiki @cakiki.bsky.social · 05/06/2023Code Dataset Visualization—11.66 million files from the Stack, a dataset sourced from permissively-licensed GitHub repositories spanning 86 programming languages (StarCoder languages subset). 1143