Sign in

Louis Maddox

@permutans.bsky.social
230 followers 134 following 4.1K posts

Combinatorially curious spin.systems

PostsRepliesMedia
Louis Maddox @permutans.bsky.social · 6h
Finished uploading a Wikidata dataset in native Parquet in 3.7% the size of the source JSON dump (total 35GB in native Parquet) huggingface.co/collections/... Source files are split by language since most people probably don't want every possible lang! 📁 Code: github.com/lmmx/wikidat...
README of the wikidata-pq repo

Table shows sizes of the component Wikidata parquet datasets
000
Louis Maddox @permutans.bsky.social · 14h
📄 Learning Multi-Level Features with Matryoshka Sparse Autoencoders (ICML 2025) arxiv.org/abs/2503.17547 📁 dictionary_learning github.com/saprmarks/di...
arxiv.org
Learning Multi-Level Features with Matryoshka Sparse Autoencoders
Sparse autoencoders (SAEs) have emerged as a powerful tool for interpreting neural networks by extracting the concepts represented in their activations. However, choosing the size of the SAE dictionar...
010
Louis Maddox @permutans.bsky.social · 14h
📄 DF-FLOPS: An Alternative to FLOPS Regularization to Effectively Productionize SPLADE-Doc (SIGIR 2025) arxiv.org/abs/2505.150...
arxiv.org
An Alternative to FLOPS Regularization to Effectively Productionize SPLADE-Doc
Learned Sparse Retrieval (LSR) models encode text as weighted term vectors, which need to be sparse to leverage inverted index structures during retrieval. SPLADE, the most popular LSR model, uses FLO...
000
Louis Maddox @permutans.bsky.social · 14h
📄 Generalizing Linear Autoencoder Recommenders with Decoupled Expected Quadratic Loss (2026) arxiv.org/abs/2603.07402 📁 DEQL github.com/coderaBruce/...
arxiv.org
Generalizing Linear Autoencoder Recommenders with Decoupled Expected Quadratic Loss
Linear autoencoders (LAEs) have gained increasing popularity in recommender systems due to their simplicity and strong empirical performance. Most LAE models, including the Emphasized Denoising Linear...
010
Louis Maddox @permutans.bsky.social · 30/09/2026
📄 ULTRA (ICLR 2024) “Towards FMs for KG reasoning” arxiv.org/abs/2310.04562
arxiv.org
Towards Foundation Models for Knowledge Graph Reasoning
Foundation models in language and vision have the ability to run inference on any textual and visual inputs thanks to the transferable representations such as a vocabulary of tokens in language. Knowl...
000
Louis Maddox @permutans.bsky.social · 30/09/2026
📄 SPARQLing Datalog for Rule-Based Reasoning over Large KGs iccl.inf.tu-dresden.de/w/images/c/c... 👨‍🏫 Work from Markus Krötzsch’s group (founded Wikidata, author of Nemo) iccl.inf.tu-dresden.de/web/Markus_K... 🦀 Nemo (2024) github.com/knowsys/nemo 🗓️ Datalog 2.0 2026 sites.google.com/view/datalog...
iccl.inf.tu-dresden.de
000
Louis Maddox @permutans.bsky.social · 29/09/2026
📄 The Geometry of Categorical and Hierarchical Concepts in LLMs (ICLR 2025, Best Paper Award at the ICML 2024 workshop on Mech Interp) arxiv.org/abs/2406.01506 📁 github.com/KihoPark/LLM... LLMs represent categorical concepts as polytopes and hierarchical relations as orthogonality
Figure 1: In the representation spaces of LLMs, hierarchically related concepts (such as plant → animal and mammal → bird) live in orthogonal subspaces, while categorical concepts are represented as polytopes. The top panel illustrates the structure; the bottom panels show the measured representation structure in the Gemma LLM. See Section 5 and Appendix A for details.Figure 5: WordNet noun hierarchy is encoded in the orthogonal structure predicted by statement (a) in Theorem 8. We plot the cosine similarity between a child-parent vector and a parent vector for each feature in the hierarchy (blue). As predicted, this value is close to 0. The left plot uses all data for representation estimation, and the right plot uses only 70% independently selected for each synset. We include baselines where a randomly selected feature is used as the parent (orange) and where the embeddings are shuffled (green) as controls for the possibility that the orthogonality is a simple byproduct of high-dimensional geometry, or of the set inclusion relationships used in estimation—see main text for details. See Appendix F for an analogous plot for statement (d).
100
Louis Maddox @permutans.bsky.social · 29/09/2026
📄 SAGA: Schema-Aware Grounding for Agentic Text-to-SPARQL Generation (July 2026) arxiv.org/abs/2607.14494
arxiv.org
SAGA: Schema-Aware Grounding for Agentic Text-to-SPARQL Generation
Complex knowledge base question answering (KBQA) is commonly approached through either information retrieval over a question-specific subgraph or semantic parsing into an executable logical form. We s...
000
Louis Maddox @permutans.bsky.social · 29/09/2026
📄 Understanding Wikidata Qualifiers: An Analysis and Taxonomy (2026) arxiv.org/abs/2603.11767
A taxonomy of qualifiers on Wikidata

First level splits into context, epistemic/uncertainty, structural, and additional.

Sublevels of context: temporal, spatial, subject modifier

Beneath epistemic: uncertainty quantification

Beneath structural: field of a structure, meta modelling

Beneath additional: sequence, provenance, causality, object/subject statement relation, subproperty, external entity description, ‘other’
100
Louis Maddox @permutans.bsky.social · 29/09/2026
Who called it knowledge graph query completion and not “now draw the rest of the OWL”
000
Louis Maddox @permutans.bsky.social · 29/09/2026
oh there’s just unadulterated Claudeslop on arXiv now huh
What we concede. The general shape is partly a consequence of the design. Backward traversal from 64 seeds through a multiparent DAG will broaden and then narrow, and one should not be impressed by broadening and narrowing as such.
What survives. The hourglass is not one shape; it is a family of shapes that differ node by node, and the differences are informative even if the general form of the family is not. Gauss, Hilbert, and Euler show thickness ratios of 15.1, 18.5, and 14.6, all constriction-and-expansion profiles. The pattern is a property of hub nodes in general, not of Leibniz specifically, and that is precisely the point: the tracer-set explanation predicts a global shape, but the variation between nodes is what carries information.
Newton, on the same graph under the same traversal with the same seeds, shows no constriction at all.
If the hourglass were purely a design artifact, Newton would have one too.
What distinguishes Leibniz within the family is not the ratio, where Hilbert is higher, but the combination of a high ratio with a large absolute path count: 47 of 64 lineages, against four for Newton. Ratio and volume are separately obtainable from the design; their conjunction is what we are pointing at.
What would settle it. The natural test is the one flagged in Section 5.2: a systematic comparison of thickness ratio and path volume across all high-traffic nodes in each century, so that the position of Leibniz within the distribution is visible rather than asserted. Our deposited data support this and we have not done it.
0120
Louis Maddox @permutans.bsky.social · 29/09/2026
Who called it dead internet theory and not asemantic web
000
Louis Maddox @permutans.bsky.social · 29/09/2026
📄 Schemas for graph data (2026) dl.acm.org/doi/pdf/10.1...
dl.acm.org
012
Reposted by Louis Maddox
David Mimno @dmimno.bsky.social · 29/09/2026
Sparsity is back! Sparsity is everywhere in language, but managing it requires overhead. For the last ~15 years it was faster to just do the dense operation, knowing most of it was useless. That sparse operations work is a huge shift in the "do something clever" vs. "spend more money" tradeoff.
2305
Louis Maddox @permutans.bsky.social · 29/09/2026
> pyarrow’s iter_batches returns an invalid nested batch here, a known weakness when it slices deep list-of-struct columns > Found the bug: RecordBatch.cast corrupts the nested struct … [bisects] … > Found it: the offset gets applied twice to the null-typed child — Opus 5.5 🐐
000
Louis Maddox @permutans.bsky.social · 29/09/2026
📄 Common Foundations for Recursive Shape Languages (2026) arxiv.org/abs/2604.20946
arxiv.org
Common Foundations for Recursive Shape Languages
As schema languages for RDF data become more mature, we are seeing efforts to extend them with recursive semantics, applying diverse ideas from logic programming and description logics. While ShEx has...
010
Louis Maddox @permutans.bsky.social · 29/09/2026
📄 Extraction of Validating Shapes from Very Large Knowledge Graphs (2023) www.vldb.org/pvldb/vol16/... 📁 QSE github.com/dkw-aau/qse Efficiently mines validating shapes [SHACL or ShEx] from very large existing KGs [English Wikidata] with support and confidence thresholds
vldb.org
001
Louis Maddox @permutans.bsky.social · 29/09/2026
📄 Characteristic sets (Neumann & Moerkotte, 2011) www.csd.uoc.gr/~hy561/paper... They define the characteristic set of a subject as the set of predicates it emits. Entities that share that set tend to be semantically alike, which gives RDF the "latent soft schema" it lacks on paper.
page 3 of the paper

V. CHARACTERISTIC SETS
A. Plain Characteristic Sets
Standard histograms do not capture correlations between join predicates at all, and even using dependent selectivities. as shown in the previous section. only helps for the first join. Therefore, we propose a radically different approach for estimating the selectivity of joins in RDF graphs.
Many of the correlations we observe during selectivity estimation stem from the fact that RD uses multinle trioles to describe the same object. Consider the following sample triples describing a book:
(os.title, The Tree and I). (o1.author,R. Pecker).
(01.author,D. Owl), (01-year, 1996).
All four triples describe the same entity, and accordingly. the individual triple patterns (derived by replacing o, with a variable ?b) are strongly correlated. The year is somewhat an exception here, but for the other three triples, searching just for one triple pattern is nearly as selective as searching for all of them. Obviously, this is true for most books. In general. many entities can be uniquely identified by a true subset of their emitting edges.
In most RDF data sets, these emitting edges exhibit a certain structure. While RDF is used usually without a fixed schema, some kind of latent soft schema in the data frequently occurs:
Books tend to have authors and titles, etc. While we might not be able to clearly classify an entity as "book" (duc to the lack of schema information), we observe that we can characterize an entity by its emitting edges. For each entity s occurring in an RDF data set R. we define its characteristic ser as follows:
Sc(s) := {p|3o: (s,p,o) € R).
For many RDF data sets, entities that have the same characteristic set tend to be semantically similar. This is not surprising, as RDF encodes all semantics using edges. But this enables us to predict selectivities based upon the involved…
000
Louis Maddox @permutans.bsky.social · 29/09/2026
[capture the lightcone voice] “we must compile the TBox”
000
Louis Maddox @permutans.bsky.social · 29/09/2026
to be held is to be carried — Clopus 5.5
000
Louis Maddox @permutans.bsky.social · 29/09/2026
Using a Sonnet model to see the dumb default response to your prompt is MTP drafting for chat UIs lol
000
Louis Maddox @permutans.bsky.social · 29/09/2026
📄 Galkin et al. (EMNLP 2020) Message Passing for Hyper-Relational Knowledge Graphs arxiv.org/abs/2009.10847 📁 StarE github.com/migalkin/StarE
arxiv.org
Message Passing for Hyper-Relational Knowledge Graphs
Hyper-relational knowledge graphs (KGs) (e.g., Wikidata) enable associating additional key-value pairs along with the main triple to disambiguate, or restrict the validity of a fact. In this work, we ...
000
Louis Maddox @permutans.bsky.social · 28/09/2026
One (1) wild brainwave acquired
000
Louis Maddox @permutans.bsky.social · 28/09/2026
📄 Wikontic: Constructing Wikidata-Aligned, Ontology-Aware Knowledge Graphs with Large Language Models (EACL 2026) aclanthology.org/2026.eacl-lo... 📁 github.com/screemix/Wik...
aclanthology.org
Wikontic: Constructing Wikidata-Aligned, Ontology-Aware Knowledge Graphs with Large Language Models
Alla Chepurova, Aydar Bulatov, Mikhail Burtsev, Yuri Kuratov. Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
010
Louis Maddox @permutans.bsky.social · 28/09/2026
📄 The Wikidata Query Logs Dataset (2026) arxiv.org/abs/2602.145... 📁 WDQL github.com/ad-freiburg/...
arxiv.org
The Wikidata Query Logs Dataset
We present the Wikidata Query Logs (WDQL) dataset, a dataset consisting of 335k question-query pairs over the Wikidata knowledge graph. It is over 11x larger than the largest existing Wikidata dataset...
000
Louis Maddox @permutans.bsky.social · 27/09/2026
Grant Sanderson of 3b1b interviewing Jane Street researcher Alok Puranik about positional encodings and group theory youtu.be/k86eUj4hgdc 📝 Using group theory to explore the space of positional encodings for attention blog.janestreet.com/using-group-...
youtu.be
Positional Encodings and Group Theory | 3Blue1Brown and Alok Puranik
YouTube video by Jane Street
010
Louis Maddox @permutans.bsky.social · 27/09/2026
Palimpsest blind decomposition link.springer.com/chapter/10.1... 📁 BIDeN github.com/bhrnprhmd-ha...
link.springer.com
Blind Image Decomposition for Recovering Overlapping Text Layers on Palimpsests
Recovering the scriptio inferior (undertext) from palimpsests remains a significant challenge in digital heritage and historical document analysis due to the severe degradation of the scraped ink and ...
000
Louis Maddox @permutans.bsky.social · 27/09/2026
MambaVision SotA on TextZoom at ICDAR link.springer.com/chapter/10.1...
link.springer.com
A MambaVision-Based Cross-Modal Feature Enhancement Network for Scene Text Super-Resolution
Scene Text Image Super-Resolution (STISR) aims to recover high-resolution textual details from low-resolution scene text images, thereby improving the accuracy of downstream text recognition tasks. Wh...
000
Louis Maddox @permutans.bsky.social · 27/09/2026
so is someone gonna give the NeurIPS programme the AI gen hype video treatment or
000
Louis Maddox @permutans.bsky.social · 27/09/2026
`tpu cuu 5`… Cthulhu coded shell incantation
000
Louis Maddox @permutans.bsky.social · 27/09/2026
Network bandwidth bottlenecked at last
000
Louis Maddox @permutans.bsky.social · 27/09/2026
2026 spoiled user mindset: "how dare this semantically useful info be computationally any less than convenient"
000
Louis Maddox @permutans.bsky.social · 26/09/2026
The amnesty on bugs as you take a library you built from zero into the correct/performant phase is… humbling
110
Louis Maddox @permutans.bsky.social · 26/09/2026
SIMDmaxxing era 🔜 I fear
110
Louis Maddox @permutans.bsky.social · 26/09/2026
is it idiomatic for data normalisation libraries to talk of invariants and determinants? well it is now
110
Louis Maddox @permutans.bsky.social · 25/09/2026
eyeballing 60% of WikiData Swahili values without a label in the language
000
Louis Maddox @permutans.bsky.social · 25/09/2026
Filleting out the entirety of English Wikidata might still change the game (Claude parquetmaxxing self-actualisation incoming)
000
Louis Maddox @permutans.bsky.social · 24/09/2026
Revisited schema inference in genson-core // polars-genson + made big speed gains + memory reductions 🏁 6+ day estimated processing time for Wikidata now <24h
Table of wall times per file for individual Wikidata files falling from 12.9 - 35.5 seconds down to 1.8 - 4.7s
100
Louis Maddox @permutans.bsky.social · 24/09/2026
Uncanny how accurately Claude can estimate the exact wall time change of individual perf optimisations
000
Louis Maddox @permutans.bsky.social · 23/09/2026
Packaged this, after some tape allocation adjustments it parses up to 3x faster than simd-json, gains most on files with large rows (Wikidata claims benchmark) 📁 cuJSON-rs github.com/lmmx/cuJSON-rs 🦀 Rust crates.io/crates/cujson 🐍 Python pypi.org/project/cujs...
Benchmarks showing speedup from simd-json (1.3-3 GB/s) to cujson (4-4.5 GB/s)
020
Louis Maddox @permutans.bsky.social · 23/09/2026
Claude yearns to tape rotate
000
Louis Maddox @permutans.bsky.social · 23/09/2026
Accidentally stumbled on what seems to be the same thing Claude Code (etc) use for displaying markdown: `cargo binstall mdcat-ng` (gives you mdcat and mdless)
110
Louis Maddox @permutans.bsky.social · 23/09/2026
📄 CuJSON: A Highly Parallel JSON Parser for GPUs (ASPLOS '26) dl.acm.org/doi/pdf/10.1... 📁 cuJSON github.com/AutomataLab/...
Bar charts of RapidJSON vs simdjson vs cuJSON performance (time in ms) on various queries
000
Louis Maddox @permutans.bsky.social · 23/09/2026
Concept: intelligent memory allocator called smartalloc
000
Louis Maddox @permutans.bsky.social · 23/09/2026
OK enough cute Opus visuals time to put some TB on the GPU
100
Louis Maddox @permutans.bsky.social · 23/09/2026
📝 hypothesmith: Hypothesis strategies for generating Python programs, something like CSmith github.com/Zac-HD/hypot...
github.com
GitHub - Zac-HD/hypothesmith: Hypothesis strategies for generating Python programs, something like CSmith
Hypothesis strategies for generating Python programs, something like CSmith - Zac-HD/hypothesmith
000
Louis Maddox @permutans.bsky.social · 21/09/2026
nvm it was only CUDA :')
000
Louis Maddox @permutans.bsky.social · 21/09/2026
Anyone else been getting the system reminder saying ABANDON ALL HOPE
000
Louis Maddox @permutans.bsky.social · 21/09/2026
ah now to ship my perfectly working softwSEGV SEGV SEGV SEGV SEGV
000
Louis Maddox @permutans.bsky.social · 20/09/2026
Snake feels like a poor demo for Jev (in that it makes the latency conspicuous)
000