Sign in

Louis Maddox

@permutans.bsky.social
231 followers 134 following 4.1K posts

Combinatorially curious spin.systems

PostsRepliesMedia
Louis Maddox @permutans.bsky.social · 2h
“Search 100M+ papers to find the strongest evidence for and against, for any scientific claim” paperclip.gxl.ai/evidence-api
paperclip.gxl.ai
Paperclip
Paperclip — search, read, and analyze 11M+ biomedical papers, regulatory documents, clinical trials, and biological databases (UniProt, PDB, ChEMBL) from the command line.
000
Louis Maddox @permutans.bsky.social · 14h
🗺️ To celebrate Polars 2.0 I cut a v1 release of polars-genson with support for its new Map data type, and made this explainer video using the example of Wikidata language labels, which'd be rather awkward as a struct 🐻‍❄️ github.com/lmmx/polars-... 📃 Docs: polars-genson.vercel.app
000
Louis Maddox @permutans.bsky.social · 06/10/2026
The greater the source set size of the analytical residue, the greater the regret at not having taken a backup on deleting it to make space for more data gen… probably a lesson in there
000
Louis Maddox @permutans.bsky.social · 06/10/2026
Core memory: becoming obsessed with this Wikipedia entry
Rovibronic coupling

Rovibronic coupling, also known as rotation/vibration-electron coupling, denotes the simultaneous interactions between rotational, vibrational, and electronic degrees of freedom in a molecule.[1] When a rovibronic transition occurs, the rotational, vibrational, and electronic states change simultaneously, unlike in rovibrational coupling. The coupling can be observed using spectroscopy, and is most easily seen in the Renner–Teller effect in which a linear polyatomic molecule is in a degenerate electronic state and bending vibrations will cause a large rovibronic coupling.
000
Reposted by Louis Maddox
Samuel @samuel.fm · 07/03/2025
dashboard of jetstream on 3€ hetzner vps my beloved
Grafana dashboard showing firehose stats from jetstream
91248
Louis Maddox @permutans.bsky.social · 06/10/2026
there’s a BlueSky firehose?
000
Louis Maddox @permutans.bsky.social · 05/10/2026
Concept: tqdm but it's able to take into account the relative sizes of the iterables
000
Louis Maddox @permutans.bsky.social · 05/10/2026
⏸️ 23x faster nested parquet field reading achieved internally (46s ⇢ 2s)
(wikidata) louis 🌟 ~/lab/wikidata $ qp scripts/refs_bench.py ~/tmp/claims-sample.parquet                  
claims-sample.parquet: 529 MB; the needed leaves:                                                          
  property                                                                                                 
  datavalue.id                                                                                                                                                                                                        
  datavalue.unit
  references.list.element.snaks.list.element.value.list.element.property
  references.list.element.snaks.list.element.value.list.element.datavalue.id                               
  references.list.element.snaks.list.element.value.list.element.datavalue.unit                             
  qualifiers.list.element.value.list.element.property                                                      
  qualifiers.list.element.value.list.element.datavalue.id                                                  
  qualifiers.list.element.value.list.element.datavalue.unit                                                
1. every column, Polars: 46 s, peak 22.3 GiB, 871,948 refs                                                 
2. needed leaves, pyarrow, Polars explode (file_refs): 29 s (read 1 s, explode and unique 29 s), peak 17.4 GiB, 871,948 refs; same refs                                                                               
3. needed leaves, pyarrow memory-mapped, Polars explode: 29 s (read 1 s, explode and unique 28 s), peak 17.4 GiB, 871,948 refs; same refs                                                                             
4. needed leaves, pyarrow, pyarrow flatten: 2 s (read 1 s, flatten 0 s, refs and unique 2 s), peak 6.8 GiB, 871,948 refs; same refs
010
Louis Maddox @permutans.bsky.social · 05/10/2026
ooh there’s actually a pyarrow backend which by default memory maps… docs.pola.rs/api/python/s...
Workaround: Use PyArrow backend
If true leaf-level Parquet projection is critical for performance (to avoid decoding entire nested structures), you can use the use_pyarrow=True option with pyarrow_options to pass PyArrow-style column paths, then convert to
polars.

df = pl. read_parquet (
    "path/to/file.parquet",
    use_pyarrow=True,
    pyarrow_options={"columns": ["main.property", "datavalue. id", "datavalue. unit"}
)
000
Louis Maddox @permutans.bsky.social · 05/10/2026
discovery du jour: > PyArrow's `read(columns=)` can project nested columns/leaves from Parquet at decode time > Polars can select top-level columns via scan_parquet(columns=...), but it doesn't expose PyArrow's fine-grained leaf-level column projection for nested Parquet fields.
Claude chat excerpt:

**The change:** Instead of reading every column from the claims Parquet file with Polars' streaming engine (`file_refs`), they switched to reading only the **"needed leaves"** with PyArrow — the specific nested fields that `file_refs` actually uses:
- `property` 
- `datavalue.id`
- `datavalue.unit`

(These are accessed at any depth in the main snak, qualifiers, and references.)

**The gain:**
- **Time:** 48 s → **30 s** (1.6× faster)
- **Peak RSS:** 21.4 GiB → **17.4 GiB** (almost 4 GiB lower)
- Same 871,948 refs produced, verified

**Why it works:** PyArrow's `read(columns=)` can project nested columns/leaves from Parquet at decode time — you only decompress and convert to Arrow what you actually need, rather than decoding the entire structure and then selecting fields afterward. The row explosions (snaks into refs) still dominate memory usage, but the payload to decode is far smaller.

The benchmark script (`refs_bench.py`) in the journal compares the two methods side-by-side and is now integrated into the pipeline for `file_refs` on the main set.

(can polars not do that?)

Not efficiently for nested structures. Polars can select top-level columns via `scan_parquet(columns=...)`, but it doesn't expose PyArrow's fine-grained **leaf-level column projection** for nested Parquet fields.

From the journal, when they tried `pl.scan_parquet(path).slice(start, n)`, it "apparently [read] far more than the row group, three times (main snaks, qualifiers, references)" — Polars' streaming engine was still decoding entire nested structures even when only a few leaves were needed.

PyArrow's `read(columns=["main.property", "datavalue.id", ...])` is a native Parquet feature: the file format supports reading just those leaf paths without touching sibling fields. PyArrow exposes this directly. Polars, built on PyArrow underneath, doesn't expose that capability at the Polars API level — you have to drop down to PyArrow to use it.
100
Louis Maddox @permutans.bsky.social · 05/10/2026
> You can make pq fast for random access with tricks like v small row groups but then scan perf becomes abysmal… AI requires multiple workflows with the same table within the same e2e workflow 🤔 > The files themselves need to be managed by the database petesoder.substack.com/p/for-20-yea...
petesoder.substack.com
For 20 years, you stored a pointer to the file. Chang She thinks those days are over.
Chang She (LanceDB) on why the data infrastructure under every AI product is about to get rebuilt — whether teams are ready or not.
000
Louis Maddox @permutans.bsky.social · 05/10/2026
Chang She on "why he walked away from Parquet to work on LanceDB" youtu.be/slZ1xg2UznQ While great for analytical access, random access for physical AI (like autonomous vehicles) performed poorly, plus complaints about slower than real time processing and data sync pipelines
youtu.be
Chang She on Why He Walked Away from Parquet to Build LanceDB
YouTube video by AI Council
110
Louis Maddox @permutans.bsky.social · 05/10/2026
Waiting for this massive pile of fresh Wikidata to upload
Pablo Escobar waiting in swing set meme
000
Louis Maddox @permutans.bsky.social · 05/10/2026
The thrill of stacking agent-won code speedups make the slow steps inversely ex — cru — ci — a — ting
000
Louis Maddox @permutans.bsky.social · 05/10/2026
[holds your agent TUI] do a breakthrough… for me
000
Reposted by Louis Maddox
Phoronix @phoronix-poster.bsky.social · 05/10/2026
Mold 3.0 High Speed Linker Released Following Rust Rewrite - www.phoronix.com/news/Mold-3.0-Rele…
phoronix.com
Mold 3.0 High Speed Linker Released Following Rust Rewrite
The Mold high speed linker alternative to the likes of GNU Gold/LD and LLVM LLD recently has been going through a rewrite from C++ to Rust as well as having high ambitions of eventually becoming the default linker on Linux systems. Today marks the release of Mold 3.0 for this Rust-based high performance linker...
071
Louis Maddox @permutans.bsky.social · 05/10/2026
Wrote what I twote: cog.spin.systems/docs-drift Termed this idea “freshness pinning”, and found some prior art. It’d be nice to have a MutaGReP-style catalogue of symbol intents mapped to the docs as a staleness signal, but AST/whole file checksums would work too
cog.spin.systems
Docs Drift
Keeping high-level docs in line as code evolves
000
Louis Maddox @permutans.bsky.social · 05/10/2026
I thought “text is the universal interface” meme was dead but the ongoing crashout over TUIs is making me reconsider
000
Louis Maddox @permutans.bsky.social · 04/10/2026
This would act like an expansion of a changelog, or a farsighted sliding window over myopic bursts. It overlaps with development journals (which in turn are internal-facing analogues to repo issues/PR documentation) but is written for user perspective not developer memory jogging
000
Louis Maddox @permutans.bsky.social · 04/10/2026
LLM docs' failure mode is probably something like recency bias (they find the context window disproportionately salient) so one solution would probably be a docs writeup of some kind from scratch on each development session, and then have a docs process take stock in the long run
100
Louis Maddox @permutans.bsky.social · 04/10/2026
Commonsense docs feature that isn’t standard: align the [high level, not API ref] docs to the codebase and record them as having been written at some checksum of each file and (written at) git commit hash (since the primary failure mode is staleness)
100
Louis Maddox @permutans.bsky.social · 02/10/2026
📄 SearchEyes: Towards Frontier Multimodal Deep Search Intelligence via Search World Simulation arxiv.org/abs/2607.05943 📁 github.com/Frostlinx/Se...
Visual abstract diagram of “method architecture” of the SearchEyes system including “Perception-Knowledge Chain Synthesis” and “Self-Contained Search World”
010
Louis Maddox @permutans.bsky.social · 01/10/2026
🪆 Wikidata ID features huggingface.co/spaces/permu... Since Wikidata links items to external catalogues, one way to discover similar items is to encode which families of these catalogues list them (with a Matryoshka SAE), ‘how the world sees things’ rather than just how Wikidata does
Results listings for Kalman filter
000
Louis Maddox @permutans.bsky.social · 01/10/2026
Wikidata induced vertigo: ✅😵‍💫
000
Louis Maddox @permutans.bsky.social · 30/09/2026
Finished uploading a Wikidata dataset in native Parquet in 3.7% the size of the source JSON dump (total 35GB in native Parquet) huggingface.co/collections/... Source files are split by language since most people probably don't want every possible lang! 📁 Code: github.com/lmmx/wikidat...
README of the wikidata-pq repo

Table shows sizes of the component Wikidata parquet datasets
000
Louis Maddox @permutans.bsky.social · 30/09/2026
📄 Learning Multi-Level Features with Matryoshka Sparse Autoencoders (ICML 2025) arxiv.org/abs/2503.17547 📁 dictionary_learning github.com/saprmarks/di...
arxiv.org
Learning Multi-Level Features with Matryoshka Sparse Autoencoders
Sparse autoencoders (SAEs) have emerged as a powerful tool for interpreting neural networks by extracting the concepts represented in their activations. However, choosing the size of the SAE dictionar...
010
Louis Maddox @permutans.bsky.social · 30/09/2026
📄 DF-FLOPS: An Alternative to FLOPS Regularization to Effectively Productionize SPLADE-Doc (SIGIR 2025) arxiv.org/abs/2505.150...
arxiv.org
An Alternative to FLOPS Regularization to Effectively Productionize SPLADE-Doc
Learned Sparse Retrieval (LSR) models encode text as weighted term vectors, which need to be sparse to leverage inverted index structures during retrieval. SPLADE, the most popular LSR model, uses FLO...
000
Louis Maddox @permutans.bsky.social · 30/09/2026
📄 Generalizing Linear Autoencoder Recommenders with Decoupled Expected Quadratic Loss (2026) arxiv.org/abs/2603.07402 📁 DEQL github.com/coderaBruce/...
arxiv.org
Generalizing Linear Autoencoder Recommenders with Decoupled Expected Quadratic Loss
Linear autoencoders (LAEs) have gained increasing popularity in recommender systems due to their simplicity and strong empirical performance. Most LAE models, including the Emphasized Denoising Linear...
010
Louis Maddox @permutans.bsky.social · 30/09/2026
📄 ULTRA (ICLR 2024) “Towards FMs for KG reasoning” arxiv.org/abs/2310.04562
arxiv.org
Towards Foundation Models for Knowledge Graph Reasoning
Foundation models in language and vision have the ability to run inference on any textual and visual inputs thanks to the transferable representations such as a vocabulary of tokens in language. Knowl...
000
Louis Maddox @permutans.bsky.social · 30/09/2026
📄 SPARQLing Datalog for Rule-Based Reasoning over Large KGs iccl.inf.tu-dresden.de/w/images/c/c... 👨‍🏫 Work from Markus Krötzsch’s group (founded Wikidata, author of Nemo) iccl.inf.tu-dresden.de/web/Markus_K... 🦀 Nemo (2024) github.com/knowsys/nemo 🗓️ Datalog 2.0 2026 sites.google.com/view/datalog...
iccl.inf.tu-dresden.de
010
Louis Maddox @permutans.bsky.social · 29/09/2026
📄 Hierarchical Concept Geometry in Language Models Emerges from Word Co-occurrence (ICML 2026) arxiv.org/abs/2605.23821 > the concept-vector orthogonality pattern approximately observed by Park et al. is also reproduced by a purely co-occurrence-driven model.
arxiv.org
Hierarchical Concept Geometry in Language Models Emerges from Word Co-occurrence
We propose a distributional theory of how hypernymy -- the ``is-a'' relation between general and specific concepts -- is encoded geometrically in language representations. Starting from the empiricall...
000
Louis Maddox @permutans.bsky.social · 29/09/2026
📄 The Geometry of Categorical and Hierarchical Concepts in LLMs (ICLR 2025, Best Paper Award at the ICML 2024 workshop on Mech Interp) arxiv.org/abs/2406.01506 📁 github.com/KihoPark/LLM... LLMs represent categorical concepts as polytopes and hierarchical relations as orthogonality
Figure 1: In the representation spaces of LLMs, hierarchically related concepts (such as plant → animal and mammal → bird) live in orthogonal subspaces, while categorical concepts are represented as polytopes. The top panel illustrates the structure; the bottom panels show the measured representation structure in the Gemma LLM. See Section 5 and Appendix A for details.Figure 5: WordNet noun hierarchy is encoded in the orthogonal structure predicted by statement (a) in Theorem 8. We plot the cosine similarity between a child-parent vector and a parent vector for each feature in the hierarchy (blue). As predicted, this value is close to 0. The left plot uses all data for representation estimation, and the right plot uses only 70% independently selected for each synset. We include baselines where a randomly selected feature is used as the parent (orange) and where the embeddings are shuffled (green) as controls for the possibility that the orthogonality is a simple byproduct of high-dimensional geometry, or of the set inclusion relationships used in estimation—see main text for details. See Appendix F for an analogous plot for statement (d).
100
Louis Maddox @permutans.bsky.social · 29/09/2026
📄 SAGA: Schema-Aware Grounding for Agentic Text-to-SPARQL Generation (July 2026) arxiv.org/abs/2607.14494
arxiv.org
SAGA: Schema-Aware Grounding for Agentic Text-to-SPARQL Generation
Complex knowledge base question answering (KBQA) is commonly approached through either information retrieval over a question-specific subgraph or semantic parsing into an executable logical form. We s...
000
Louis Maddox @permutans.bsky.social · 29/09/2026
📁 Code: github.com/cui-ke/wikid...
github.com
GitHub - cui-ke/wikidata-qualifiers: Models and tools for analyzing wikidata qualifiers and using them in queries and inferences
Models and tools for analyzing wikidata qualifiers and using them in queries and inferences - cui-ke/wikidata-qualifiers
000
Louis Maddox @permutans.bsky.social · 29/09/2026
📄 Understanding Wikidata Qualifiers: An Analysis and Taxonomy (2026) arxiv.org/abs/2603.11767
A taxonomy of qualifiers on Wikidata

First level splits into context, epistemic/uncertainty, structural, and additional.

Sublevels of context: temporal, spatial, subject modifier

Beneath epistemic: uncertainty quantification

Beneath structural: field of a structure, meta modelling

Beneath additional: sequence, provenance, causality, object/subject statement relation, subproperty, external entity description, ‘other’
100
Louis Maddox @permutans.bsky.social · 29/09/2026
Who called it knowledge graph query completion and not “now draw the rest of the OWL”
000
Louis Maddox @permutans.bsky.social · 29/09/2026
oh there’s just unadulterated Claudeslop on arXiv now huh
What we concede. The general shape is partly a consequence of the design. Backward traversal from 64 seeds through a multiparent DAG will broaden and then narrow, and one should not be impressed by broadening and narrowing as such.
What survives. The hourglass is not one shape; it is a family of shapes that differ node by node, and the differences are informative even if the general form of the family is not. Gauss, Hilbert, and Euler show thickness ratios of 15.1, 18.5, and 14.6, all constriction-and-expansion profiles. The pattern is a property of hub nodes in general, not of Leibniz specifically, and that is precisely the point: the tracer-set explanation predicts a global shape, but the variation between nodes is what carries information.
Newton, on the same graph under the same traversal with the same seeds, shows no constriction at all.
If the hourglass were purely a design artifact, Newton would have one too.
What distinguishes Leibniz within the family is not the ratio, where Hilbert is higher, but the combination of a high ratio with a large absolute path count: 47 of 64 lineages, against four for Newton. Ratio and volume are separately obtainable from the design; their conjunction is what we are pointing at.
What would settle it. The natural test is the one flagged in Section 5.2: a systematic comparison of thickness ratio and path volume across all high-traffic nodes in each century, so that the position of Leibniz within the distribution is visible rather than asserted. Our deposited data support this and we have not done it.
0110
Louis Maddox @permutans.bsky.social · 29/09/2026
Who called it dead internet theory and not asemantic web
011
Louis Maddox @permutans.bsky.social · 29/09/2026
📄 Schemas for graph data (2026) dl.acm.org/doi/pdf/10.1...
dl.acm.org
012
Reposted by Louis Maddox
David Mimno @dmimno.bsky.social · 29/09/2026
Sparsity is back! Sparsity is everywhere in language, but managing it requires overhead. For the last ~15 years it was faster to just do the dense operation, knowing most of it was useless. That sparse operations work is a huge shift in the "do something clever" vs. "spend more money" tradeoff.
2315
Louis Maddox @permutans.bsky.social · 29/09/2026
> pyarrow’s iter_batches returns an invalid nested batch here, a known weakness when it slices deep list-of-struct columns > Found the bug: RecordBatch.cast corrupts the nested struct … [bisects] … > Found it: the offset gets applied twice to the null-typed child — Opus 5.5 🐐
000
Louis Maddox @permutans.bsky.social · 29/09/2026
📄 Common Foundations for Recursive Shape Languages (2026) arxiv.org/abs/2604.20946
arxiv.org
Common Foundations for Recursive Shape Languages
As schema languages for RDF data become more mature, we are seeing efforts to extend them with recursive semantics, applying diverse ideas from logic programming and description logics. While ShEx has...
010
Louis Maddox @permutans.bsky.social · 29/09/2026
📄 Extraction of Validating Shapes from Very Large Knowledge Graphs (2023) www.vldb.org/pvldb/vol16/... 📁 QSE github.com/dkw-aau/qse Efficiently mines validating shapes [SHACL or ShEx] from very large existing KGs [English Wikidata] with support and confidence thresholds
vldb.org
001
Louis Maddox @permutans.bsky.social · 29/09/2026
📄 Characteristic sets (Neumann & Moerkotte, 2011) www.csd.uoc.gr/~hy561/paper... They define the characteristic set of a subject as the set of predicates it emits. Entities that share that set tend to be semantically alike, which gives RDF the "latent soft schema" it lacks on paper.
page 3 of the paper

V. CHARACTERISTIC SETS
A. Plain Characteristic Sets
Standard histograms do not capture correlations between join predicates at all, and even using dependent selectivities. as shown in the previous section. only helps for the first join. Therefore, we propose a radically different approach for estimating the selectivity of joins in RDF graphs.
Many of the correlations we observe during selectivity estimation stem from the fact that RD uses multinle trioles to describe the same object. Consider the following sample triples describing a book:
(os.title, The Tree and I). (o1.author,R. Pecker).
(01.author,D. Owl), (01-year, 1996).
All four triples describe the same entity, and accordingly. the individual triple patterns (derived by replacing o, with a variable ?b) are strongly correlated. The year is somewhat an exception here, but for the other three triples, searching just for one triple pattern is nearly as selective as searching for all of them. Obviously, this is true for most books. In general. many entities can be uniquely identified by a true subset of their emitting edges.
In most RDF data sets, these emitting edges exhibit a certain structure. While RDF is used usually without a fixed schema, some kind of latent soft schema in the data frequently occurs:
Books tend to have authors and titles, etc. While we might not be able to clearly classify an entity as "book" (duc to the lack of schema information), we observe that we can characterize an entity by its emitting edges. For each entity s occurring in an RDF data set R. we define its characteristic ser as follows:
Sc(s) := {p|3o: (s,p,o) € R).
For many RDF data sets, entities that have the same characteristic set tend to be semantically similar. This is not surprising, as RDF encodes all semantics using edges. But this enables us to predict selectivities based upon the involved…
000
Louis Maddox @permutans.bsky.social · 29/09/2026
[capture the lightcone voice] “we must compile the TBox”
000
Louis Maddox @permutans.bsky.social · 29/09/2026
to be held is to be carried — Clopus 5.5
000
Louis Maddox @permutans.bsky.social · 29/09/2026
Using a Sonnet model to see the dumb default response to your prompt is MTP drafting for chat UIs lol
000
Louis Maddox @permutans.bsky.social · 29/09/2026
📄 Galkin et al. (EMNLP 2020) Message Passing for Hyper-Relational Knowledge Graphs arxiv.org/abs/2009.10847 📁 StarE github.com/migalkin/StarE
arxiv.org
Message Passing for Hyper-Relational Knowledge Graphs
Hyper-relational knowledge graphs (KGs) (e.g., Wikidata) enable associating additional key-value pairs along with the main triple to disambiguate, or restrict the validity of a fact. In this work, we ...
000
Louis Maddox @permutans.bsky.social · 28/09/2026
One (1) wild brainwave acquired
000
Louis Maddox @permutans.bsky.social · 28/09/2026
📄 Wikontic: Constructing Wikidata-Aligned, Ontology-Aware Knowledge Graphs with Large Language Models (EACL 2026) aclanthology.org/2026.eacl-lo... 📁 github.com/screemix/Wik...
aclanthology.org
Wikontic: Constructing Wikidata-Aligned, Ontology-Aware Knowledge Graphs with Large Language Models
Alla Chepurova, Aydar Bulatov, Mikhail Burtsev, Yuri Kuratov. Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
010