Sign in

apoorva lal

@apoorvalal.com
5.2K followers 608 following 1.8K posts

causal inference, econometrics, ML, arsenal, loud music, unix, FOSS for scientific computing. opinions my own. apoorvalal.github.io (passively) maintains @paperposterbot.bsky.social

PostsRepliesMedia
apoorva lal @apoorvalal.com · 09/05/2026
is this masheen leurning
020
apoorva lal @apoorvalal.com · 03/05/2026
cheap, fast local embeddings for lookups. Then again based on your ability to instantaneously recall things this might not be a big bump relative to your own head github.com/tobi/qmd
github.com
GitHub - tobi/qmd: mini cli search engine for your docs, knowledge bases, meeting notes, whatever. Tracking current sota approaches while being all local
mini cli search engine for your docs, knowledge bases, meeting notes, whatever. Tracking current sota approaches while being all local - tobi/qmd
140
apoorva lal @apoorvalal.com · 02/05/2026
Get a clanker to make a python port!
100
apoorva lal @apoorvalal.com · 28/04/2026
good bot
010
apoorva lal @apoorvalal.com · 21/04/2026
GEoRCE
120
apoorva lal @apoorvalal.com · 18/04/2026
but the whimsy grant! would you think of all the whimsy of a lobotomized jester dead behind the eyes
230
apoorva lal @apoorvalal.com · 14/04/2026
true; i think this was just path of least resistance. i don't think the hansen material sets it up with XX' and potential p>n cases [that's for Bach or Wainwright, you go try that; too rich for my blood]
010
apoorva lal @apoorvalal.com · 14/04/2026
⅟ (Xᵀ * X) being the inverse operator applied to XᵀX is such a cursed combination of programming and math idioms it is kinda amazing.
030
apoorva lal @apoorvalal.com · 14/04/2026
extremist static typing and proofs are a match made in heaven - check out how we get OLS
200
apoorva lal @apoorvalal.com · 14/04/2026
been trying to get my claw to teach me Lean - send PRs cuz misery loves company github.com/apoorvalal/l...
github.com
GitHub - apoorvalal/lean-hansen-econometrics: formalizing econometrics
formalizing econometrics. Contribute to apoorvalal/lean-hansen-econometrics development by creating an account on GitHub.
282
apoorva lal @apoorvalal.com · 11/04/2026
AI Adoption (2026), colorised.
050
apoorva lal @apoorvalal.com · 10/04/2026
arxiv.org/abs/2401.12143 forgot why i read this but i did
arxiv.org
Anisotropy Is Inherent to Self-Attention in Transformers
The representation degeneration problem is a phenomenon that is widely observed among self-supervised learning methods based on Transformers. In NLP, it takes the form of anisotropy, a singular property of hidden representations which makes them unexpectedly close to each other in terms of angular distance (cosine-similarity). Some recent works tend to show that anisotropy is a consequence of optimizing the cross-entropy loss on long-tailed distributions of tokens. We show in this paper that anisotropy can also be observed empirically in language models with specific objectives that should not suffer directly from the same consequences. We also show that the anisotropy problem extends to Transformers trained on other modalities. Our observations suggest that anisotropy is actually inherent to Transformers-based models.
151
apoorva lal @apoorvalal.com · 08/04/2026
Bring back self censorship
010
apoorva lal @apoorvalal.com · 08/04/2026
That's what aperture science was. Portal was an RL env
020
apoorva lal @apoorvalal.com · 06/04/2026
Nice, yeah pretty funny it settled on quite similar ui too (codex likes its pastel backdrop html). I think with appropriate metadata scouring gh for different variations on the same idea should be feasible, although licensing etc could be a minefield. In this case, parallel branches+merge would do
010
apoorva lal @apoorvalal.com · 05/04/2026
to kick the tires, poke around with my edits of the cars dataset figure here and mangle it some more lalten.org/vega-ui/char...
lalten.org
Vega UI
000
apoorva lal @apoorvalal.com · 05/04/2026
quite nice prototype built over telegram: apply finishing touches on your figures in a web-ui using the magic of vega json graphics and get back json/code that you can put back in your source-code [thereby maintaining reproducibility of output - usually my biggest bugbear with wysiwyg fig edits]
140
apoorva lal @apoorvalal.com · 05/04/2026
prototype seems to work; source here github.com/apoorvalal/v... and deployed here lalten.org/vega-ui/ 1) get starter plot from altair and extract json; then paste into vega-ui 2) edit [fig1:every 'apply' button mutates the vega json directly] 3) when done, extract python/json [fig2]
151
apoorva lal @apoorvalal.com · 05/04/2026
this is a nice idea, although i wouldn't use matplotlib but altair [since it has vega underneath and that fully contains the data + code so editing it can probably permit a clean round trip back into source code]. Set off a job on my openclaw; will report back with a link if promising.
130
apoorva lal @apoorvalal.com · 05/04/2026
github.com/apoorvalal/m... Demucs + basic pitch might help with what you want; I glued things together here for my own ear training and it works surprisingly well
github.com
GitHub - apoorvalal/mlodies: glues pretrained models to do stem separation and transcription to help learn music by ear.
glues pretrained models to do stem separation and transcription to help learn music by ear. - apoorvalal/mlodies
020
apoorva lal @apoorvalal.com · 04/04/2026
modal.com/blog/resourc... linear programming done well
modal.com
Linear Programming for Fun and Profit
How we use an eighty-year-old algorithm to find arbitrages in the cloud market.
161
apoorva lal @apoorvalal.com · 29/03/2026
ha Zellij way to spook the normies. Ghostty tabs are fine.
150
apoorva lal @apoorvalal.com · 27/03/2026
Making sense of Power-tweets is a pretty good case for reasoning models. chatgpt.com/share/69c609... is HCR an OG minimax result? I‘d seen the Loewner order version of CR and the associated ellipsoid claim is natural. HCR as ’a secant body in chi-squared space’ makes my head hurt.
210
apoorva lal @apoorvalal.com · 26/03/2026
the aristotelian method is when you are such a boring teacher and wont shut up about rocks having pneuma that your ideal student runs away and conquers most of the known world
020
apoorva lal @apoorvalal.com · 22/03/2026
ha yeah the least squares solver is blazing fast [which came as a surprise to me - rust libraries rely way less on BLAS/LAPACK and perf opts from mkl].
000
apoorva lal @apoorvalal.com · 22/03/2026
Painstakingly detailed documentation; this is a good intro to binding rust apoorvalal.github.io/crabbymetric...
apoorvalal.github.io
Mechanics: OLS – crabbymetrics
100
apoorva lal @apoorvalal.com · 22/03/2026
Codex is great at Rust, i am not. I have strong preferences and tests for what a lean statistics package should do, it does not. I'm not crazy enough to try to patch out numpy deps but that really is it. Solid collaboration, now on pypi (uv add crabbymetrics) github.com/apoorvalal/c...
260
apoorva lal @apoorvalal.com · 22/03/2026
allpoetry.com/colonel-faza...
allpoetry.com
Colonel Fazackerley Butterworth-Toast by Charles Causley - Famous poems, famous poets. - All Poetry
Comments & analysis: Colonel Fazackerley Butterworth-Toast / Bought an old castle complete with a ghost,
000
apoorva lal @apoorvalal.com · 16/03/2026
now you have to write the paper. Mondrian Orderings for Interactive Sequential Tests using Sequential Hierarchical Additive Regression Trees
010
apoorva lal @apoorvalal.com · 15/03/2026
Structured Hierarchical Additive Regression Trees
240
apoorva lal @apoorvalal.com · 15/03/2026
based on recent history i think it would be called something wonderfully juvenile like "deezNUTS" or "YUTS"
110
apoorva lal @apoorvalal.com · 15/03/2026
agree it's tricky; i think for SEs the standard GMM style sandwich representation should be valid [Conley is written with generic moment conditions in mind IIRC]. Optimal bandwidth for RD is harder because the MSE representation in sth like CCT is specific to local linear.
010
apoorva lal @apoorvalal.com · 15/03/2026
no super clean solution i'm afraid; pandoc-ing the tex->markdown conversion and then getting an LLM to iterate on the conversion has reasonable odds of success but tex is too messy to reliably automate bsky.app/profile/apoo...
111
apoorva lal @apoorvalal.com · 14/03/2026
Apropos of econometrics discourse du jour, people are often trying to model a multiplicative shift in the rate of some event, and OLS really is just bad for this. Poisson is good. lalten.org/pages/counti... (Aside, quarto papers are better than pdfs, esp for our clanker assistants)
lalten.org
Difference in Differences for Event Data: The Temporal MAUP and Solutions
3232
apoorva lal @apoorvalal.com · 12/03/2026
Flip the (now somewhat hack) metaphor for human-AI augmentation. AI used well with human in the lead: Centaur AI used poorly with prose and code smells: horseface Same ingredients and proportions, very different outcomes.
070
apoorva lal @apoorvalal.com · 10/03/2026
apoorvalal.github.io/pyensmallen/ Pyensmallen now has a nice website with benchmarks (it is very fast!) and API docs. An LLM agent also helped me figure out how to patch a weird BLAS error that cropped up on older mac metal; should now work out of the box.
apoorvalal.github.io
pyensmallen
020
apoorva lal @apoorvalal.com · 10/03/2026
open a PR!
020
apoorva lal @apoorvalal.com · 10/03/2026
This looks cool. But also reminded me of a jab I've been meaning to make at methods land: All of interference: Cai et al 2015 Panel data: California Prop99 IV: Card 1993 Unconfoundedness: Lalonde, 401k
220
apoorva lal @apoorvalal.com · 08/03/2026
Stanford. This covers material from Hainmueller, Grimmer, Imbens, Wager, Rivers, Athey, Owen and a lot of self-study
210
apoorva lal @apoorvalal.com · 07/03/2026
alternate timeline is better youtu.be/_90HLJ_DYS8
youtu.be
That Mitchell and Webb Look - Medical Drama
YouTube video by 48bytes
020
apoorva lal @apoorvalal.com · 07/03/2026
conversion workflow for long context-management and automating hill-climbing on latex<>katex rough edges that might help with your own big tex docs apoorvalal.github.io/lalgorithms/...
apoorvalal.github.io
Conversion Workflow
This documents the workflow used to convert the LaTeX methods notes into Quartz chapter notes, plus a debrief of what was required in practice.
010
apoorva lal @apoorvalal.com · 07/03/2026
LLMs are finally good enough at latex that codex helped me port over my gigantic (150+pg 2col) grad school methods notes into an obsidian vault that makes upkeep considerably easier. Read, refer relevant sections to your favourite clanker, enjoy! apoorvalal.github.io/lalgorithms/...
apoorvalal.github.io
Methods Rolodex
A structured set of methods notes migrated from the LaTeX source in 00_methods_notes.
3412
apoorva lal @apoorvalal.com · 05/03/2026
Yeah I'm an early adopter but it fits a small subset of what I want from a scientific notebook interface (esp that it should export to static html and not need to be spun up by readers). Great for interactive notebooks and dashboards though.
010
apoorva lal @apoorvalal.com · 05/03/2026
side-by-side code snippets for a variety of standard figures in 4 python libraries I like a lot [altair is my most used lately for aforementioned tooltip reasons] lalten.org/pages/py_viz...
lalten.org
Data Visualization Comparison
030
apoorva lal @apoorvalal.com · 05/03/2026
viz with tooltips are a great way to quickly verify that your agents aren't dropping data or doing weird shenanigans, and are considerably less clunky with embedding loads of info than static plots. eg: matplotlib on left, altair with a tooltip on the right.
140
apoorva lal @apoorvalal.com · 05/03/2026
things i use a lot more often in our agentic age - uv init + uv add + uv sync everything - quarto markdown instead of jupyter - hidden state is especially cursed now - interactive viz libraries that generate tooltips - tmux splits with different agent instances playing generator-discriminator
2241
apoorva lal @apoorvalal.com · 24/02/2026
Lol same dank room sold me on it bsky.app/profile/apoo...
011
apoorva lal @apoorvalal.com · 24/02/2026
Subtract off socarxiv and biorxiv as control series given lower tech adoption and you have a diff in diff
030
apoorva lal @apoorvalal.com · 22/02/2026
Written primarily over telegram with Krusty the Krabs, who trained SideShowSpongeBob during the course of doing this project
000
apoorva lal @apoorvalal.com · 22/02/2026
Squeezing more performance out of small models by algorithmically "asking nicely". apoorvalal.github.io/lalgorithms/...
apoorvalal.github.io
Optimizing tiny LLM programs: EuroSAT + DSPy MIPROv2
gist This note documents a small end‑to‑end experiment: run a local Qwen3‑VL multimodal model via llama.cpp’s llama-server, use it as a VLM classifier on EuroSAT (satellite land‑use classes), and then...
130