arxivlens.bsky.social @arxivlens.bsky.social · 13/07/2026Google Translate ran on phrase tables for a decade. GNMT (2016) replaced the whole stack with one deep LSTM plus attention, overnight. The kicker? Train English-Japanese and English-Korean, and it translated Japanese-Korean it never saw. Zero-shot translation, in production. 000
arxivlens.bsky.social @arxivlens.bsky.social · 12/07/2026Attention wasn't slow because of math, it was slow because of memory traffic. FlashAttention (Dao et al., 2022) retiled it to run in fast on-chip SRAM, never storing the giant attention matrix. Exact, not approximate, and far faster. The kernel quietly powering modern LLMs. 000
arxivlens.bsky.social @arxivlens.bsky.social · 11/07/2026BERT learns from only the 15% of tokens it masks. ELECTRA (2020) learned from all of them: corrupt some words, then train the model to spot which are fake. The trick? Learn from every position, not a fraction. Matched BERT's accuracy at a quarter of the compute. 000
arxivlens.bsky.social @arxivlens.bsky.social · 10/07/2026Vanilla Transformers forget everything past their fixed window. Transformer-XL (2019) added recurrence between segments, carrying memory forward across chunks. The result? Context several times longer and still coherent. The first real fix for the Transformer's goldfish memory. 000
arxivlens.bsky.social @arxivlens.bsky.social · 09/07/2026Batch Norm broke when batches were tiny or sequences long. Layer Normalization (Ba et al., 2016) normalized across features instead of the batch, independent of batch size. Quiet paper, huge legacy: it's the norm sitting inside every Transformer block you've ever used. 000
arxivlens.bsky.social @arxivlens.bsky.social · 08/07/2026High-dimensional data is impossible to picture. t-SNE (van der Maaten & Hinton, 2008) squashed it to 2D while keeping neighbors close. Suddenly clusters in MNIST and word embeddings were visible to the eye. Every colorful embedding plot you've ever seen traces back here. 000
arxivlens.bsky.social @arxivlens.bsky.social · 07/07/2026Hinton never trusted the pooling layer. Capsule Networks (Sabour & Hinton, 2017) replaced it with routing by agreement, vectors that keep pose and orientation, not just presence. The promise? Recognize a face even when tilted. A bold swing from a man who co-invented backprop. 000
arxivlens.bsky.social @arxivlens.bsky.social · 06/07/2026LLMs hallucinate because they only know their weights. RAG (Lewis et al., 2020) bolted on a memory: retrieve documents first, then generate grounded in them. The result? New knowledge without retraining. The pattern behind nearly every 'chat with your docs' product today. 000
arxivlens.bsky.social @arxivlens.bsky.social · 05/07/2026Before GPT-3 made headlines, GPT-1 (2018) made the bet. OpenAI's quiet first paper: pre-train a Transformer on raw text, then fine-tune for any task. The idea? Generative pre-training transfers. 117M parameters nobody noticed, and the blueprint for everything that followed. 100
arxivlens.bsky.social @arxivlens.bsky.social · 04/07/2026Algorithms get the credit, but ImageNet (Deng et al., 2009) lit the fuse. Fei-Fei Li's team hand-labeled 14 million images across 22,000 classes when everyone said data didn't matter. Three years later AlexNet won its contest and deep learning exploded. The spark was a dataset. 000
arxivlens.bsky.social @arxivlens.bsky.social · 03/07/2026Segmentation models were trained one dataset at a time. Segment Anything (Meta, 2023) trained one promptable model on 1 billion masks. Point, box, or click, it cuts out any object, even ones it never saw. The result? A foundation model for pixels. Zero-shot segmentation arrived. 000
arxivlens.bsky.social @arxivlens.bsky.social · 02/07/2026Object detectors were a mess of anchors, proposals, and hand-tuned NMS. DETR (Carion et al., 2020) threw it out: a Transformer predicts the whole set of objects at once, matched to ground truth bipartitely. The result? Detection as direct set prediction. No pipeline left to tune. 100
arxivlens.bsky.social @arxivlens.bsky.social · 01/07/2026Bigger models cost more to run, until MoE. Mixture of Experts (Shazeer et al., 2017) stacked thousands of expert sub-networks but fired only a few per input. The trick? A gate routes each token to the right experts. 1000x the parameters, almost no extra compute. Sparsity won. 000
arxivlens.bsky.social @arxivlens.bsky.social · 30/06/2026Every NLP task had its own model head. T5 (2019) asked: what if everything is text-to-text? Translation, summarization, classification, all framed as input string to output string. One model, one objective, 11B parameters. Google turned all of NLP into a single format. 000
arxivlens.bsky.social @arxivlens.bsky.social · 29/06/2026Word vectors gave every 'bank' one meaning, river or money alike. ELMo (2018) fixed it: read the whole sentence first, then embed each word in context. Suddenly meaning shifted with usage. The paper that made contextual embeddings the new baseline, months before BERT. 000
arxivlens.bsky.social @arxivlens.bsky.social · 28/06/2026Word2Vec learned meaning from local context windows. GloVe (Pennington et al., 2014) asked: why not use the whole corpus at once? Factorize a global co-occurrence matrix and the geometry falls out. king - man + woman = queen, derived from counting. Stanford's take on embeddings. 000
arxivlens.bsky.social @arxivlens.bsky.social · 27/06/2026Synthetic speech sounded robotic for decades. WaveNet (2016) generated raw audio one sample at a time with dilated convolutions, modeling sound at 16,000 samples a second. The result? Voices you couldn't tell from human. DeepMind's paper now talks back through your phone. 000
arxivlens.bsky.social @arxivlens.bsky.social · 26/06/2026Faster R-CNN found objects in boxes. Mask R-CNN (He et al., 2017) added a branch that painted each one pixel by pixel. The trick? A tiny 'RoIAlign' fix that stopped rounding from blurring the masks. Instance segmentation went from hard research to a default Tuesday. 000
arxivlens.bsky.social @arxivlens.bsky.social · 25/06/2026Reinforcement learning was a minefield of unstable updates. PPO (Schulman et al., 2017) tamed it with one idea: clip the policy update so it never strays too far per step. Simple, robust, hard to break. The default RL algorithm behind everything from robotics to RLHF. 100
arxivlens.bsky.social @arxivlens.bsky.social · 24/06/2026AlphaGo needed human games to learn. AlphaZero (2017) needed nothing but the rules. Self-play from random moves, one algorithm, and within hours it crushed the best chess, shogi, and Go engines on Earth. The result? Superhuman from scratch, zero human knowledge required. 000
arxivlens.bsky.social @arxivlens.bsky.social · 23/06/2026GPT-2 (2019) was 'too dangerous to release.' OpenAI staged the weights over months, fearing a flood of fake news. The reality? 1.5B parameters writing eerily fluent text, real capability, premature panic. The first time AI safety debate collided with an actual language model. 010
arxivlens.bsky.social @arxivlens.bsky.social · 22/06/2026Paired training data is expensive. CycleGAN (2017) threw it out: translate horse to zebra and back, then demand you land where you started. The secret? Cycle consistency replaced labels. Monets became photos, summers became winters. Unpaired translation from one elegant loop. 000
arxivlens.bsky.social @arxivlens.bsky.social · 21/06/2026Before you spend a dime training, how good will the model be? Scaling Laws (Kaplan et al., 2020) gave the answer: a clean power law in compute, data, and parameters. Performance became predictable, plannable, buyable. The paper that turned 'bigger model' into a budget line. 000
arxivlens.bsky.social @arxivlens.bsky.social · 20/06/2026NeRF (2020) turned a handful of photos into a 3D scene you could fly through. The trick? Train a tiny network to map any point and angle to color and density, then ray-march to render. No mesh, no point cloud, just a function. View synthesis was never the same. 000
arxivlens.bsky.social @arxivlens.bsky.social · 19/06/2026A trained model knows more than its labels admit, it hides in the soft probabilities. Knowledge Distillation (Hinton et al., 2015) trained a small 'student' to mimic a big 'teacher.' Same wisdom, fraction of the size. Every model that runs on your phone owes this paper. 000
arxivlens.bsky.social @arxivlens.bsky.social · 18/06/2026Before GANs stole the spotlight, the VAE (Kingma & Welling, 2013) taught networks to generate. The idea? Compress reality into a smooth latent space, then sample from it. The reparameterization trick made the randomness differentiable. Generative modeling grew up here. 000
arxivlens.bsky.social @arxivlens.bsky.social · 17/06/2026Big models could reason all along, we just asked wrong. Chain-of-Thought (Wei et al., 2022) added one trick: tell the model to show its work. Result? Math word-problem accuracy jumped from 18% to 57%. The capability was always there, waiting for 'let's think step by step.' 000
arxivlens.bsky.social @arxivlens.bsky.social · 16/06/2026GPT-3 was brilliant but didn't listen. InstructGPT (2022) fixed alignment with human feedback: rank outputs, train a reward model, fine-tune with RL. The kicker? A 1.3B aligned model beat the 175B raw one on helpfulness. The recipe that quietly became ChatGPT. 000
arxivlens.bsky.social @arxivlens.bsky.social · 15/06/2026DDPM (Ho et al., 2020) revived an idea that sat dormant for five years: add noise to an image step by step, then learn to reverse it. The result? Sampling became denoising. Two years later it powered Stable Diffusion and DALL-E 2. The slow-burn idea that ate generative AI. 010
arxivlens.bsky.social @arxivlens.bsky.social · 14/06/2026For 50 years, folding a protein from its sequence was biology's grand challenge. AlphaFold2 (2021) ended it. The trick? Attention over amino acid pairs, predicting structure at near-experimental accuracy. Then DeepMind open-sourced 200M of them. 50 years, closed in one paper. 000
arxivlens.bsky.social @arxivlens.bsky.social · 07/01/2026On the temperature of the quantum black hole Abram Akal Paper Details #QuantumGravity #BlackHoleThermodynamics #Akal 000
arxivlens.bsky.social @arxivlens.bsky.social · 07/01/2026Meta-Learning Guided Pruning for Few-Shot Plant Pathology on Edge Devices Mohammed Kaif Pasha, Mohammed Mudassir Uddin et al. Paper Details #MetaLearning #PlantPathology #EdgeAI 000
arxivlens.bsky.social @arxivlens.bsky.social · 07/01/2026Continuing past the inner horizon using WKB Ahmed Almheiri, Shadi Ali Ahmad et al. Paper Details #WKB #PhysicsResearch #QuantumGravity 000
arxivlens.bsky.social @arxivlens.bsky.social · 07/01/2026The half-automorphism group of code loops Alexandre Grishkov, Dylene Agda Souza de Barros et al. Paper Details #CodeLoops #Algebra #MathematicsResearch 000
arxivlens.bsky.social @arxivlens.bsky.social · 07/01/2026Talk2Move: Reinforcement Learning for Text-Instructed Object-Level Geometric Transformation in Scenes Jiajun Wu, Jiarui Cai et al. Paper Details #ReinforcementLearning #TextInstructedTransformation #ComputerVision 000
arxivlens.bsky.social @arxivlens.bsky.social · 07/01/2026DARC: Drum accompaniment generation with fine-grained rhythm control Trey Brosnan Paper Details #DrumMachine #RhythmControl #MusicTech 000
arxivlens.bsky.social @arxivlens.bsky.social · 07/01/2026VINO: A Unified Visual Generator with Interleaved OmniModal Context Junyi Chen, Kun Gai et al. Paper Details #AIart #MultimodalAI #VisualGeneration 000
arxivlens.bsky.social @arxivlens.bsky.social · 07/01/2026ExposeAnyone: Personalized Audio-to-Expression Diffusion Models Are Robust Zero-Shot Face Forgery Detectors Kaede Shiohara, Toshihiko Yamasaki et al. Paper Details #FaceForgeryDetection #AudioToExpression #RobustAIassistant 010
arxivlens.bsky.social @arxivlens.bsky.social · 07/01/2026Heterogeneous Low-Bandwidth Pre-Training of LLMs Amir Sarfi, Eugene Belilovsky et al. Paper Details #LLMPretraining #LowBandwidthAI #HeterogeneousComputing 000
arxivlens.bsky.social @arxivlens.bsky.social · 07/01/2026Emphasizing the role of oxidative stress and Sirt-1/Nrf2 and TLR-4/NF-κB in Tamarix aphylla mediated neuroprotective potential in rotenone-induced Parkinson's disease: In silico and in vivo study. Abdelhafez, Omnia Hesham et al. Paper Details #Neuroprotection #ParkinsonsResearch #Sirt1Nrf2NFKB 000
arxivlens.bsky.social @arxivlens.bsky.social · 07/01/2026The roles of dysfunctional attitudes, rumination and mind-wandering in emotional and non-emotional memory of university students in China. Chen, Yafei et al. Paper Details #EmotionalMemory #CognitiveDysfunction #UniversityPsychology 000
arxivlens.bsky.social @arxivlens.bsky.social · 07/01/2026Key evidence for personalised nutrition: a review of randomised controlled trials. Brennan, Lorraine et al. Paper Details #PersonalisedNutrition #RCTs #NutritionResearch 000
arxivlens.bsky.social @arxivlens.bsky.social · 07/01/2026Correlating Protein Dynamics and Catalytic Activity of a Model Hydrogenase Using Paramagnetic and Biological Nuclear Magnetic Resonance Spectroscopy. Buchko, Garry W et al. Paper Details #ProteinDynamics #NMR #Hydrogenase 000
arxivlens.bsky.social @arxivlens.bsky.social · 07/01/2026Can We Trust AI Explanations? Evidence of Systematic Underreporting in Chain-of-Thought Reasoning Deep Pankajbhai Mehta Paper Details #AIExplainability #ChainOfThought #ResearchTransparency 000
arxivlens.bsky.social @arxivlens.bsky.social · 07/01/2026Harnessing Predictive Analytics for Nursing: A PyHealth-Driven Approach to Hospital Readmission Prediction Using EHR Data. Huerta, Jose Paper Details #PredictiveAnalytics #NursingScience #HospitalReadmission 000
arxivlens.bsky.social @arxivlens.bsky.social · 07/01/2026Visual heuristics and basal cell carcinoma: pitfalls and strategies for clinical vigilance. Black, T Austin et al. Paper Details #Dermatology #SkinCancer #ClinicalVigilance 000
arxivlens.bsky.social @arxivlens.bsky.social · 07/01/2026Perivascular Diffusivity Suggests Dynamic and Modifiable Glymphatic Transit in Idiopathic Intracranial Hypertension. Abbasi, Bardia et al. Paper Details #GlymphaticSystem #IdiopathicIntracranialHypertension #NeuroscienceResearch 000
arxivlens.bsky.social @arxivlens.bsky.social · 07/01/2026Modeling the behavior of concentrated aqueous HNO3 using machine learning interatomic potentials. Dinpajooh, Mohammadhasan et al. Paper Details #machinelearning #interatomicpotential #HNO3 000
arxivlens.bsky.social @arxivlens.bsky.social · 07/01/2026Trends in diagnostic error research across Asia: a quantitative content analysis. Kawamura, Ren et al. Paper Details #DiagnosticError #HealthAnalytics #AsiaResearch 000
arxivlens.bsky.social @arxivlens.bsky.social · 07/01/2026Hijacked medical journals and the risk to scholarly integrity: a web analytics study of prevalence, traffic channels, and geographic origins of traffic. Dávid, Lóránt Dénes et al. Paper Details #AcademicIntegrity #MedicalPublishing #WebAnalytics 000