Sign in

Kyle Kastner

@kastnerkyle.bsky.social
395 followers 789 following 77 posts

computers and music are (still) fun

PostsRepliesMedia
Reposted by Kyle Kastner
Motonobu Kanagawa @motonobu-kanagawa.bsky.social · 04/09/2025
ProbNum 2025 Keynote 2 ``Gradient Flows on the Maximum Mean Discrepancy'' by @arthurgretton.bsky.social ( @gatsbyucl.bsky.social and Google DeepMind. Slides available here: probnum25.github.io/keynotes
182
Reposted by Kyle Kastner
Tim Duffy @timfduffy.com · 22/07/2025
Surprising new results from Owain Evans and Anthropic: Training on the outputs of a model can change the model's behavior, even when those outputs seem unrelated. Training only on completions of 3-digit numbers was able to transmit a love of owls. alignment.anthropic.com/2025/sublimi...
5335
Reposted by Kyle Kastner
Catherine Arnett @catherinearnett.bsky.social · 10/07/2025
MorphScore got an update! MorphScore now covers 70 languages 🌎🌍🌏 We have a new-preprint out and we will be presenting our paper at the Tokenization Workshop @tokshop.bsky.social at ICML next week! @marisahudspeth.bsky.social @brenocon.bsky.social
1134
Reposted by Kyle Kastner
Harry Thasarathan @hthasarathan.bsky.social · 01/05/2025
Our work finding universal concepts in vision models is accepted at #ICML2025!!! My first major conference paper with my wonderful collaborators and friends @matthewkowal.bsky.social @thomasfel.bsky.social @Julian_Forsyth @csprofkgd.bsky.social Working with y'all is the best 🥹 Preprint ⬇️!!
0164
Reposted by Kyle Kastner
Jack Greenhalgh @wildaudiojack.bsky.social · 09/06/2025
Contribute to the first global archive of soniferous freshwater life, The Freshwater Sounds Archive, and receive recognition as a co-author in a resulting data paper! Pre-print now available. New deadline: 31st Dec, 2025. See link 👇4 more fishsounds.net/freshwater.js
44116
Reposted by Kyle Kastner
Daniel Tanneberg @dantanvii.bsky.social · 19/05/2025
🚀 Interested in Neuro-Symbolic Learning and attending #ICRA2025? 🧠🤖 Do not miss Leon Keller presenting “Neuro-Symbolic Imitation Learning: Discovering Symbolic Abstractions for Skill Learning”. Joint work of Honda Research Institute EU and @jan-peters.bsky.social (@ias-tudarmstadt.bsky.social).
1102
Reposted by Kyle Kastner
arxiv cs.CL @arxiv-cs-cl.bsky.social · 23/05/2025
Prasoon Bajpai, Tanmoy Chakraborty Multilingual Test-Time Scaling via Initial Thought Transfer arxiv.org/abs/2505.15508
011
Reposted by Kyle Kastner
AI Firehose @ai-firehose.column.social · 23/05/2025
A study shows in-context learning in spoken language models can mimic human adaptability, reducing word error rates by nearly 20% with just a few utterances, especially aiding low-resource language varieties and enhancing recognition across diverse speakers. arxiv.org/abs/2505.14887
arxiv.org
In-Context Learning Boosts Speech Recognition via Human-like Adaptation to Speakers and Language Varieties
ArXiv link for In-Context Learning Boosts Speech Recognition via Human-like Adaptation to Speakers and Language Varieties
011
Reposted by Kyle Kastner
‏ deepfates @deepfates.com.deepfates.com.deepfates.com.deepfates.com.deepfates.com · 22/05/2025
"Interdimensional Cable", shorts made with Veo 3 ai. By CodeSamurai on Reddit
1117228
Reposted by Kyle Kastner
arxiv cs.CV @arxiv-cs-cv.bsky.social · 16/05/2025
Bingda Tang, Boyang Zheng, Xichen Pan, Sayak Paul, Saining Xie Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis arxiv.org/abs/2505.10046
011
Reposted by Kyle Kastner
arXiv Sound @arxiv-sound.bsky.social · 16/05/2025
A neural ODE model combined modal decomposition with a neural network to model nonlinear string vibrations, generating synthetic data and sound examples.
arxiv.org
Learning Nonlinear Dynamics in Physical Modelling Synthesis using Neural Ordinary Differential Equations
Victor Zheleznov, Stefan Bilbao, Alec Wright, Simon King
021
Reposted by Kyle Kastner
AI Firehose @ai-firehose.column.social · 15/05/2025
Research unveils Omni-R1, a fine-tuning method for audio LLMs that boosts audio performance via text training, achieving MMAU results. Findings reveal how enhanced text reasoning affects audio capacities, suggesting new model optimization directions. arxiv.org/abs/2505.09439
arxiv.org
Omni-R1: Do You Really Need Audio to Fine-Tune Your Audio LLM?
ArXiv link for Omni-R1: Do You Really Need Audio to Fine-Tune Your Audio LLM?
011
Reposted by Kyle Kastner
Alexander Doria @dorialexander.bsky.social · 13/05/2025
Yeah we finally have a model report with an actual data section. Thanks Qwen 3! github.com/QwenLM/Qwen3...
15310
Reposted by Kyle Kastner
AI Firehose @ai-firehose.column.social · 10/05/2025
FLAM, a novel audio-language model, enables frame-wise localization of sound events in an open-vocabulary format. With large-scale synthetic data and advanced training methods, FLAM enhances audio understanding and retrieval, aiding multimedia indexing and access. arxiv.org/abs/2505.05335
arxiv.org
FLAM: Frame-Wise Language-Audio Modeling
ArXiv link for FLAM: Frame-Wise Language-Audio Modeling
021
Reposted by Kyle Kastner
Ahmad Beirami @abeirami.bsky.social · 09/05/2025
#ICML2025 Is standard RLHF optimal in view of test-time scaling? Unsurprisingly no. We show a simple change to standard RLHF framework that involves 𝐫𝐞𝐰𝐚𝐫𝐝 𝐜𝐚𝐥𝐢𝐛𝐫𝐚𝐭𝐢𝐨𝐧 and 𝐫𝐞𝐰𝐚𝐫𝐝 𝐭𝐫𝐚𝐧𝐬𝐟𝐨𝐫𝐦𝐚𝐭𝐢𝐨𝐧 (suited to test-time procedure) is optimal!
1176
Reposted by Kyle Kastner
Dylan Foster 🐢 @djfoster.bsky.social · 03/05/2025
Is Best-of-N really the best we can do for language model inference? New paper (appearing at ICML) led by the amazing Audrey Huang (ahahaudrey.bsky.social) with Adam Block, Qinghua Liu, Nan Jiang, and Akshay Krishnamurthy (akshaykr.bsky.social). 1/11
1215
Reposted by Kyle Kastner
Tim G. J. Rudner @timrudner.bsky.social · 29/04/2025
Congratulations to the #AABI2025 Workshop Track Outstanding Paper Award recipients!
0238
Reposted by Kyle Kastner
Sung Kim @sungkim.bsky.social · 30/04/2025
Why not? Reinforcement Learning for Reasoning in Large Language Models with One Training Example Applying RLVR to the base model Qwen2.5-Math-1.5B, they identify a single example that elevates model performance on MATH500 from 36.0% to 73.6%,
2202
Reposted by Kyle Kastner
AI Firehose @ai-firehose.column.social · 29/04/2025
Instruct-LF merges LLMs' instruction-following with statistical models, enhancing interpretability in noisy datasets and improving task performance up to 52%. arxiv.org/abs/2502.15147
arxiv.org
Latent Factor Models Meets Instructions: Goal-conditioned Latent Factor Discovery without Task Supervision
ArXiv link for Latent Factor Models Meets Instructions: Goal-conditioned Latent Factor Discovery without Task Supervision
021
Reposted by Kyle Kastner
Sung Kim @sungkim.bsky.social · 27/04/2025
An incomplete list of Chinese AI: - DeepSeek: www.deepseek.com. You can also access AI models via API. - Moonshot AI's Kimi: www.kimi.ai - Alibaba's Qwen: chat.qwen.ai. You can also access AI models via API. - ByteDance's Doubaob (only in Chinese): www.doubao.com/chat/
1227
Reposted by Kyle Kastner
lebellig @lebellig.bsky.social · 26/04/2025
I really liked this approach by @matthieuterris.bsky.social et al.They propose learning a unique lightweight model for multiple inverse problems by conditioning it with the forward operator A. Thanks to self-supervised fine-tuning, it can tackle unseen inverse pb. 📰 arxiv.org/abs/2503.08915
071
Reposted by Kyle Kastner
Mattie Fellows @mattieml.bsky.social · 25/04/2025
Excited to be presenting our spotlight ICLR paper Simplifying Deep Temporal Difference Learning today! Join us in Hall 3 + Hall 2B Poster #123 from 3pm :)
arxiv.org
071
Reposted by Kyle Kastner
speechpapers.bsky.social @speechpapers.bsky.social · 26/04/2025
Balinese text-to-speech dataset as digital cultural heritage pubmed.ncbi.nlm.nih.gov/40275973
011
Reposted by Kyle Kastner
Sung Kim @sungkim.bsky.social · 25/04/2025
Kimi.ai releases Kimi-Audio! Our new open-source audio foundation model advances capabilities in audio understanding, generation, and conversation. Paper: github.com/MoonshotAI/K... Repo: github.com/MoonshotAI/K... Model: huggingface.co/moonshotai/K...
1132
Reposted by Kyle Kastner
lebellig @lebellig.bsky.social · 25/04/2025
Very cool article from Panagiotis Theodoropoulos et al: arxiv.org/abs/2410.14055 Feedback Schrödinger Bridge Matching introduces a new method to improve transfer between two data distributions using only a small number of paired samples!
042
Reposted by Kyle Kastner
Arno Solin @arnosolin.bsky.social · 21/04/2025
Our #ICLR2025 poster "Discrete Codebook World Models for Continuous Control" (Aidan Scannell, Mohammadreza Nakhaeinezhadfard, Kalle Kujanpää, Yi Zhao, Kevin Luck, Arno Solin, Joni Pajarinen) 🗓️ Hall 3 + Hall 2B #415, Thu 24 Apr 10 a.m. +08 — 12:30 p.m. +08 📄 Preprint: arxiv.org/abs/2503.00653
2113
Reposted by Kyle Kastner
arxiv cs.CV @arxiv-cs-cv.bsky.social · 21/04/2025
Andrew Kiruluta Wavelet-based Variational Autoencoders for High-Resolution Image Generation arxiv.org/abs/2504.13214
011
Reposted by Kyle Kastner
Sakana AI @sakanaai.bsky.social · 21/04/2025
7/ Large Language Models to Diffusion Finetuning Paper: openreview.net/forum?id=Wu5... Workshop: workshop-llm-reasoning-planning.github.io New finetuning method empowering pre-trained LLMs with some of the key properties of diffusion models and the ability to scale test-time compute.
143
Reposted by Kyle Kastner
Sakana AI @sakanaai.bsky.social · 21/04/2025
10/ Sakana AI Co-Founder and CEO, David Ha, will be giving a talk at the #ICLR2025 World Models Workshop, at a panel to discuss the Current Development and Future Challenges of World Models. Workshop Website: sites.google.com/view/worldmo...
0113
Reposted by Kyle Kastner
arXiv cs.LG Machine Learning @cslg-bot.bsky.social · 21/04/2025
Duy A. Nguyen, Quan Huu Do, Khoa D. Doan, Minh N. Do: Are you SURE? Enhancing Multimodal Pretraining with Missing Modalities through Uncertainty Estimation arxiv.org/abs/2504.13465 arxiv.org/pdf/2504.13465 arxiv.org/html/2504.13465
111
Reposted by Kyle Kastner
arXiv cs.LG Machine Learning @cslg-bot.bsky.social · 21/04/2025
Yixuan Even Xu, Yash Savani, Fei Fang, Zico Kolter: Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning arxiv.org/abs/2504.13818 arxiv.org/pdf/2504.13818 arxiv.org/html/2504.13818
123
Reposted by Kyle Kastner
Richard McElreath 🐈‍⬛ @rmcelreath.bsky.social · 18/04/2025
I have learned that some people have still not heard the Good News about Hamiltonian Monte Carlo. My gentle animated interactive explanation: elevanth.org/blog/2017/11...
elevanth.org
Markov Chains: Why Walk When You Can Flow?
In 1989, Depeche Mode was popular, the first version of Microsoft Office was released, large demonstrations brought down the wall separating East and West Germany, and a group of statisticians in the ...
28627
Reposted by Kyle Kastner
arXiv cs.LG Machine Learning @cslg-bot.bsky.social · 18/04/2025
Akira Tamamori: Kernel Ridge Regression for Efficient Learning of High-Capacity Hopfield Networks arxiv.org/abs/2504.12561 arxiv.org/pdf/2504.12561 arxiv.org/html/2504.12561
113
Reposted by Kyle Kastner
Lynn Cherny @arnicas.bsky.social · 18/04/2025
Jack Morris’s embzip github.com/jxmorris12/e... “efficiently compressing and decompressing embeddings using Product Quantization”
github.com
GitHub - jxmorris12/embzip
Contribute to jxmorris12/embzip development by creating an account on GitHub.
063
Reposted by Kyle Kastner
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 16/04/2025
Pre training for reasoning by doing RL on synthetic tasks, a position paper I’ve really enjoyed and am upset by arxiv.org/abs/2502.19402
arxiv.org
General Reasoning Requires Learning to Reason from the Get-go
Large Language Models (LLMs) have demonstrated impressive real-world utility, exemplifying artificial useful intelligence (AUI). However, their ability to reason adaptively and robustly -- the hallmar...
1436
Reposted by Kyle Kastner
AI Firehose @ai-firehose.column.social · 16/04/2025
DiTSE revolutionizes speech enhancement using latent diffusion transformers, delivering studio-quality audio while preserving speaker identity and minimizing content hallucination, transforming audio content creation and telecommunications. arxiv.org/abs/2504.09381
arxiv.org
DiTSE: High-Fidelity Generative Speech Enhancement via Latent Diffusion Transformers
ArXiv link for DiTSE: High-Fidelity Generative Speech Enhancement via Latent Diffusion Transformers
011
Reposted by Kyle Kastner
Serge Belongie @serge.belongie.com · 16/04/2025
A must-read for CV/NLP/ML grad students seeking wisdom on writing conference papers
2333
Reposted by Kyle Kastner
AI Firehose @ai-firehose.column.social · 15/04/2025
A key study presents Dynamic Importance Sampling for Constrained Decoding (DISC), enhancing efficiency and accuracy in large language models. This method reduces bias and optimizes constrained generation, broadening AI's practical applications in real-world uses. arxiv.org/abs/2504.09135
arxiv.org
Efficient and Asymptotically Unbiased Constrained Decoding for Large Language Models
ArXiv link for Efficient and Asymptotically Unbiased Constrained Decoding for Large Language Models
011
Reposted by Kyle Kastner
Hacker News 100 @hn100.atproto.rocks · 14/04/2025
NoProp: Training neural networks without back-propagation or forward-propagation arxiv.org/abs/2503.24322 news.ycombinator.com/item?id=436768…
arxiv.org
NoProp: Training Neural Networks without Back-propagation or Forward-propagation
The canonical deep learning approach for learning requires computing a gradient term at each layer by back-propagating the error signal from the output towards each learnable parameter. Given the stacked structure of neural networks, where each layer builds on the representation of the layer below, this approach leads to hierarchical representations. More abstract features live on the top layers of the model, while features on lower layers are expected to be less abstract. In contrast to this, we introduce a new learning method named NoProp, which does not rely on either forward or backwards propagation. Instead, NoProp takes inspiration from diffusion and flow matching methods, where each layer independently learns to denoise a noisy target. We believe this work takes a first step towards introducing a new family of gradient-free learning methods, that does not learn hierarchical representations -- at least not in the usual sense. NoProp needs to fix the representation at each layer beforehand to a noised version of the target, learning a local denoising process that can then be exploited at inference. We demonstrate the effectiveness of our method on MNIST, CIFAR-10, and CIFAR-100 image classification benchmarks. Our results show that NoProp is a viable learning algorithm which achieves superior accuracy, is easier to use and computationally more efficient compared to other existing back-propagation-free methods. By departing from the traditional gradient based learning paradigm, NoProp alters how credit assignment is done within the network, enabling more efficient distributed learning as well as potentially impacting other characteristics of the learning process.
011
Reposted by Kyle Kastner
Max Slater @thenumb.at · 12/04/2025
Monte Carlo methods require randomly sampling complicated domains, which can be difficult in of itself. Part three (thenumb.at/Sampling/) discusses how to create samplers using rejection, inversion, and changes of coordinates.
37010
Kyle Kastner @kastnerkyle.bsky.social · 12/04/2025
Really, really cool work. You can try it yourself. Check the hygiene comments for some "adversarial" collaboration huggingface.co/papers/2504....
152
Kyle Kastner @kastnerkyle.bsky.social · 12/04/2025
Really enjoyed ICASSP this year. Hopefully will get a chance to give a summary/shout out to some interesting work I saw there in the upcoming days. Looking forward to the next one!
010
Reposted by Kyle Kastner
Sung Kim @sungkim.bsky.social · 11/04/2025
DDT: Decoupled Diffusion Transformer They've come up with a more efficient way to use diffusion models to generate high-quality images by breaking down the process into separate, specialized components.
1131
Reposted by Kyle Kastner
arxiv cs.CL @arxiv-cs-cl.bsky.social · 10/04/2025
Gabriel Grand, Joshua B. Tenenbaum, Vikash K. Mansinghka, Alexander K. Lew, Jacob Andreas Self-Steering Language Models arxiv.org/abs/2504.07081
011
Kyle Kastner @kastnerkyle.bsky.social · 08/04/2025
Dynamic evaluation (cf Generating Sequences with Recurrent Neural Networks) making a comeback under the umbrella of test-time methods
010
Reposted by Kyle Kastner
Tommy Thompson @tommy.aiandgames.com · 05/04/2025
To clarify: Microsoft hasn't created an AI version of Quake, but an AI model that simulates the behaviour of Quake based on existing play data at a reduced resolution. I just covered this AI trend on @aiandgames.com last month. youtu.be/9_oTroD9nzM?...
youtu.be
Explaining the Rise of AI Generated 'Games' | AI and Games #78
YouTube video by AI and Games
13510
Reposted by Kyle Kastner
Kosta Derpanis @csprofkgd.bsky.social · 07/04/2025
151
Reposted by Kyle Kastner
Sarath Chandar @sarath-chandar.bsky.social · 04/04/2025
Can better architectures & representations make self-play enough for zero-shot coordination? 🤔 We explore this in our ICLR 2025 paper: A Generalist Hanabi Agent. We develop R3D2, the first agent to master all Hanabi settings and generalize to novel partners! 🚀 #ICLR2025 1/n
1134
Reposted by Kyle Kastner
garreth @garrethlee.bsky.social · 16/12/2024
🚀 With Meta's recent paper replacing tokenization in LLMs with patches 🩹, I figured that it's a great time to revisit how tokenization has evolved over the years using everyone's favourite medium - memes! Let's take a trip down memory lane! [1/N]
4339
Reposted by Kyle Kastner
arxiv cs.CV @arxiv-cs-cv.bsky.social · 03/04/2025
Xiaohua Qi, Renda Li, Long Peng, Qiang Ling, Jun Yu, Ziyi Chen, Peng Chang, Mei Han, Jing Xiao Data-free Knowledge Distillation with Diffusion Models arxiv.org/abs/2504.00870
011