Tanishq Mathew Abraham @iscienceluvr.bsky.social · 04/09/2025Has anyone successfully done RL post-training of GPT-oss with meaningful performance gains? What libraries even support it? I guess technically TRL/axolotl, maybe Unsloth... but there are no good examples of doing it... 291
Tanishq Mathew Abraham @iscienceluvr.bsky.social · 01/04/2025I have EXCITING news: I've started a company! Introducing Sophont We’re building open multimodal foundation models for the future of healthcare. We need a DeepSeek for medical AI, and @sophontai.bsky.social will be that company! Check out our website & blog post for more info (link below) 1302
Tanishq Mathew Abraham @iscienceluvr.bsky.social · 17/03/2025NEW BLOG POST: LLMs in medicine: evaluations, advances, and the future www.tanishq.ai/blog/posts/l... A short blog post discussing how LLMs are evaluated for medical capabilities and what's the future for LLMs in medicine (spoiler: it's reasoning!) 1212
Tanishq Mathew Abraham @iscienceluvr.bsky.social · 19/02/2025I restarted my blog a few weeks ago. The 1st post was: Debunking DeepSeek Delusions I discussed 5 main myths that I saw spreading online back during the DeepSeek hype. It may be a little less relevant now, but hopefully still interesting to folks. Check it out → www.tanishq.ai/blog/posts/d... 4211
Tanishq Mathew Abraham @iscienceluvr.bsky.social · 20/01/2025Okay so this is so far the most important paper in AI of the year 2231
Tanishq Mathew Abraham @iscienceluvr.bsky.social · 11/01/2025Anthropic, please add a higher tier plan for unlimited messages 😭🙏 4160
Tanishq Mathew Abraham @iscienceluvr.bsky.social · 10/01/2025Decentralized Diffusion Models UC Berkeley and Luma AI introduce Decentralized Diffusion Models, a way to train diffusion models on decentralized compute with no communication between nodes. abs: arxiv.org/abs/2501.05450 project page: decentralizeddiffusion.github.io 0202
Tanishq Mathew Abraham @iscienceluvr.bsky.social · 10/01/2025The GAN is dead; long live the GAN! A Modern Baseline GAN This is a very interesting paper, exploring making GANs simpler and more performant. abs: arxiv.org/abs/2501.05441 code: github.com/brownvc/R3GAN 0132
Tanishq Mathew Abraham @iscienceluvr.bsky.social · 10/01/2025Happy birthday to my incredible and awesome Mamma! 🥳🎉🎂 To many more years of health and happiness. Tiara (my sister) and I love you very much ❤️❤️❤️ 190
Tanishq Mathew Abraham @iscienceluvr.bsky.social · 28/12/2024Happy 19th birthday to my amazing sister Tiara Abraham! 🥳🎉 🎂 Proud of you graduating with your Master's degree at 18 and starting your doctorate in music degree this past year! Excited to see what this final teen year holds for you! 0150
Tanishq Mathew Abraham @iscienceluvr.bsky.social · 10/12/2024Inventors of flow matching have released a comprehensive guide going over the math & code of flow matching! Also covers variants like non-Euclidean & discrete flow matching. A PyTorch library is also released with this guide! This looks like a very good read! 🔥 arxiv: arxiv.org/abs/2412.06264 110927
Tanishq Mathew Abraham @iscienceluvr.bsky.social · 10/12/2024Normalizing Flows are Capable Generative Models Apple introduces TarFlow, a new Transformer-based variant of Masked Autoregressive Flows. SOTA on likelihood estimation for images, quality and diversity comparable to diffusion models. arxiv.org/abs/2412.06329arxiv.orgNormalizing Flows are Capable Generative ModelsNormalizing Flows (NFs) are likelihood-based models for continuous inputs. They have demonstrated promising results on both density estimation and generative modeling tasks, but have received relati... 1549
Tanishq Mathew Abraham @iscienceluvr.bsky.social · 10/12/2024Refusal Tokens: A Simple Way to Calibrate Refusals in Large Language Models "We introduce a simple strategy that makes refusal behavior controllable at test-time without retraining: the refusal token." arxiv.org/abs/2412.06748 061
Tanishq Mathew Abraham @iscienceluvr.bsky.social · 10/12/2024Can foundation models actively gather information in interactive environments to test hypotheses? "Our experiments with Gemini 1.5 reveal significant exploratory capabilities" arxiv.org/abs/2412.06438 0101
Tanishq Mathew Abraham @iscienceluvr.bsky.social · 10/12/2024Training Large Language Models to Reason in a Continuous Latent Space Introduces a new paradigm for LLM reasoning called Chain of Continuous Thought (COCONUT) Directly feed the last hidden state (a continuous thought) as the input embedding for the next token. arxiv.org/abs/2412.06769 2528
Tanishq Mathew Abraham @iscienceluvr.bsky.social · 10/12/2024[MASK] is All You Need New paper from CompVis group, introduces a new method called Discrete Interpolants that builds on top of discrete flow matching. Achieves SOTA performance on MS-COCO, competitive results on ImageNet 256. arxiv.org/abs/2412.06787 1326
Tanishq Mathew Abraham @iscienceluvr.bsky.social · 09/12/2024A new tutorial on RL by Kevin Patrick Murphy, a Research Scientist at Google DeepMind who also wrote several comprehensive, well-regarded textbooks on ML/DL. This ought to be a good read 👀 arxiv.org/abs/2412.05265 1385
Tanishq Mathew Abraham @iscienceluvr.bsky.social · 09/12/2024Birth and Death of a Rose abs: arxiv.org/abs/2412.05278 Generating temporal object intrinsics - temporally evolving sequences of object geometry, reflectance, and texture, such as blooming of a rose - from pre-trained 2D foundation models. 0130
Tanishq Mathew Abraham @iscienceluvr.bsky.social · 09/12/2024Frontier Models are Capable of In-context Scheming abs: arxiv.org/abs/2412.04984 "Our results show that o1, Claude 3.5 Sonnet, Claude 3 Opus, Gemini 1.5 Pro, and Llama 3.1 405B all demonstrate in-context scheming capabilities" 0111
Tanishq Mathew Abraham @iscienceluvr.bsky.social · 09/12/2024BigDocs: An Open and Permissively-Licensed Dataset for Training Multimodal Models on Document and Code Tasks abs: arxiv.org/abs/2412.04626 project page: bigdocs.github.io BigDocs-7.5M is a high-quality, open-access dataset comprising 7.5 million multimodal documents across 30 tasks. 0121
Tanishq Mathew Abraham @iscienceluvr.bsky.social · 09/12/2024Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling abs: arxiv.org/abs/2412.05271 model: huggingface.co/OpenGVLab/In... Introduces new InternVL-2.5 model, the first open-source MLLMs to surpass 70% on the MMMU benchmark 0101
Tanishq Mathew Abraham @iscienceluvr.bsky.social · 06/12/2024NVILA: Efficient Frontier Visual Language Models abs: arxiv.org/abs/2412.04468 NVIDIA introduces NVILA, a family of open VLMs designed to optimize both efficiency and accuracy. 0121
Tanishq Mathew Abraham @iscienceluvr.bsky.social · 06/12/2024Infinity: Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis abs: arxiv.org/abs/2412.04431 New visual autoregression framework that performs bitwise token prediction w/ an infinite-vocabulary tokenizer & classifier, a new record for autoregressive text-to-image models. 082
Reposted by Tanishq Mathew AbrahamNick Stracke @rmsnorm.bsky.social · 04/12/2024🤔 Why do we extract diffusion features from noisy images? Isn’t that destroying information? Yes, it is - but we found a way to do better. 🚀 Here’s how we unlock better features, no noise, no hassle. 📝 Project Page: compvis.github.io/cleandift 💻 Code: github.com/CompVis/clea... 🧵👇 24210
Tanishq Mathew Abraham @iscienceluvr.bsky.social · 04/12/2024Leading computer vision researchers Lucas Beyer (@giffmana.ai), Alexander Kolesnikov (@kolesnikov.ch), Xiaohua Zhai have left Google DeepMind to join OpenAI! They were behind recent SOTA vision approaches and open-source models like ViT, SigLIP, PaliGemma 1190
Tanishq Mathew Abraham @iscienceluvr.bsky.social · 02/12/2024the restrictions on post and video length is gonna make it harder to paper-post here ngl 3101
Tanishq Mathew Abraham @iscienceluvr.bsky.social · 02/12/2024Reverse Thinking Makes LLMs Stronger Reasoners abs: arxiv.org/abs/2411.19865 Train an LLM to be able to generate forward reasoning from question, backward question, and backward reaoning from backward question Shows an average 13.53% improvement over the student model’s zero-shot performance 05410
Tanishq Mathew Abraham @iscienceluvr.bsky.social · 02/12/2024GaussianSpeech: Audio-Driven Gaussian Avatars abs: arxiv.org/abs/2411.18675 project page: shivangi-aneja.github.io/projects/gau... 080
Tanishq Mathew Abraham @iscienceluvr.bsky.social · 02/12/2024some people managed to find some AoC-solving code from qianxyz in a github repo that has now been deleted seems like an automated pipeline using gpt-4o-mini with a pretty basic prompt 1121
Reposted by Tanishq Mathew AbrahamTanishq Mathew Abraham @iscienceluvr.bsky.social · 01/12/2024how does someone solve Advent of Code problem in 9 seconds??!! 4121
Tanishq Mathew Abraham @iscienceluvr.bsky.social · 01/12/2024how does someone solve Advent of Code problem in 9 seconds??!! 4121
Tanishq Mathew Abraham @iscienceluvr.bsky.social · 28/11/2024At #XPANSE in Abu Dhabi last week: - Met @anilseth.bsky.social backstage between our talks, discussed studying the nature of consciousness w/ neuroimaging. Appreciated him gifting me a signed copy of his book! - Met @seanmcarroll.bsky.social who gave a great talk about entropy vs. complexity 1181
Tanishq Mathew Abraham @iscienceluvr.bsky.social · 27/11/2024Many SOTA image generation models use an adversarial loss (VAE for latent diffusion for example), which counts I would say... 170
Tanishq Mathew Abraham @iscienceluvr.bsky.social · 25/11/2024My Bluesky follower count (1.6k followers) has now surpassed my Threads follower count (1.1k). I still see a few AI folks on Threads but it seems so much more dead compared to BlueSky. 4351
Tanishq Mathew Abraham @iscienceluvr.bsky.social · 25/11/2024Every time conference reviews and rebuttals come in we hear complaints about how bad the process is. Which ML conference has the best review process and what's stopping other conferences from improving their processes? 1100
Reposted by Tanishq Mathew AbrahamYoshitomo Matsubara @yoshitomo-matsubara.net · 21/11/2024Here is a list of ML OSS & Open Source / Science enthusiasts I found on Bluesky 🦋 go.bsky.app/8MFcfXd Let me know if you find such people here! I'm still new here and probably the list misses many must-add people, so let's built it together💪 4011249
Tanishq Mathew Abraham @iscienceluvr.bsky.social · 21/11/2024I hope I get added to some starter packs 4271
Tanishq Mathew Abraham @iscienceluvr.bsky.social · 21/11/2024Why does he follow me if no one wants me there? 🤔 2100
Tanishq Mathew Abraham @iscienceluvr.bsky.social · 19/11/2024What should I post here? Anything different from Twitter? 4140
Tanishq Mathew Abraham @iscienceluvr.bsky.social · 18/11/2024Excited to talk at XPANSE WORLD in Abu Dhabi, looking forward to meeting other esteemed speakers! 020
Reposted by Tanishq Mathew AbrahamGabriele Corso @gcorso.bsky.social · 17/11/2024Thrilled to announce Boltz-1, the first open-source and commercially available model to achieve AlphaFold3-level accuracy on biomolecular structure prediction! An exciting collaboration with Jeremy, Saro, and an amazing team at MIT and Genesis Therapeutics. A thread! 18610204
Tanishq Mathew Abraham @iscienceluvr.bsky.social · 30/06/2023A reminder: if you want it done right do it yourself 000
Tanishq Mathew Abraham @iscienceluvr.bsky.social · 07/06/2023Posting on LinkedIn or Instagram can be extremely infuriating and I am grateful that Twitter managed to get it correct and luckily remains fairly functional over this past year. Other social media platforms would do well to imitate some of Twitter's awesome posting features. 000
Tanishq Mathew Abraham @iscienceluvr.bsky.social · 05/06/2023Twitter and YouTube heavily pushing Across the SpiderVerse spoilers 😩 Wasn't even this bad when NWH was released... 000
Tanishq Mathew Abraham @iscienceluvr.bsky.social · 01/06/2023I'm really excited to share MedARC's first paper since our public launch 🥳 🧠👁️ MindEye! Our state-of-the-art fMRI-to-image approach that retrieves and reconstructs images from brain activity! Project page: medarc-ai.github.io/mindeye arXiv: arxiv.org/abs/2305.18274 040