Sign in

Jonathan Lorraine

@jonlorraine.bsky.social
578 followers 1.9K following 44 posts

Research scientist @NVIDIA | PhD in machine learning @UofT. Previously @Google / @MetaAI. Opinions are my own. 🤖 💻 ☕️

PostsRepliesMedia
Reposted by Jonathan Lorraine
Masha Sh. @shumash.bsky.social · 28/07/2026
🦎 Did you know axolotls can regenerate complex body parts, including their brain? 🧠 Our new #AI model 𝗔𝘅𝗼𝗹𝗼𝘁𝗹𝟯𝗗 (#ECCV26) faithfully completes #3D objects based on partial views and incomplete point clouds: research.nvidia.com/labs/sil/pro... Kudos Anita Hu! #MachineLearning #NVIDIA
021
Reposted by Jonathan Lorraine
bioRxiv Neuroscience @biorxiv-neursci.bsky.social · 17/11/2025
From Tasks to Topology: Dorsal and Ventral Streams Emerge in Optimized Neural Networks www.biorxiv.org/content/10.1101/202…
022
Jonathan Lorraine @jonlorraine.bsky.social · 09/10/2025
Apply here: nvidia.eightfold.ai/careers?star... I'm personally interested in multimodal generation and the tools that power it.
nvidia.eightfold.ai
NVIDIA 2026 Internships: PhD Generative AI Research - US | NVIDIA Corporation
By submitting your resume, you're expressing interest in one of our 2026 Generative AI focused Research Internships. We'll review resumes on an ongoing basis, and a recruiter may reach out if your exp...
010
Jonathan Lorraine @jonlorraine.bsky.social · 09/10/2025
🔍 New NVIDIA Spatial Intelligence Lab internship postings for 2026. Come work with us to advance foundational technologies that enable AI systems to model and interact meaningfully with the world! Topics on our homepage: research.nvidia.com/labs/sil/ Application link below
research.nvidia.com
NVIDIA Spatial Intelligence Lab (SIL)
Advancing foundational technologies enabling AI systems to perceive, model, and interact with the world.
150
Reposted by Jonathan Lorraine
Masha Sh. @shumash.bsky.social · 06/06/2025
Join us at #CVPR2025 for a preview of this #NVIDIA tech during a live-coding session. A #GPU back end will be reserved for all attending – just don’t forget to bring your laptop for some hands-on fun! Wed, Jun 11, 8am-noon, or join in at 10:20 after the break. tinyurl.com/nv-kaolin-cv...
031
Jonathan Lorraine @jonlorraine.bsky.social · 09/05/2025
We find a new set of use cases for Stable Audio Open ( @jordiponsdotme.bsky.social, @stabilityai.bsky.social, @hf.co) and other large pretrained audio generative models, like AudioLDM and beyond!
010
Jonathan Lorraine @jonlorraine.bsky.social · 09/05/2025
Our work is inspired by and builds on the SDS update of DreamFusion (dreamfusion3d.github.io/, @benmpoole.bsky.social , @ajayjain9.bsky.social , @jonbarron.bsky.social), and related updates (VSD, SDI @vincentsitzmann.bsky.social, SJC, many more!)
110
Jonathan Lorraine @jonlorraine.bsky.social · 09/05/2025
💡 SDS treats any differentiable parameter set as optimizable from a prompt. Source-guided separation emerged when we brainstormed novel uses. We hope for similarly practical tasks to surface—e.g., automatic Foley layering?—as the community experiments.
110
Jonathan Lorraine @jonlorraine.bsky.social · 09/05/2025
🚀 Vision of the Future: Content designers easily use one video + audio diffusion backbone with SDS-style updates to nudge any differentiable task—impacts, lighting, cloth, fluids—until the joint model says “looks & sounds right” given powerful user controls, like text.
110
Jonathan Lorraine @jonlorraine.bsky.social · 09/05/2025
⚠️ Limitations ⚠️ Clip-Length Budget: We optimized on ≤10 s clips; minute-scale audio may have artifacts or blow up memory. A hierarchical/windowed Audio-SDS could help here.
110
Jonathan Lorraine @jonlorraine.bsky.social · 09/05/2025
⚠️ Limitations ⚠️ Audio-Model Bias: We rely on Stable Audio Open, so when this struggles, e.g., on rare instruments, speech, audio without silence at the end, or out-of-domain SFX, our method can have difficulties. Other diffusion models can help here.
110
Jonathan Lorraine @jonlorraine.bsky.social · 09/05/2025
This project was led by the great work of @jrichterpowell.bsky.social along with Antonio Torralba. See more work from the NVIDIA Spatial Intelligence Lab: research.nvidia.com/labs/toronto... Work supported indirectly by MIT CSAIL, @vectorinstitute.ai #nvidia #mit
110
Jonathan Lorraine @jonlorraine.bsky.social · 09/05/2025
Results on Prompt-Guided Source Separation: We report an improved SDR to ground-truth sources when available and show improved CLAP scores after training.
110
Jonathan Lorraine @jonlorraine.bsky.social · 09/05/2025
Results on Tuning FM Synthesizers & Impact Synthesis: We improve CLAP scores over training for prompts, along with qualitative results. Impact synthesis shows improved performance on impact-oriented prompts.
110
Jonathan Lorraine @jonlorraine.bsky.social · 09/05/2025
Results on Fully-Automatic In-the-Wild Source Separation: We demonstrate a pipeline that takes a video from the internet, captions the audio with a model (like AudioCaps), and provides that to an LLM-assistant who suggests source decompositions. We run our method on the suggested decompositions.
110
Jonathan Lorraine @jonlorraine.bsky.social · 09/05/2025
Modifications to SDS for Audio Diffusion: 🅰 We use an augmented Decoder-SDS in audio space, 🅱 using a spectrogram emphasis to better weight transients, and 🅲️ multiple denoising steps to increase fidelity. This image highlights these in red in the detailed overview of our update.
110
Jonathan Lorraine @jonlorraine.bsky.social · 09/05/2025
③ Prompt-Guided Source Separation: A prompt-conditioning source separation for a given audio, such as separating a “sax …” and “cars …” from a music recording on a road, by using the audio-SDS update for each channel while forcing the sum of channels to reconstruct the audio.
110
Jonathan Lorraine @jonlorraine.bsky.social · 09/05/2025
② Physical Impact Synthesis: We generate impacts consistent with prompts like “hitting pot with wooden spoon” by convolving an impact with a learned object and reverb impulse. We learn the parametrized forms of the object and reverb impulses.
110
Jonathan Lorraine @jonlorraine.bsky.social · 09/05/2025
① FM Synthesis: A toy setup where we generate settings aligning with prompts like “kick drum, bass, reverb” using sine oscillators modulating each other’s frequency as in a synthesizer. We visualize the final optimized parameters as the dial settings on a synthesizer instrument's user interface.
110
Jonathan Lorraine @jonlorraine.bsky.social · 09/05/2025
We propose three novel audio tasks: ① FM Synthesis, ② Physical Impact Synthesis, and ③ Prompt-Guided Source Separation. This image briefly summarizes the use case, optimizable parameters, rendering function, and parameter update.
110
Jonathan Lorraine @jonlorraine.bsky.social · 09/05/2025
Intuitively, our update finds a direction to move the audio to increase its probability given the prompt, by noising and denoising with our diffusion model, then “nudging” our audio towards it by propagating the update through our differentiable rendering to our audio parameters.
110
Jonathan Lorraine @jonlorraine.bsky.social · 09/05/2025
🔊 New NVIDIA paper: Audio-SDS 🔊 We repurpose Score Distillation Sampling (SDS) for audio, turning any pretrained audio diffusion model into a tool for diverse tasks, including source separation, impact synthesis & more. 🎧 Demos, audio examples, paper: research.nvidia.com/labs/toronto... 🧵below
161
Reposted by Jonathan Lorraine
Chih-Hao Lin @chih-hao.bsky.social · 02/05/2025
What if you could control the weather in any video — just like applying a filter? Meet WeatherWeaver, a video model for controllable synthesis and removal of diverse weather effects — such as 🌧️ rain, ☃️ snow, 🌁 fog, and ☁️ clouds — for any input video.
131
Jonathan Lorraine @jonlorraine.bsky.social · 12/12/2024
We envision a future where LLMs are universal generative tools capable of seamlessly producing content across multiple modalities, including text, images, videos, and 3D structures.
040
Jonathan Lorraine @jonlorraine.bsky.social · 12/12/2024
Integrating 3D mesh generation into LLMs opens exciting possibilities for interactive design. Users can converse with a model to create and manipulate 3D objects in real time.
120
Jonathan Lorraine @jonlorraine.bsky.social · 12/12/2024
We're excited to scale LLaMA-Mesh to handle more complex and detailed meshes by extending context lengths. Integrating textures and physical properties, exploring larger base models, part-based generation, and enabling dynamic generation are interesting ways forward!
100
Jonathan Lorraine @jonlorraine.bsky.social · 12/12/2024
Due to context length constraints, we're currently limited to meshes with up to 500 faces. We generate one 3D object per dialog due to our fine-tuning dataset construction. We see a slight degradation in language ability, perhaps due to using UltraChat in fine-tuning.
100
Jonathan Lorraine @jonlorraine.bsky.social · 12/12/2024
This project was led by Zhengyi Wang with Yikai Wang, Hang Su, Jun Zhu, Sanja Fidler, and Xiaohui Zeng. See more work from the #NVIDIA Toronto AI Lab here: research.nvidia.com/labs/toronto... Work supported by Tsinghua University, @vectorinst.bsky.social, @uoft.bsky.social #UofT #Tsinghua
110
Jonathan Lorraine @jonlorraine.bsky.social · 12/12/2024
We generate diverse and high-quality 3D meshes directly from textual prompts without expanding the vocabulary or introducing new tokenizers.
110
Jonathan Lorraine @jonlorraine.bsky.social · 12/12/2024
The model retains its language-understanding abilities, demonstrating coherent and contextually appropriate dialogues and being able to describe meshes in natural language.
110
Jonathan Lorraine @jonlorraine.bsky.social · 12/12/2024
LLaMA-Mesh achieves mesh generation quality comparable to specialized models trained from scratch on 3D data, as evidenced by qualitative comparisons with state-of-the-art methods like MeshXL.
110
Jonathan Lorraine @jonlorraine.bsky.social · 12/12/2024
We construct a dataset of text-mesh pairs and interleaved text-3D dialogues. We fine-tune a pre-trained LLaMA-3.1-8B-Instruct model on our curated dataset, allowing it to generate 3D meshes directly from text prompts and engage in conversational 3D content creation.
110
Jonathan Lorraine @jonlorraine.bsky.social · 12/12/2024
We quantize vertex coordinates in the OBJ files, reducing the token count, with minimal impact on geometric fidelity.
110
Jonathan Lorraine @jonlorraine.bsky.social · 12/12/2024
We represent 3D meshes using the OBJ file format, converting vertex coordinates and face definitions into plain text sequences that LLMs can process directly without modifying tokenizers or vocabularies.
110
Jonathan Lorraine @jonlorraine.bsky.social · 12/12/2024
We unify text and 3D mesh in a uniform format by representing the numerical values of vertex coordinates and face definitions of a 3D mesh as plain text. We train using text and 3D interleaved data end-to-end. With a single, unified model, we can generate both text and 3D meshes.
110
Jonathan Lorraine @jonlorraine.bsky.social · 12/12/2024
Check out the blender addon powered by LLaMA-Mesh from @dylanebert.bsky.social! It’s an impressive example of how mesh generation can be integrated into familiar creative workflows, streamlining the design process 🔥 🧩 Blender Addon: github.com/huggingface/...
120
Jonathan Lorraine @jonlorraine.bsky.social · 12/12/2024
Explore our interactive LLaMA-Mesh demo on @hf.co , built with @gradio-hf.bsky.social. Experiment with generating/understanding meshes from text input and examine the model’s performance firsthand—a step towards conversational 3D design workflows. 🕹️ Demo: huggingface.co/spaces/Zheng...
110
Jonathan Lorraine @jonlorraine.bsky.social · 12/12/2024
🦙New #NVIDIA paper: LLaMA-Mesh 🦙 We enable LLMs to generate 3D meshes by representing them as plain text and fine-tuning, unifying 3D and text modalities in a single model. 🔎 Webpage research.nvidia.com/labs/toronto... 🕹️ Interactive Demo huggingface.co/spaces/Zheng... 💾 Model checkpoint available
160
Jonathan Lorraine @jonlorraine.bsky.social · 04/12/2024
Check out our new #NVIDIA paper: ⚡️Multi-student Diffusion Distillation ⚡️ We make single-step distilled generators better and faster using our new method, multi-student distillation (MSD)! Explore the project page to learn more: research.nvidia.com/labs/toronto...
research.nvidia.com
Multi-Student Distillation
040
Jonathan Lorraine @jonlorraine.bsky.social · 27/11/2024
Huge thanks to my amazing collaborators. This project was led by Juhan Bae along with Wu Lin and @rogergrosse.bsky.social Supported (indirectly) by @anthropic.com , NVIDIA, @vectorinst.bsky.social, @uoft.bsky.social/ @uoftartsci.bsky.social
020
Jonathan Lorraine @jonlorraine.bsky.social · 27/11/2024
Want to learn more? Our work is featured at ml-data-tutorial.org. It's an excellent resource for anyone interested in training data attribution techniques.
ml-data-tutorial.org
Data Attribution at Scale | ICML 2024
Notes accompanying our ICML tutorial
120
Jonathan Lorraine @jonlorraine.bsky.social · 27/11/2024
By removing the top-k influential training points identified by SOURCE and retraining, we more accurately predicted changes in the model's behavior compared to other methods.
120
Jonathan Lorraine @jonlorraine.bsky.social · 27/11/2024
We tested SOURCE against other methods using the linear datamodeling score (LDS). SOURCE outperforms others, especially when models haven't fully converged or are trained in multiple stages!
120
Jonathan Lorraine @jonlorraine.bsky.social · 27/11/2024
Why does this matter? Understanding the impact of each training data point helps in: ·Diagnosing and fixing biases ·Improving model performance ·Enhancing transparency and trust in AI systems
120
Jonathan Lorraine @jonlorraine.bsky.social · 27/11/2024
How does SOURCE work? It divides the training process into segments and treats the gradients and Hessians as stable within each segment. This approximation gives us a formula that's both accurate and computationally efficient!
120
Jonathan Lorraine @jonlorraine.bsky.social · 27/11/2024
We develop SOURCE, combining the strengths of influence functions and unrolling methods without their limitations. It efficiently handles: ·Models that haven't converged ·Multi-stage training pipelines (like fine-tuning foundation models) ·Different training phases / optimizers (SGD, Adam, etc.)
120
Jonathan Lorraine @jonlorraine.bsky.social · 27/11/2024
🚨 New #NeurIPS2025 paper “Training Data Attribution via Approximate Unrolling” 🚨 Introducing SOURCE: A method to understand how individual training examples influence neural net behavior, allowing us to make AI models more transparent and trustworthy! 📄 Full paper: openreview.net/pdf?id=3NaqG...
1192