Cyrus Rashtchian @cyroid.bsky.social · 25/04/2025Curious about fine-grained text-to-image model evaluation? Come see our spotlight paper on Gecko 🦎 in the afternoon poster session at #ICLR25 🏆Hall 3 + Hall 2B #359 🎖️Friday 3pm ICLR: iclr.cc/virtual/2025... Paper: arxiv.org/abs/2404.16820 Prompts: github.com/google-deepm...iclr.ccICLR Poster Revisiting text-to-image evaluation with Gecko: on metrics, prompts, and human ratingICLR 2025 000
Cyrus Rashtchian @cyroid.bsky.social · 25/04/2025Why do LLMs hallucinate with RAG?! 🤔 Find out at my #ICLR25 poster on Sufficient Context! 👋🏼 📍Hall 3 + Hall 2B #230 ⏰ Fri 25 Apr 10 a.m. to 12:30 p.m. 000
Cyrus Rashtchian @cyroid.bsky.social · 25/04/2025Happy to chat with anyone at ICLR about RAG, LLMs, Factuality! 000
Reposted by Cyrus RashtchianHailey Joren @haileyjoren.bsky.social · 24/04/2025When RAG systems hallucinate, is the LLM misusing available information or is the retrieved context insufficient? In our #ICLR2025 paper, we introduce "sufficient context" to disentangle these failure modes. Work w Jianyi Zhang, Chun-Sung Ferng, Da-Cheng Juan, Ankur Taly, @cyroid.bsky.social 1116
Reposted by Cyrus RashtchianHossein Mobahi @thegradient.bsky.social · 20/12/2024 1/2 Just a reminder about Google Research Scholar Program, providing up to $60K unrestricted gifts to recognize early-career professors and support world-class research at institutions around the world. This year, we are particularly interested in the following research areas... 1125
Cyrus Rashtchian @cyroid.bsky.social · 13/12/2024[6/6] The other idea is to do the weighted combination at an instance level. We look at intermediate layers for *each token* and slightly modify the overall distribution. This leads to consistent accuracy improvements for many models and datasets! Would love to see some theory on why this works! 000
Cyrus Rashtchian @cyroid.bsky.social · 13/12/2024[5/6] Here's a nice example. We want to do some math. Greedy decoding leads to 5 x $10 = $50 for the overtime pay. This is cus A x B = C is a common pattern. But we really need A x B x C = D to get the answer. SLED can help with this because the internal layers happen to predict 'x' instead of '='. 100
Cyrus Rashtchian @cyroid.bsky.social · 13/12/2024[4/6] Our main decoding trick is to use a weighted combination of *all of the layers*. Precisely, we project the layers into the same output distribution (over vocab tokens). Then we combine the intermediate "logits" with the output logits based on our estimate of the LLM's internal knowledge 100
Cyrus Rashtchian @cyroid.bsky.social · 13/12/2024[3/6] The key observation is that LLMs "know" a lot more than they "tell" -- basically the training process can favor more popular tokens (in the dataset) rather than more accuracy predictions for the query at hand. So we can utilize this during decoding time... 100
Cyrus Rashtchian @cyroid.bsky.social · 13/12/2024[2/6] Joint work with Jianyi Zhang · Da-Cheng Juan · Chun-Sung Ferng · Heinrich Jiang · Yiran Chen ArXiv paper: arxiv.org/abs/2411.02433 Project page: jayzhang42.github.io/sled_page/ GitHub: github.com/JayZhang42/S... But how does it work you ask? 110
Cyrus Rashtchian @cyroid.bsky.social · 13/12/2024Longer thread about our new factuality decoding method SLED at NeurIPS 2024. Main idea: freeze the model, but be thoughtful about the decoding. With a small amount of extra inference-time compute, we increase accuracy by 3% on several benchmarks! SLED helps for all major open source models! 120
Cyrus Rashtchian @cyroid.bsky.social · 13/12/2024ArXiv paper: arxiv.org/abs/2411.02433 Project page: jayzhang42.github.io/sled_page/ GitHub: github.com/JayZhang42/S... 010
Cyrus Rashtchian @cyroid.bsky.social · 13/12/2024First shameless plug -- our new factuality decoding method, SLED gets SOTA improvements on 14+ models (Llama 2/3, Gemma, Mistral) & 9 benchmarks! See our #NeurIPS2024 poster today (Friday) in the East Exhibit Hall A-C #3311 100
Reposted by Cyrus RashtchianInês🩷 @loverines.swifties.social · 02/12/2024Hi friends!🩷 I have never done this but i’m making a list so and i can keep in touch with all of you more easily🫶🏻 please like this or say hi if i can add you🥰 Thank🫶🏻 22541
Reposted by Cyrus RashtchianPablo Samuel Castro @pcastr.bsky.social · 02/12/2024Everyone I spoke to at @rl-conference.bsky.social last summer agreed on it being one of the best conferences ever for an RL researcher... So many great RL-focused papers! CFP is out, send your work here! 14513
Cyrus Rashtchian @cyroid.bsky.social · 02/12/2024Excited to try out bluesky and chat about GenAI and ML theory! 170