Kush Jain @kjain14.bsky.social · 19/12/2024Thrilled to announce our new work TestGenEval, a benchmark that measures unit test generation and test completion capabilities. This work was done in collaboration with the FAIR CodeGen team. Preprint: arxiv.org/abs/2410.00752 Leaderboard: testgeneval.github.io/leaderboard.... 1177
Reposted by Kush JainCatarina Gamboa @catarinavgamboa.bsky.social · 26/11/2024Hi, Bluesky! 👋 I’m Catarina, a dual PhD student in 🖥️ Software Engineering with the CMU Portugal program ( @carnegiemellon.bsky.social and U. Lisbon). Imagine a world with reliable software and user-friendly verification tools. Let’s build it together! 🚀 #PhDlife #SE #PL #HCI #CMU-Portugal 0164
Reposted by Kush JainDr. Claire Le Goues @clegoues.bsky.social · 26/11/2024And now that we’re all here, some work!🚨 Are Large Language Models Memorizing Bug Benchmarks? 🚨 There’s growing concern that LLMs for SE are prone to data leakage, but no one has quantified it... until now. 🕵️♂️ 1/arxiv.org 26411