Sebastian Farquhar @sebfar.bsky.social · 23/01/2025By default, LLM agents with long action sequences use early steps to undermine your evaluation of later steps; a big alignment risk. Our new paper mitigates this, keeps the ability for long-term planning, and doesnt assume you can detect the undermining strategy. 👇 0131
Sebastian Farquhar @sebfar.bsky.social · 25/11/2024Help me grow this starter pack for technical researchers working on AGI safety! go.bsky.app/D6P44sC Some flex, but aiming for mostly technical research rather than governance/strategy. Who am I missing? 15289
Sebastian Farquhar @sebfar.bsky.social · 18/11/2024Starting to prepare yourself to submit to ICML? Here are my tips on how to write well for an ML research audience. sebastianfarquhar.com/on-research/...sebastianfarquhar.comHow to Write ML PapersThis doc is aimed at students learning to write ML papers as well as more experienced writers. It isn’t about how to do the research itself, but about how to present it in a way that makes it impactfu... 2152
Sebastian Farquhar @sebfar.bsky.social · 15/11/2024Entertaining essay about how the decline in practical engineering education has been devastating for *checks notes* professional criminal safe crackers. (Ok, mostly just a fun history of safe cracking.) www.timhunkin.com/94_illegal_e...timhunkin.comtimhunkin/illegal engineering 010
Sebastian Farquhar @sebfar.bsky.social · 13/11/2024Something I loved most about the internet in the 2000s was the idiosyncratic personal webpages that some people had put a crazy amount of time and effort into. These pages must still exist right? What are the best ones you know of? 161