Sign in

Sergey Feldman

@sergeyf.bsky.social
312 followers 405 following 35 posts

ML/AI at AI2 semanticscholar.org, alongside.care, data-cowboys.com

PostsRepliesMedia
Reposted by Sergey Feldman
Jonathan Bragg @jbragg.bsky.social · 06/11/2025
Agent benchmarks don't measure true *AI* advances We built one that's hard & trustworthy: 👉 AstaBench tests agents w/ *standardized tools* on 2400+ scientific research problems 👉 SOTA results across 22 agent *classes* 👉 AgentBaselines agents suite 🆕 arxiv.org/abs/2510.21652 🧵👇
AstaBench with abstract measurement icons
171
Reposted by Sergey Feldman
Ai2 @ai2.bsky.social · 26/03/2025
Meet Ai2 Paper Finder, an LLM-powered literature search system. Searching for relevant work is a multi-step process that requires iteration. Paper Finder mimics this workflow — and helps researchers find more papers than ever 🔍
Screenshot of the Ai2 Paper Finder interface
611723
Reposted by Sergey Feldman
Ai2 @ai2.bsky.social · 05/03/2025
Hope you’re enjoying Ai2 ScholarQA as your literature review helper 🥳 We’re excited to share some updates: 🗂️ You can now sign in via Google to save your query history across devices and browsers. 📚 We added 108M+ paper abstracts to our corpus - expect to get even better responses! More below…
Ai2 ScholarQA logo, with a red sign that says "Updated!"
1124
Reposted by Sergey Feldman
Ai2 @ai2.bsky.social · 21/01/2025
Can AI really help with literature reviews? 🧐 Meet Ai2 ScholarQA, an experimental solution that allows you to ask questions that require multiple scientific papers to answer. It gives more in-depth and contextual answers with table comparisons and expandable sections 💡 Try it now: scholarqa.allen.ai
Ai2 ScholarQA logo
13412
Sergey Feldman @sergeyf.bsky.social · 08/01/2025
000
Reposted by Sergey Feldman
Scott McGrath @smcgrath.phd · 13/12/2024
Building off the story I shared yesterday about fighting potential Insurance Company AI with AI: Claimable uses AI to tackle insurance claim denials. With an 85% success rate, it generates tailored appeals via clinical research and policy analysis. 🩺 #HealthPolicy
businessinsider.com
The CEO using AI to fight insurance-claim denials says he wants to remove the 'fearfulness' around getting sick
Claimable has helped patients file hundreds of health-insurance appeals. Its CEO says its success rate of overturning denials is about 85%.
1145
Sergey Feldman @sergeyf.bsky.social · 13/12/2024
Super awesome paper that directly addresses questions I've had for a while: arxiv.org/abs/2412.02674 Their experiments: (1) They get 128 responses from a LLM for some prompt. p = 0.9, t = 0.7, max length of 512 and 4-shot in-context samples 1/n
120
Reposted by Sergey Feldman
Katie Keith @katakeith.bsky.social · 11/12/2024
Check out our #NeurIPS2024 poster (presented by my collaborators Jacob Chen and Rohit Bhattacharya) about “Proximal Causal Inference With Text Data” at 5:30pm tomorrow (Weds)! neurips.cc/virtual/2024...
1124
Reposted by Sergey Feldman
Oregon 🕎🎲 @oregonthedm.bsky.social · 26/11/2024
Windows has issue: Person: fuck this I'm going to Linux Narrator: and they quickly learned to hate two operating systems.
3569745737
Sergey Feldman @sergeyf.bsky.social · 22/11/2024
Here are some research questions I'd like to get answers to. We are using LLMs to make training data for smaller, portable search or retrieval relevance models. (thread)
120
Reposted by Sergey Feldman
John Penniman @historiographos.bsky.social · 27/11/2023
I will never recover from this student email.
Image of an email from a student asking if sources "from the late 1900s" are acceptable.
37993822331
Sergey Feldman @sergeyf.bsky.social · 20/11/2023
www.semanticscholar.org/paper/Ground... I really like this paper. They study whether LLMs do reasonable things like ask follow-up questions and acknowledge what the users are saying. The answer is "not really".
111
Reposted by Sergey Feldman
Maria Antoniak @mariaa.bsky.social · 24/10/2023
An actually useful task for GPT-4: formatting my bibliography.
051
Reposted by Sergey Feldman
Maria Antoniak @mariaa.bsky.social · 20/10/2023
Do crowdworkers use ChatGPT to write their responses? I'm still not sure, but when I asked on Reddit, I got a flood of fascinating responses from the workers themselves, including some practical tips for researchers looking to prevent this.
reddit.com
Do you use ChatGPT or similar tools to complete tasks? : r/ProlificAc
3136
Sergey Feldman @sergeyf.bsky.social · 19/10/2023
Here's an interview with me about my work at alongside.care where I argue that everyone should take every August off. medium.com/authority-ma... #mlsky
010
Reposted by Sergey Feldman
Luca Soldaini 🎀 @soldaini.net · 13/10/2023
yayyy @ai2.bsky.social is on bsky 🥹
2103
Sergey Feldman @sergeyf.bsky.social · 11/10/2023
"Why Do We Need Weight Decay in Modern Deep Learning?" overparameterized deep networks -> WD changes enhances the implicit regularization of SGD underparameterized models trained with nearly online SGD -> WD balances bias/variance and lowers training loss. #mlsky arxiv.org/abs/2310.04415
arxiv.org
Why Do We Need Weight Decay in Modern Deep Learning?
Weight decay is a broadly used technique for training state-of-the-art deep networks, including large language models. Despite its widespread usage, its role remains poorly understood. In this...
041
Reposted by Sergey Feldman
Kyle Lo @ ICML2026 🇰🇷 @kylelo.bsky.social · 11/10/2023
manually curating a list of Twitter-Bluesky account mappings of NLP HCI folks, thinking it’ll be helpful for newcomers to Bluesky with migration friction. sinking way too much time into defining what qualifies as NLP HCI accounts. realizing I’m the sole annotator for my own annotation task 😵
4111
Reposted by Sergey Feldman
Vicente Valentim @valentimvicente.bsky.social · 24/09/2023
I think this is worth noting. Your language preferences on here may keep you from seeing posts containing only one word that is not in your chosen languages. This may keep you from seeing posts about feeds you’re following (e.g., pokisky, behavioursky). Turning all languages off fixes the issue.
45332
Reposted by Sergey Feldman
Kyle Lo @ ICML2026 🇰🇷 @kylelo.bsky.social · 06/10/2023
Don't forget to apply by *Oct 15* for AI2 research internships! Interested in language models of science, evaluating AI-generated text, challenging retrieval settings, and human-AI collaborative reading/writing? Come work with meeee! 😸 Learn more: kyleclo.github.io/mentorship
kyleclo.github.io
mentorship | Kyle Lo
Researcher at AI2 in Seattle. NLP + HCAI for scholars and scientists.
065
Reposted by Sergey Feldman
Alex Rubinsteyn @alexr.bsky.social · 06/10/2023
In what scenarios / model classes has double descent not been observed? (eg adding trees to an RF has marginal impact on accuracy, feature selection gets worse past the optimal number of features, &c)
211
Sergey Feldman @sergeyf.bsky.social · 05/10/2023
XGBoost @ v2.0. I love ensembles of GBDTs - feels like a mature technology where new features are not so exciting. Still it's nice that they are adding multi-label trees out-of-the-box. LightGBM (my preferred package) still can't do it. Do GBDTs work well in high dimensions or for text data?
121
Sergey Feldman @sergeyf.bsky.social · 27/09/2023
So hard!!!
A screenshot of an article about Linda Yaccarino (CEO of X) in the financial times. She's laying there in a read dress with bows with a slight smile. The quote from her is "It's hard on me"
020
Reposted by Sergey Feldman
Alex Rubinsteyn @alexr.bsky.social · 15/09/2023
Made an MLSky feed for machine learning research: bsky.app/profile/did:... Tag your posts with #mlsky I'll probably also add some ML jargon keywords later.
385
Reposted by Sergey Feldman
Alex Rubinsteyn @alexr.bsky.social · 15/09/2023
New platform, new semaglutide post. I try to be very public about using semaglutide to counteract the influence of misinformation: it's a great drug that saved me from a downward spiral of decreasing life quality (obesity -> hyperlipidemia -> statins -> joint & muscle injuries -> worse obesity) ⚕️💊
241
Reposted by Sergey Feldman
Kyle Lo @ ICML2026 🇰🇷 @kylelo.bsky.social · 14/07/2023
impressions @ACL23: * retrieval-augmented / multimodal getting crowded * better eval requires moving beyond “autometric correlate w humans” * ppl aspire to try systems HCI style work but need time to adopt research methods * corpus linguistics / CSS w LLMs exciting * smaller confs better but FOMO 🤷🏻‍♂️
053
Sergey Feldman @sergeyf.bsky.social · 03/07/2023
Best historical parallels to the OpenAI/Anthropic "have the best tech and everyone is trying hard to catch up" duopoly? How long do such monopolies/duopolies tend to last before everyone else figures it out?
010