Sign in

Shantanu Singh

@shantanu-singh.cc
3.8K followers 886 following 62 posts

computation biology, drug discovery, computer vision, microscopy, statistics, machine learning, all happening at carpenter-singh-lab.broadinstitute.…

PostsRepliesMedia
Shantanu Singh @shantanu-singh.cc · 19/12/2024
Taking pictures of cells with a microscope, then extracting thousands of features from them is uncannily effective for quantifying cell state, esp. for genes and chemicals (e.g., Cell Painting). But we often average the rich single-cell data to simplify analysis. Can we do better? #bioML 🧪 1/n
(a) Human U2OS cells treated with dimethyl sulfoxide (DMSO) and stained using the Cell Painting assay, which employs six dyes in five channels to label eight cellular compartments. The top row (from left to right) shows mitochondrial staining; actin, Golgi, and plasma membrane staining; and nucleolar and cytoplasmic RNA staining. The bottom row (from left to right) displays endoplasmic reticulum staining, DNA staining, and a montage of all five channels (from Cimini et al. [21]). (b) Thousands of features are extracted from each segmented cell in microscopy images of wells. A learned function f(x) (CytoSummaryNet) aggregates this data into a single feature vector: the sample’s profile. (c) An in-depth look at the model architecture used in this study. The model consists of three elements: a function φ(x), which maps the input data from ℝD to ℝL space, a summation, which collapses the cell dimension, and ρ(z), which maps the collapsed representation from ℝN to ℝL space. (d) During training, replicate compound profiles are forced to attract each other (green arrows) and simultaneously repel every other compound (red arrows) in the learned feature space. Here, all forces are drawn for a single profile of compound B.
17318
Shantanu Singh @shantanu-singh.cc · 13/12/2024
🧪 So proud of this work by the dream team of @johnarevalo.bsky.social and Ellen Su: a new graph dataset for predicting drug-target interactions, using information from Cell Painting. Stop by their poster in a few hours @ #NeurIPS! (details below) PS: John is on the job market 🚀 #bioML #MLSky
Diagram showing 3.6k drugs, 11.5k genes, and the counts of "ground truth" relationships among them.
0203
Shantanu Singh @shantanu-singh.cc · 22/11/2024
Pasting this here for those who coming looking for the takeaway of this paper
Twelve critical questions to be asked by readers and reviewers when confronted with prediction models that are based on AI:

Is AI needed to solve the targeted medical problem?
How does the AI prediction model fit in the existing clinical workflow?
Are the data for prediction model development and testing representative for the targeted patient population and intended use?
Is the (time)point of prediction clear and aligned with the feature measurements?
Is the outcome variable labelling procedure reliable, replicable, and independent?
Was the sample size sufficient for AI prediction model development and testing?
Is optimism of predictive performance of the AI prediction model avoided?
Was the AI model’s performance evaluated beyond simple classification statistics?
Were the relevant reporting guidelines for AI prediction model studies followed?
Is algorithmic (un)fairness considered and appropriately addressed?
Is the developed AI prediction model open for use, further testing, critical appraisal, and updating and use in daily practice?
Are presented relations between individual features and the outcome not overinterpreted?
210
Shantanu Singh @shantanu-singh.cc · 19/11/2024
Absolutely delighted to see this paper and your explainer! The U-turn paper (and that neat schematic) had me intrigued but a bit dismayed that it couldn't (yet) shed light on NN learning. Can't wait to read the paper, but will likely rely on your explainer to get there faster :D
A key schematic from the  U-turn on Double Descent paper, which tried to demystify the double descent phenomenon, but it wasn't (yet) extensible to neural network learning
140
Shantanu Singh @shantanu-singh.cc · 11/11/2024
Related: I used NotebookLM to search through reviewer comments, internal discussions, and early versions of a 2023 manuscript (all in gdocs thankfully) to figure out if we had discussed testing a specific configuration. It did an excellent job of summarizing and citing the relevant sections!
An image containing text responding to the question "Did we ever discuss doing splits based on morphology, not just chemical scaffolds?". The text explains that the sources discuss using morphology to split data, but not for the same purpose as splitting by chemical scaffolds. It describes how splits based on morphology were used as a control experiment to identify potential biases in the data and reduce performance across data modalities, but were not intended for predicting assay activity of compounds with dissimilar morphology.
020
Shantanu Singh @shantanu-singh.cc · 15/01/2024
Thanks, both – yep, makes sense to use it to highlight the distinction between likelihood vs. density. For completeness for whomever comes looking, it's worth noting that the support for the two is different (0 to n vs. k to ∞).
The support of the two distributions is different. It is 0 to n for Binomial, and k to infinity for Negative Binomial.
000
Shantanu Singh @shantanu-singh.cc · 14/01/2024
Q for #stats folks: A 🤖 and I looked up intuitive terms for Binomial and Neg Binomial distributions: -'Fixed-Trials Binomial' (for Binomial, focus on fixed # of trials, n) -'Fixed-Successes Binomial' (for Neg Binomial, focus on fixed # of successes, k) Is this symmetry emphasized in courses?
Equations for binomial and negative binomial, emphasizing their similarity with each other
261