sina b @sina.bio · 29/06/2026Does your paper really suck? Oded Rechavi, at QED Science, believes that if your paper is not in the top 1% of their QED score then it "sucks". But what is this QED score and what is its purpose? If a paper is not in the 1% does it really suck? My thoughts here: www.sina.bio/posts/does-y... 120
Reposted by sina bGennady Gorin @goringennady.bsky.social · 02/03/2026Hi everyone, I am looking for a new industry role in computational biology! Check out my portfolio of genomics, statistics, ML, and biophysics work at gennadygorin.github.io, and reach out if you have any suggestions or open roles!gennadygorin.github.ioGennady Gorin, Ph.D.Senior Scientist applying stochastic models for therapeutic discovery 121
Reposted by sina bRoland Gromes @gromesroland.bsky.social · 20/02/2026"We therefore argue that publication systems should optimize separately for the dissemination of data and results versus the conveying of novel ideas, and the former should be machine-readable." - This is such a good point. Machine readability is a goal too often viewed as just an addon to > 121
sina b @sina.bio · 11/02/202614/ In practice, extract passages from documents, align structured fields (gene names, lab values, etc.) with `taln`, and flag extractions that don't align. `taln` recovers alignments that other tools miss, reducing the number of extractions that need manual review. 010
sina b @sina.bio · 11/02/2026 13/ The BOAT datasets and `taln` tool can be found here: github.com/sbooeshaghi/...github.comGitHub - sbooeshaghi/taln: non-contiguous token alignmentnon-contiguous token alignment. Contribute to sbooeshaghi/taln development by creating an account on GitHub. 110
sina b @sina.bio · 11/02/202612/ To summarize, verifying LLM-extracted text doesn't require another LLM. Subword tokenization + ordered alignment recovers ~50% more phrases than word-level tokenization, deterministically, in milliseconds. We release `taln`, BOAT, and BIO-BOAT as open-source tools and benchmarks. 110
sina b @sina.bio · 11/02/202611/ Unlike LCS and difflib, `taln` enumerates all valid alignments. On BOAT, that's ~8 and on BIO-BOAT ~90 (likely because biomedical terms like "VEGF-A165b" get split into 5+ subword tokens, combinatorially increasing matches.) Domain-specific tokenizers would help here. 100
sina b @sina.bio · 11/02/202610/ The remaining errors mostly come from subword boundaries that don't match how the phrase was originally selected. Edge tokens often have inconsistent start/end positions that change where the tokenizer splits the phrase. 100
sina b @sina.bio · 11/02/20269/ For non-contiguous phrases, we ablated one interior token from each target and reran alignment. Naive matching failed almost entirely (<1%). Word-level tokenization recovered ~50%. Subword tokenization recovered >96%. 110
sina b @sina.bio · 11/02/20268/ We tested word-level and subword tokenization with four alignment tools (taln, LCS, difflib, and naive matching.) On contiguous phrases, naive succeeds as expected, but word-level tokenization yielded ~50% missed alignments. Subword tokenization brought accuracy back to ~99%. 100
sina b @sina.bio · 11/02/20267/ With ordered alignment selected, we asked, how much does tokenization matter? We built the BOAT dataset (Berkeley Ordered Alignment of Text) from the Stanford Question Answering Dataset (SQuAD, by Pranav Rajpurkar and Percy Liang), containing 35K source-target pairs (and BIO-BOAT from biorxiv). 100
sina b @sina.bio · 11/02/20266/ Ordered alignment sits at the practical sweet spot in this hierarchy. It allows gaps while preserving order, keeping the number of alignments manageable. We implemented this in a Python tool called `taln`, inspired by pseudoalignment tools like kallisto. github.com/sbooeshaghi/talngithub.comGitHub - sbooeshaghi/taln: non-contiguous token alignmentnon-contiguous token alignment. Contribute to sbooeshaghi/taln development by creating an account on GitHub. 100
sina b @sina.bio · 11/02/20265/ Alignment methods form a hierarchy based on the constraints they impose on this map. Contiguous alignment -> ordered -> permutation -> rearrangement. Each relaxation admits more alignments but the search space grows fast—from linear to factorial to exponential. 100
sina b @sina.bio · 11/02/20264/ There's a long history of alignment methods in genomics and NLP. Pseudoalignment (@lpachter, @pmelsted uses set inclusion + k-mers. LCS uses order-preserving subsequences. To select an alignment method, we formalize alignment as a map from target to source. 110
sina b @sina.bio · 11/02/20263/ To find these phrases, you need an alignment method that allows gaps, and a tokenization strategy that doesn't merge the text you're looking for with surrounding punctuation (both matter!) We studied how much. 100
sina b @sina.bio · 11/02/20262/ Take this sentence from a scRNA-seq paper (by @ADHildreth.) An LLM correctly extracts "natural killer cells" and "CD96" as a cell-type marker gene pair. Both exist in the sentence, but the parenthetical "(NK)" breaks contiguity, so naive string matching fails. 100
sina b @sina.bio · 11/02/20261/ LLMs are great at text extraction, but sometimes they hallucinate. A simple way to catch hallucinations is to check if the extracted text actually exists in the source. Turns out this is harder than it sounds. (new paper with Aaron Streets) www.biorxiv.org/content/10.6...biorxiv.org 112
Reposted by sina bLior Pachter @lpachter.bsky.social · 11/02/2026Interesting article on LLM text extraction and the connection of that problem to (computational biology) sequence alignment. www.biorxiv.org/content/10.6... by @sina.bio and Aaron Streets.biorxiv.org 081
sina b @sina.bio · 04/02/2026I agree that machine readability enables these analyses in principle. Developing robust, grounded benchmarks will be essential for evaluating the utility of these kinds of analysis. 000
Reposted by sina bRichard Sever @richardsever.bsky.social · 03/02/2026"publication systems [should] distinguish between dissemination of results & communication of ideas, and optimize them separately. Results should be in explicit, machine-readable form, while narrative text serves as an interpretive layer for human readers" www.biorxiv.org/content/10.6...biorxiv.org 1156
sina b @sina.bio · 03/02/2026If you work in AI for Science, take a moment to familiarize yourself with a common failure mode: paranormal citations (or paracites). Our paper describes them. 082
Reposted by sina bLior Pachter @lpachter.bsky.social · 03/02/2026AI hallucinations in science manuscripts are a nuisance. Paranormal citations, or paracites, will be a nightmare. www.biorxiv.org/content/10.6... (w/ @sina.bio & @lauraluebbert.com). 23211
sina b @sina.bio · 03/02/2026Peer review is often opaque and confusing. @elife.bsky.social worked to change that. In a new preprint, we show how eLife’s Publish, Review, Curate model makes it possible to evaluate AI-generated reviews (with OpenEval) against human peer review. w/ @lauraluebbert.com and @lpachter.bsky.social 2189
Reposted by sina bGennady Gorin @goringennady.bsky.social · 06/03/2025single-cell is a fun field. for instance, one of the heavily curated bixbench scenarios is about interpreting the results of a sc analysis and comparing to ground truth. this ground truth is, of course, based on DE analyses with some truly remarkable p-values for n=5 221
Reposted by sina bC. Brandon Ogbunu @cbo.bsky.social · 06/03/2025"No one is coming out of the sky to give you your grant money. Your citation portfolio won’t survive this market crash. Your credentials mean nothing. Everything is going to change." New for @undark.org undark.org/2025/03/06/o...undark.orgHow Science Can Adapt to a New NormalOpinion | In the wake of attacks on the research enterprise, scientists need to focus on protecting its fragile infrastructure. 516879
Reposted by sina bDavide CIttaro @daweonline.bsky.social · 23/02/2025Not really my field but I love the abstract! www.biorxiv.org/content/10.1...biorxiv.orgIs Tanimoto a metric?No. However, here we show how to generate a metric consistent with the Tanimoto similarity. We also explore new properties of this index, and how it relates to other popular alternatives. ### Competi... 021
sina b @sina.bio · 12/02/2025“blocking retro nasal sensation with a nose clip significantly reduces the subjective and objective neural responses to sucrose taste” 000
Reposted by sina bRichard McElreath 🐈⬛ @rmcelreath.bsky.social · 10/02/2025Amazing to me how useful looking at data in 2D PCA continues to be, even though the approach sounds crazy on paper—"p-dimensional ellipsoid", rantings of a madman. PCA is the cockroach of dimension reduction. I expect it to be present in any advanced galactic civilization. 38812
Reposted by sina bC. Brandon Ogbunu @cbo.bsky.social · 06/02/2025"One thing is certain: The changes we make ourselves will be healthier than the ones our adversaries demand." New work for @undark.org: undark.org/2025/02/06/o...undark.orgThe End of Science’s PeacetimeOpinion | Defending the practice of science from its adversaries will require dealing with some uncomfortable truths. 613278
sina b @sina.bio · 08/02/2025Lastly, the fact that this announcement, from a historically legitimate scientific agency, was made exclusively on Twitter, on a Friday evening, is frankly, shameful. 020
sina b @sina.bio · 08/02/2025What are the implications of this rate cut? I'm no expert but these outcomes may be impacted: - fewer faculty hires - department closures - fewer students (undergrad/grad), harder for int'l students - loss of US scientific hegemony 110
sina b @sina.bio · 08/02/2025In FY2022, Caltech received $342,234,517 in federal funding (representing 80% of contract grant funding) 19.7% of which went to the NIH. A cut to 15% indirect costs, from 70%, means a loss of ~ $37.1M annually from NIH grants alone. Across all grants, $188M poof, gone. 100
sina b @sina.bio · 08/02/2025As can be seen in the graphic, indirect costs vary per institution but previously were around 60-70%. This means for every dollar a professor was awarded from the NIH, the university got 60-70 cents. My prior institution, @caltech.edu had a 70% indirect cost rate. 100
sina b @sina.bio · 08/02/2025When a professor at a university receives a research grant, they don't just get money for experiments and equipment. A percentage of the grant goes to the university as "indirect costs" - money used to "keep the lights on", administrative overhead, facilities, and infra. 100
sina b @sina.bio · 08/02/2025🧵 On a Friday night, the NIH twitter account announced the most significant change to research funding in decades. What are indirect costs, how are universities funded and what are the impacts? 130
sina b @sina.bio · 06/02/2025Is cause of “heavy traffic” known? Has this happened before? The thought of losing these resources is terrifying,. 110
sina b @sina.bio · 23/01/2025Why are machine learning networks described as "neural"? The nomenclature arises because of the visual similarity between neural network architectures and connected neurons in the mammalian brain. 010
sina b @sina.bio · 23/01/2025Reminder: there is little evidence that brain computation works in the same way as neural networks. Quote from "Understanding Deep Learning by Simon Prince (@simonprinceai.bsky.social)" 140
sina b @sina.bio · 23/01/2025why is this desirable? Language models are helpful in parsing unstructured data, but QC reports are already structured... 020
sina b @sina.bio · 17/01/2025An often overlooked point in genomics: "Molecular omics resources should require sex annotation: a call for action" by @gliomath.bsky.social www.nature.com/articles/s41...nature.comMolecular omics resources should require sex annotation: a call for action - Nature MethodsThe most commonly used omics databases are a compilation of results from primarily male-only and sex-agnostic studies. The pervasive use of these databases critically hinders progress toward fully acc... 141
sina b @sina.bio · 17/12/2024Important to clearly define "genomic data"- certain data types may get better performance than others. Metadata storage seems to work well for my use-case. cc @manzt.sh 120
sina b @sina.bio · 17/12/2024TIL about the watch command in the terminal: it reruns a command at set intervals, and is perfect for monitoring GPU usage or tracking real-time system updates. 010
Reposted by sina bEMBL-EBI @ebi.embl.org · 17/12/2024In this year's user survey, 89% of respondents said that EMBL-EBI data resources empowered them to undertake work that would otherwise not have been possible 💪 Big thank you to everyone who filled in the survey - we appreciate your input! Explore further findings: www.ebi.ac.uk/about/news/a...ebi.ac.ukFuelling discovery together: 2024 user survey learningsBlog post by Eleni Tzampatzopoulou, EMBL-EBI Impact Manager In summer 2024, EMBL-EBI ran a user survey, inviting our community to let us know how they use the open data resources we jointly manage wit... 0153
sina b @sina.bio · 13/12/2024Cool to see our open source syringe pumps being made in the wild ! Original article: www.nature.com/articles/s41...nature.comPrinciples of open source bioinstrumentation applied to the poseidon syringe pump system - Scientific ReportsScientific Reports - Principles of open source bioinstrumentation applied to the poseidon syringe pump system 040
Reposted by sina bJase Gehring @skyjase.bsky.social · 08/12/2024gonna post up in a cafe and speedrun a new protein diffusion model "from scratch" with Claude live poasting my way thru it, public Git repo last did this in June and it's really at the edge of both of our capabilities 041
Reposted by sina bJase Gehring @skyjase.bsky.social · 06/12/2024I often make microfluidic emulsions and microparticles for scalable biology experiments, but the other day I was thinking, what if we polymerize the continuous phase of an emulsion? come learn about a delightful little branch of science, polyHIPEs! 121
sina b @sina.bio · 27/11/2024I am honestly stoked by the ability to customize one's algorithm on @bsky.app. It's a killer feature (over Twitter) that makes it actually useable for scientific communication. Simply pick the feed you want to follow and boom you get the relevant content. 010
Reposted by sina bFranciska de Vries 🟥 @frantecol.bsky.social · 27/11/2024It’s out!! We subjected soils from 30 different locations across Europe to extreme events and found that soil fungal and bacterial communities showed consistent responses that could be predicted from their origin! With @knightjar.bsky.social and many collaborators! www.nature.com/articles/s41...nature.comSoil microbiomes show consistent and predictable responses to extreme events - NatureSoils from 30 grasslands across Europe were subjected to 4 contrasting extreme climatic events under drought, flood, freezing and heat conditions, with the results suggesting that soil microbiomes fro... 16571164