Sign in

sina b

@sina.bio
203 followers 62 following 62 posts

HHMI Hanna Gray Fellow at @UCBerkeley w/ @airstreets. PhD @caltech, Math & ME BS @mit. I enjoy drinking tea, riding bikes, taking photos, and exploring nature.

PostsRepliesMedia
sina b @sina.bio · 29/06/2026
Does your paper really suck? Oded Rechavi, at QED Science, believes that if your paper is not in the top 1% of their QED score then it "sucks". But what is this QED score and what is its purpose? If a paper is not in the 1% does it really suck? My thoughts here: www.sina.bio/posts/does-y...
120
sina b @sina.bio · 11/02/2026
9/ For non-contiguous phrases, we ablated one interior token from each target and reran alignment. Naive matching failed almost entirely (<1%). Word-level tokenization recovered ~50%. Subword tokenization recovered >96%.
110
sina b @sina.bio · 11/02/2026
8/ We tested word-level and subword tokenization with four alignment tools (taln, LCS, difflib, and naive matching.) On contiguous phrases, naive succeeds as expected, but word-level tokenization yielded ~50% missed alignments. Subword tokenization brought accuracy back to ~99%.
100
sina b @sina.bio · 11/02/2026
7/ With ordered alignment selected, we asked, how much does tokenization matter? We built the BOAT dataset (Berkeley Ordered Alignment of Text) from the Stanford Question Answering Dataset (SQuAD, by Pranav Rajpurkar and Percy Liang), containing 35K source-target pairs (and BIO-BOAT from biorxiv).
100
sina b @sina.bio · 11/02/2026
4/ There's a long history of alignment methods in genomics and NLP. Pseudoalignment (@lpachter, @pmelsted uses set inclusion + k-mers. LCS uses order-preserving subsequences. To select an alignment method, we formalize alignment as a map from target to source.
110
sina b @sina.bio · 11/02/2026
2/ Take this sentence from a scRNA-seq paper (by @ADHildreth.) An LLM correctly extracts "natural killer cells" and "CD96" as a cell-type marker gene pair. Both exist in the sentence, but the parenthetical "(NK)" breaks contiguity, so naive string matching fails.
100
sina b @sina.bio · 03/02/2026
If you work in AI for Science, take a moment to familiarize yourself with a common failure mode: paranormal citations (or paracites). Our paper describes them.
082
sina b @sina.bio · 03/02/2026
Peer review is often opaque and confusing. @elife.bsky.social worked to change that. In a new preprint, we show how eLife’s Publish, Review, Curate model makes it possible to evaluate AI-generated reviews (with OpenEval) against human peer review. w/ @lauraluebbert.com and @lpachter.bsky.social
2189
sina b @sina.bio · 08/02/2025
In FY2022, Caltech received $342,234,517 in federal funding (representing 80% of contract grant funding) 19.7% of which went to the NIH. A cut to 15% indirect costs, from 70%, means a loss of ~ $37.1M annually from NIH grants alone. Across all grants, $188M poof, gone.
100
sina b @sina.bio · 08/02/2025
As can be seen in the graphic, indirect costs vary per institution but previously were around 60-70%. This means for every dollar a professor was awarded from the NIH, the university got 60-70 cents. My prior institution, @caltech.edu had a 70% indirect cost rate.
100
sina b @sina.bio · 08/02/2025
🧵 On a Friday night, the NIH twitter account announced the most significant change to research funding in decades. What are indirect costs, how are universities funded and what are the impacts?
130
sina b @sina.bio · 23/01/2025
Why are machine learning networks described as "neural"? The nomenclature arises because of the visual similarity between neural network architectures and connected neurons in the mammalian brain.
010
sina b @sina.bio · 23/01/2025
Reminder: there is little evidence that brain computation works in the same way as neural networks. Quote from "Understanding Deep Learning by Simon Prince (@simonprinceai.bsky.social)"
140
sina b @sina.bio · 17/12/2024
TIL about the watch command in the terminal: it reruns a command at set intervals, and is perfect for monitoring GPU usage or tracking real-time system updates.
010
sina b @sina.bio · 06/12/2024
141
sina b @sina.bio · 26/11/2024
The dreaded gene name turns date in Excel strikes again (in a published paper!).. this happens because MARCH3 is interpreted as March 3rd by Excel's automatic date formatting. (previously I was bamboozled by SEPT1 (Septin 1)!
000
sina b @sina.bio · 25/11/2024
I visited Hinxton, UK for an incredible conference ( #GenomeInformatics24, thx to @pmelsted.bsky.social, @zaminiqbal.bsky.social, and Nicky Mulder). And I got to see where the magic happens.
040
sina b @sina.bio · 24/11/2024
344 words you can spell on a calculator www.mathsquad.com/calculatorwo...
000
sina b @sina.bio · 23/11/2024
ChatGPT usage in the wild. @ebi.embl.org is using it to describe cell types on their Cell Type Ontologies site. www.ebi.ac.uk/ols4/ontolog...
000
sina b @sina.bio · 22/11/2024
I'm giving a talk in 30 min @ucberkeleyofficial.bsky.social on The Human Commons Cell Atlas- a project where we develop code and tools to build cell atlases from publicly available data. All are welcome to attend in person! (no zoom :/)
000
sina b @sina.bio · 21/11/2024
Hi, I’m Sina! A bioinformatics PhD @ucberkeleyofficial.bsky.social developing tools for single-cell data 🧬. When I’m not coding in a café ☕️ or biking around the bay area 🚴‍♂️, I’m planning experiments, staying active, and building things with my hands. Follow for science and life updates! 🌟
Sina writing code on a canoe. In the background is the campus of Kings College at Cambridge University in the UK.
2142