Sign in

Nathan Schaefer

@nkschaefer.bsky.social
29 followers 38 following 65 posts

UCSF postdoc, human, mammal

PostsRepliesMedia
Nathan Schaefer @nkschaefer.bsky.social · 12/03/2026
When species diverge, gene networks can accumulate sets of new mutations that are compatible and preserve network function within each species, but which can cause dysfunction when brought together in a hybrid. Does GC help alleviate this problem in P. formosa? www.nature.com/scitable/con...
100
Nathan Schaefer @nkschaefer.bsky.social · 12/03/2026
We then looked for evidence of positive selection in noncoding sequence and found that gene conversion copies sequences positively selected (low Tajima's D, high Fay & Wu's H, high Zeng et al's E) in their parental species of origin. P. formosa ends up with the handiest sequences from both parents.
100
Nathan Schaefer @nkschaefer.bsky.social · 12/03/2026
We asked what types of mutations gene conversion (GC) is likeliest to copy versus overwrite and found evidence that it aids purifying selection. GC likes to eliminate young mutations and those with high variant impact (likely to be deleterious).
100
Nathan Schaefer @nkschaefer.bsky.social · 24/03/2025
Thanks for reading, and good luck checking IDs and keeping the rifraff out of your single cell data sets. www.biorxiv.org/content/10.1... github.com/nkschaefer/c...
030
Nathan Schaefer @nkschaefer.bsky.social · 24/03/2025
Back-mutations to the ancestral state at this type are uncommon, at a frequency typically seen in mitochondrial protein-coding or disease-implicated mutations. This suggests that this mutation may be one of the changes affecting gene regulation at this locus.
110
Nathan Schaefer @nkschaefer.bsky.social · 24/03/2025
Mitochondrial genes are expressed as polycistronic transcripts, then cleaved and selectively degraded. We looked at species differences in this process, from two causes: nuclear and mitochondrial mutations. Interestingly, the biggest differences we found were compensatory, with little net effect.
110
Nathan Schaefer @nkschaefer.bsky.social · 24/03/2025
After demultiplexing with CellBouncer, we found that composite cells mostly inherit only one species’ mitochondria: human, for human/chimpanzee cells, and bonobo, for chimpanzee/bonobo cells. Not always, though: some cells retained both mitochondria, or those from the less common species.
110
Nathan Schaefer @nkschaefer.bsky.social · 24/03/2025
We take CellBouncer for a spin on a cool data set: inter-species composite iPSCs we created by cell fusion (www.nature.com/articles/s41...) for studying species differences in gene regulation. Here, we asked if there were biases in which species’ mitochondria were inherited by the composite cells.
130
Nathan Schaefer @nkschaefer.bsky.social · 24/03/2025
doublet_dragon takes assignments from the other programs and infers a global doublet rate that encompasses both homotypic doublets (invisible to individual programs) and heterotypic ones. This can help with QC (given expectation based on cell loading density) and serve as a prior for other tools.
110
Nathan Schaefer @nkschaefer.bsky.social · 24/03/2025
demux_tags assigns custom labels (e.g. MULTIseq/HTO data), or sgRNAs (CRISPR guide capture data) to cells. Our method considers the distribution of all tag counts together, rather than considering each tag independently, and handles noisy/low-count data better than some alternatives.
110
Nathan Schaefer @nkschaefer.bsky.social · 24/03/2025
bulkprops takes genotypes and bulk data (or single cell data, ignoring cell barcodes) and infers the proportion of each individual in the pool. This can cross-check the other programs, and we provide a method to bootstrap proportions and get p-values when comparing two sets of proportions.
110
Nathan Schaefer @nkschaefer.bsky.social · 24/03/2025
quant_contam quantifies ambient RNA by measuring how often cells mismatch their expected genotypes. This introduces an external ground truth (genotype data), avoids the need to consider empty droplets, and can find ambient RNA in data lacking cell type diversity.
120
Nathan Schaefer @nkschaefer.bsky.social · 24/03/2025
demux_mt answers this problem by simultaneously clustering mitochondrial haplotypes and inferring the number of individuals in the pool. It takes only a BAM file. There is also a way to plot the haplotypes to see how well clustering worked.
110
Nathan Schaefer @nkschaefer.bsky.social · 24/03/2025
demux_vcf assigns cells to individuals using genotypes and is fast, accurate, and robust to deep population structure. It groups SNPs by allelic state in each pair of individuals and compares the likelihood of each pair of IDs for each cell, improving speed over methods that filter or refine SNPs.
120
Nathan Schaefer @nkschaefer.bsky.social · 24/03/2025
demux_species uses an alignment-free k-mer counting strategy to save time and memory and assigns cells to species using a statistical model instead of a cutoff. Users can plot the clustered k-mer counts to see if it worked. demux_species also separates reads by species for downstream processing.
120
Nathan Schaefer @nkschaefer.bsky.social · 24/03/2025
CellBouncer provides fast, compiled, self-contained, interacting programs with methods to validate results where possible (e.g. you can visually compare two sets of IDs for the same cells, and you can visualize inferred mitochondrial haplotypes to determine how well the clustering worked).
120
Nathan Schaefer @nkschaefer.bsky.social · 24/03/2025
Introducing CellBouncer, a toolkit for pooled single cell data that assigns cells to species or individual of origin, performs genotype-free ids using mitochondrial haplotypes, assigns sgRNAs and custom tags to cells, and models ambient RNA using external genotype data as a ground truth.
171
Nathan Schaefer @nkschaefer.bsky.social · 24/03/2025
Back-mutations to the ancestral state at this type are uncommon, at a frequency typically seen in mitochondrial protein-coding or disease-implicated mutations. This supports the idea that this mutation could be one of the changes affecting gene regulation at this locus.
100
Nathan Schaefer @nkschaefer.bsky.social · 24/03/2025
Mitochondrial genes are expressed as polycistronic transcripts, then cleaved and selectively degraded. We looked at species differences in this, from two causes: nuclear and mitochondrial genome mutations. Interestingly, the biggest differences we found were compensatory, with little net effect.
100
Nathan Schaefer @nkschaefer.bsky.social · 24/03/2025
Cells with two species’ mitochondria have significantly altered gene expression related to cell cycle arrest and apoptosis relative to other cells, suggesting they’re in trouble. They also express fewer mitochondrial transcripts overall and have abnormal post-expression transcriptional regulation.
100
Nathan Schaefer @nkschaefer.bsky.social · 24/03/2025
After demultiplexing with CellBouncer, we found that composite cells mostly inherit only one species’ mitochondria: human, for human/chimpanzee cells, and bonobo, for chimpanzee/bonobo cells. Not always, though: some cells retained both mitochondria, or those from the less common species.
100
Nathan Schaefer @nkschaefer.bsky.social · 24/03/2025
We take CellBouncer for a spin on a cool data set: inter-species composite iPSCs we created by cell fusion (www.nature.com/articles/s41...) for studying species differences in gene regulation. Here, we asked if there were biases in which species’ mitochondria were inherited by the composite cells.
100
Nathan Schaefer @nkschaefer.bsky.social · 24/03/2025
doublet_dragon takes assignments from the other programs and infers a global doublet rate that encompasses both homotypic doublets (invisible to individual programs) and heterotypic ones. This can help with QC (given expectation based on cell loading density) and serve as a prior for other tools.
100
Nathan Schaefer @nkschaefer.bsky.social · 24/03/2025
demux_tags assigns custom labels (e.g. MULTIseq/HTO data), or sgRNAs (CRISPR guide capture data) to cells. Our method considers the distribution of all tag counts together, rather than considering each tag independently, and handles noisy/low-count data better than some alternatives.
100
Nathan Schaefer @nkschaefer.bsky.social · 24/03/2025
bulkprops takes genotypes and bulk data (or single cell data, ignoring cell barcodes) and infers the proportion of each individual in the pool. This can cross-check the other programs, and we provide a method to bootstrap proportions and get p-values when comparing two sets of proportions.
100
Nathan Schaefer @nkschaefer.bsky.social · 24/03/2025
quant_contam quantifies ambient RNA by measuring how often cells mismatch their expected genotypes. This introduces an external ground truth (genotype data), avoids the need to consider empty droplets, and can find ambient RNA in data lacking cell type diversity.
100
Nathan Schaefer @nkschaefer.bsky.social · 24/03/2025
demux_vcf assigns cells to individuals using genotypes and is fast, accurate, and robust to deep population structure. It groups SNPs by allelic state in each pair of individuals and compares the likelihood of each pair of IDs for each cell, improving speed over methods that filter or refine SNPs.
000
Nathan Schaefer @nkschaefer.bsky.social · 24/03/2025
demux_species uses an alignment-free k-mer counting strategy to save time and memory and assigns cells to species using a statistical model instead of a cutoff. Users can plot the clustered k-mer counts to see if it worked. demux_species also separates reads by species for downstream processing.
100
Nathan Schaefer @nkschaefer.bsky.social · 24/03/2025
CellBouncer provides fast, compiled, self-contained, interacting programs with methods to validate results where possible (e.g. you can visually compare two sets of IDs for the same cells, and you can visualize inferred mitochondrial haplotypes to determine how well the clustering worked).
100
Nathan Schaefer @nkschaefer.bsky.social · 24/03/2025
Introducing CellBouncer, a toolkit for pooled single cell data that assigns cells to species or individual of origin, performs genotype-free ids using mitochondrial haplotypes, assigns sgRNAs and custom tags to cells, and models ambient RNA using external genotype data as a ground truth.
100