Sign in

Vanja

@halfacrocodile.bsky.social
115 followers 169 following 54 posts

bald but bearded bioinformatics buddy👨‍🦲 autosome.org

PostsRepliesMedia
Reposted by Vanja
Waggoner Lab @labwaggoner.bsky.social · 05/08/2026
An expanded codebook of human transcription factor DNA-binding specificity @nature.com @bartdeplancke.bsky.social @halfacrocodile.bsky.social @arttujolma.bsky.social @vorontsovie.bsky.social @downbythebayes.bsky.social @uoftpress.bsky.social #Hughes www.nature.com/articles/s41...
331
Vanja @halfacrocodile.bsky.social · 11/04/2026
10/10 The FANTOMUS atlas is available at fantomus.autosome.org: an interactive browser of skeletal muscles’ gene expression, protein abundance, TRE expression, and motif activity. Read more: www.biorxiv.org/content/10.6...
010
Vanja @halfacrocodile.bsky.social · 11/04/2026
9/10 Finally, allele-specific analysis revealed 6653 single-nucleotide variants with a significant allelic imbalance, which were enriched by GTEx eQTLs, ADASTRA ASBs, and associated with total thigh muscle volume.
110
Vanja @halfacrocodile.bsky.social · 11/04/2026
8/10 To identify the key regulators driving the observed muscle heterogeneity, we performed motif activity response analysis. 353 motifs have non-uniform activity across muscles, including MEF2, SIX, and HSF proteins, and eye-specific ZBTB7A.
100
Vanja @halfacrocodile.bsky.social · 11/04/2026
7/10 Besides gene expression, CAGE-Seq revealed alternative promoter usage: for 1648 TREs of 1264 genes, including TPM3 and PITX2, the contribution to the total gene expression was significantly varying across muscles.
100
Vanja @halfacrocodile.bsky.social · 11/04/2026
6/10 Non-uniformly expressed genes include sarcomere components, HOX transcription factors, and atrophy- and dystrophy-associated genes, e.g., FKRP, reflecting the differences in the tissue etiology, fiber contraction, and varying susceptibility to myopathies.
110
Vanja @halfacrocodile.bsky.social · 11/04/2026
5/10 Surprisingly, more than 80% of genes and proteins demonstrated significant differential expression across muscles. Extraocular muscles, tongue, and diaphragm have the most uncommon activity patterns of promoters and enhancers.
100
Vanja @halfacrocodile.bsky.social · 11/04/2026
4/10 Further, with mass spectrometry, we assessed the abundance of 1804 protein groups encompassing 1895 proteins, and estimated the concordance of transcriptomic and proteomic readouts.
100
Vanja @halfacrocodile.bsky.social · 11/04/2026
3/10 Inspired by FANTOM5 (fantom.gsc.riken.jp/5/), FANTOMUS took advantage of CAGE-Seq to identify and measure the activity of 37001 TREs, including promoters of 18329 genes.
100
Vanja @halfacrocodile.bsky.social · 11/04/2026
2/10 FANTOMUS delineates activity of transcribed regulatory elements (TREs = promoters + transcribed enhancers) and protein abundance profiles in 75 (CAGE-Seq) and 22 (LC-MS/MS) distinct human skeletal muscles.
100
Vanja @halfacrocodile.bsky.social · 11/04/2026
1/10 Did you know that molecular phenotypes of human skeletal muscles are vastly different? Meet FANTOMUS, the skeletal muscles promoterome-proteome atlas: doi.org/10.64898/202... made with @andreybuyan.bsky.social @nikitagryzunov.bsky.social @alforrest.bsky.social @sevamakeev.bsky.social
165
Reposted by Vanja
Dmitry Penzar @pensarata.bsky.social · 18/11/2025
(1/13) Excited to share the outcome of the IBIS Challenge! The IBIS challenge united dozens of teams across the world in tackling the problem of modeling transcription factor (TF) binding specificity using a diverse collection of experimental datasets for understudied human TFs.
1117
Reposted by Vanja
vorontsovie.bsky.social @vorontsovie.bsky.social · 17/11/2025
Our paper on LARGE-scale benchmarking of motif discovery tools is published! nature.com/articles/s42... It was a long, 7 years long journey, which coordinated efforts of 50+ researchers, proud to be on of them. More results from Codebook about poorly studied TFs are coming soon.
nature.com
Cross-platform motif discovery and benchmarking to explore binding specificities of poorly studied human transcription factors - Communications Biology
Cross-platform benchmarking of DNA binding specificity models highlights top-performing motif discovery methods and demonstrates the potential of advanced models to capture alternative binding modes o...
066
Vanja @halfacrocodile.bsky.social · 19/02/2025
(15/15) MIXALIME is written in Python, getting the stable version is as easy as 'pip install mixalime', and the source code plus tutorial are freely available at github.com/autosome-ru/...
020
Vanja @halfacrocodile.bsky.social · 19/02/2025
(14/15) We hope you will find MIXALIME and UDACHA useful in your study of regulatory sequence variants.
110
Vanja @halfacrocodile.bsky.social · 19/02/2025
(13/15) According to stratified LD score regression, these rSNPs correspond to regulatory regions involved in cell type-specific phenotypes. Most importantly, the collection of significant rSNPs can be fully explored at udacha.autosome.org
120
Vanja @halfacrocodile.bsky.social · 19/02/2025
(12/15) AS-chromatin variants are predominantly located in promoter and enhancer regions and significantly overlap ADASTRA ASBs and GTEx eQTLs.
110
Vanja @halfacrocodile.bsky.social · 19/02/2025
(11/15) Finally, we used MIXALIME to analyze 5858 chromatin accessibility datasets from gtrd.biouml.org. In the end, we identified >200 thousand allele-specific chromatin accessibility variants.
120
Vanja @halfacrocodile.bsky.social · 19/02/2025
(10/15) In most cases, MIXALIME outperforms other popular tools for AS calling, offering a good sensitivity/specificity tradeoff.
110
Vanja @halfacrocodile.bsky.social · 19/02/2025
(9/15) Using heart CAGE-Seq data from Deviatiiarov et al., we benchmarked MIXALIME by calling AS variants in CAGE-Seq and comparing them to allele-specific transcription factor binding from ADASTRA and eQTLs from GTEx.
110
Vanja @halfacrocodile.bsky.social · 19/02/2025
(8/15) Copy-number variation and aneuploidy are accounted for by fitting a mixture model assuming that reads originate from haplotypes with different copy numbers.
110
Vanja @halfacrocodile.bsky.social · 19/02/2025
(7/15) MIXALIME also handles reference mapping bias and aneuploidy, see the underlying math in arxiv.org/abs/2306.08287. To counter the mapping bias, MIXALIME uses separate fits for Alt read counts with the fixed number of Ref reads and vice versa.
110
Vanja @halfacrocodile.bsky.social · 19/02/2025
(6/15) MIXALIME provides a variety of statistical models to fulfill particular use cases, from a standard binomial model to the beta negative binomial (BetaNB) model that accounts for extra overdispersion.
110
Vanja @halfacrocodile.bsky.social · 19/02/2025
(5/15) MIXALIME is a novel toolbox that uses MIXture models for ALlelic IMbalance Estimation. In the paper, we describe a general workflow from FASTQ files to allelic read counts and SNV-level allele-specific statistics.
110
Vanja @halfacrocodile.bsky.social · 19/02/2025
(4/15) Technically, the allele specificity is revealed by counting the number of reads supporting each of the alleles and estimating the statistical significance of the observed allelic imbalance.
110
Vanja @halfacrocodile.bsky.social · 19/02/2025
(3/15) High-throughput sequencing allows tracking chromatin state, gene expression, protein-DNA interactions, and more. Eventually, all methods yield short reads that can be used to call single-nucleotide variants and assess the allele specificity.
110
Vanja @halfacrocodile.bsky.social · 19/02/2025
(2/15) Check the original Tweetorial describing the respective preprint x.com/halfacrocodi... or follow the thread below.
110
Vanja @halfacrocodile.bsky.social · 19/02/2025
(1/15) Yet another sweet bioinformatics "software+database" couple from our team: Meet MIXALIME, a framework for assessing allelic imbalance, and UDACHA, a database of allele-specific chromatin accessibility, read more at www.nature.com/articles/s41...
191
Vanja @halfacrocodile.bsky.social · 17/02/2025
Last but not least: this update became possible thanks to the experimental data & motif analysis performed within the Codebook/GRECO-BIT collaboration, ibis.autosome.org/docs/about_us
ibis.autosome.org
IBIS Challenge
011
Vanja @halfacrocodile.bsky.social · 17/02/2025
(4/4) Don't hesitate to grab a fresh release from hocomoco.autosome.org and remember that we also provide a fancy online motif scanner, MoLoTool, in all its interactive JS-powered beauty.
101
Vanja @halfacrocodile.bsky.social · 17/02/2025
(3/4) Also do not forget that HOCOMOCO also provides motifs for orthologous mouse TFs, >800 of those are covered in v13.
101
Vanja @halfacrocodile.bsky.social · 17/02/2025
(2/4) v13 covers >1100 of ~1600 human TFs with >1600 primary motifs and subtypes. Since v12 we also provide a reduced non-redundant set of motifs, which are often shared between TFs with similar DBDs.
101
Vanja @halfacrocodile.bsky.social · 17/02/2025
(1/4) Thrilled to announce another major release of the HOCOMOCO motif collection, well-known for its silly name and rigorous approach to constructing and benchmarking DNA sequence motifs recognized by human and mouse transcription factors. hocomoco.autosome.org
23113
Reposted by Vanja
Hani Goodarzi @genophoria.bsky.social · 02/01/2025
The first preprint of 2025! Together with Matvei, @halfacrocodile.bsky.social, & our amazing team, we are excited to share PARADE: an AI framework for designing mRNA UTRs with enhanced cell-type specificity & stability. www.biorxiv.org/content/10.1...
biorxiv.org
A generative framework for enhanced cell-type specificity in rationally designed mRNAs
mRNA delivery offers new opportunities for disease treatment by directing cells to produce therapeutic proteins. However, designing highly stable mRNAs with programmable cell type-specificity remains ...
18138
Vanja @halfacrocodile.bsky.social · 28/11/2024
(6/6) 👏 Once again, congratulations to the winners and their mentors including @callitmagic.bsky.social y.social @yaronorenstein.bsky.social @salimovdr.bsky.social @german-roev.bsky.social
022
Vanja @halfacrocodile.bsky.social · 28/11/2024
(5/6) 🏁 This is the finishing line of the challenge but more remains to be done to analyze the solutions in detail and perform extra tests against other models for the post-challenge manuscript.
110
Vanja @halfacrocodile.bsky.social · 28/11/2024
(4/6) 🗣️ Check the IBIS website (ibis.autosome.org/challenge_te...) to see the detailed rankings and read the short interviews with the winners!
110
Vanja @halfacrocodile.bsky.social · 28/11/2024
(3/6) 🏆 We congratulate the runner-up teams: Medici, Salimov & Frolov lab (Russia), callitmagic (Armenia); and, most importantly, the winners: teams Bench Pressers (US), mj (Russia), and Biology Impostor (Israel).
110
Vanja @halfacrocodile.bsky.social · 28/11/2024
(2/6) 😮 Since the launch, more than 100 teams registered to participate, and 19 teams (with members from 12 countries across the globe) survived till the final stage. The complete IBIS data on Zenodo: zenodo.org/records/1418... and zenodo.org/records/1417...
110
Vanja @halfacrocodile.bsky.social · 28/11/2024
(1/6) 🐦‍🔥 In IBIS #ibischallenge, we challenged teams from all over the world to decipher the DNA recognition code of human transcription factors. The IBIS Final Conference took place on November 27, 2024. Recordings and slides: disk.yandex.ru/d/82FEnwPn15...
1105
Vanja @halfacrocodile.bsky.social · 26/11/2024
A brief reminder: the IBIS conference is going to happen on Nov 27 = today (or tomorrow, depending on your timezone). Grab the Zoom link at ibis.autosome.org/news and enjoy the show!
ibis.autosome.org
IBIS Challenge
000
Reposted by Vanja
timhughesto.bsky.social @timhughesto.bsky.social · 15/11/2024
(1/7) Thrilled to reveal the results of the Codebook Project, an international effort to identify accurate DNA-binding motifs and genomic binding loci for the >300 uncharacterized human transcription factors (TFs) doi.org/10.1101/2024...
doi.org
Perspectives on Codebook: sequence specificity of uncharacterized human transcription factors
We describe an effort (“Codebook”) to determine the sequence specificity of 332 putative and largely uncharacterized human transcription factors (TFs), as well as 61 control TFs. Nearly 5,000 independ...
14111
Reposted by Vanja
Isaac Yellan @downbythebayes.bsky.social · 15/11/2024
(1/8) 🚀 Excited to share our findings on a large-scale ChIP-seq assay for 166 previously uncharacterized human transcription factors (TFs) and their roles in both regulatory regions, and more strikingly, the “dark matter” genome. 🌌 doi.org/10.1101/2024....
doi.org
Extensive binding of uncharacterized human transcription factors to genomic dark matter
Most of the human genome is thought to be non-functional, and includes large segments often referred to as “dark matter” DNA. The genome also encodes hundreds of putative and poorly characterized tran...
39439
Vanja @halfacrocodile.bsky.social · 14/11/2024
(12/12) Kudos to the Codebook and GRECO-BIT consortium members including @bartdeplancke.bsky.social, see the whole team at ibis.autosome.org/docs/about_us
ibis.autosome.org
IBIS Challenge
023
Vanja @halfacrocodile.bsky.social · 14/11/2024
(11/12) This project is only one of the multiple facets of the larger Codebook initiative, check dx.doi.org/10.1101/2024...
112
Vanja @halfacrocodile.bsky.social · 14/11/2024
(10/12) Want something more than simple PWMs? Reaching the next level with IBIS: join the online IBIS conference on November 27, more details soon at ibis.autosome.org.
111
Vanja @halfacrocodile.bsky.social · 14/11/2024
(9/12) Check the interactive Codebook/GRECO-BIT Motif Explorer (mex.autosome.org), which provides motifs, performance metrics, ranks, logos, top-performing motifs, and structured metadata.
mex.autosome.org
GRECO-BIT/Codebook Motif Explorer
101
Vanja @halfacrocodile.bsky.social · 14/11/2024
(8/12) With that many PWMs at hand, we also demonstrate that several motifs can be easily combined into a better model with ArChIPelago: regression or Random Forest on top of PWM scans.
111
Vanja @halfacrocodile.bsky.social · 14/11/2024
(7/12) Comparison of different tools yielded many surprises. On the one hand, underused motif discovery tools such as Dimont excel across different types of experimental data. On the other hand, no single tool or platform is enough to get the most for each TF and each data type.
112
Vanja @halfacrocodile.bsky.social · 14/11/2024
(6/12) In total, we processed data from 4,237 experiments and generated 219,939 motifs. 164,350 motifs (for 236 TFs) from curated and approved datasets were taken for benchmarking. The motifs and datasets are available at ZENODO zenodo.org/records/1018....
zenodo.org
Codebook Motif Explorer Supplementary Dataset
This is the supplementary dataset of the GRECO-BIT/Codebook Motif Explorer (MEX),which is a detailed online catalog of DNA motifs built using a huge collection of experimental data produced by the Cod...
101