Sign in

Dongwook Kim

@dongwookkim.bsky.social
202 followers 134 following 15 posts

Developing fast and easy methods for #phylogenetics and #bioinformatics | PhD in Bioinformatics | Postdoc @ Comparative Genomics Lab, UNIL/SIB🇨🇭| Formerly @ Steinegger Lab, SNU🇰🇷 | he/him

PostsRepliesMedia
Reposted by Dongwook Kim
Bioinformatics Advances @bioinfoadv.bsky.social · 18/08/2026
🌿 Just published in Bioinformatics Advances: "AmpliPhy improves gene trees by adding homologous sequences without affecting alignments"  Find it here: doi.org/10.1093/bioadv/vbag222 Authors include: @dongwookkim.bsky.social, @dessimoz.bsky.social
143
Dongwook Kim @dongwookkim.bsky.social · 13/08/2026
AI-driven protein structure prediction lets us revisit the question once impossible to tackle at scale: What can we learn from the evolution of protein structures? Our new preprint reviews how to properly assess the quality of structure-based alignments and trees. 🧵1/5 📄 doi.org/10.32942/X2T10D
1133
Dongwook Kim @dongwookkim.bsky.social · 10/08/2026
AmpliPhy is now reviewed and published in @bioinfoadv.bsky.social! AmpliPhy improves your gene trees by adding homologs with a single command, based on the observations from a quantitative benchmark design. More details in quoted 🧵 📄 doi.org/10.1093/bioadv/vbag222 💾 github.com/DessimozLab/ampliphy
doi.org
AmpliPhy improves gene trees by adding homologous sequences without affecting alignments
AbstractMotivation. In phylogenomics, gene tree reconstruction depends on multiple sequence alignment and tree inference, and ongoing work continues to imp
074
Reposted by Dongwook Kim
Jaebeom Kim @jbeom.bsky.social · 09/04/2026
Metabuli & Metabuli App v1.2 improve novel species classification with higher precision and recall. New light mode is 1.8× faster and requires 50% less storage while keeping precision. New RefSeq, GTDB, HRGM, and HROM databases added. 💾 github.com/steineggerla... 📄 doi.org/10.64898/2026.03.13.711249
13117
Reposted by Dongwook Kim
Chan Yeong Kim @chanyeong-kim.bsky.social · 09/02/2026
I am pleased to share that our paper is now published in Cell! www.cell.com/cell/fulltex... I am deeply grateful to all co-authors for making this possible. This work was made possible through the guidance of Dr. Peer Bork. I share this in grateful memory and with deep respect for his mentorship.
cell.com
Planetary microbiome structure and generalist-driven gene flow across disparate habitats
A planetary-scale analysis of over 85,000 metagenomes establishes a framework for exploring the structure and drivers of global microbial habitats, revealing that generalist species bridge ecological ...
13114
Reposted by Dongwook Kim
Martin Steinegger 🇺🇦 @martinsteinegger.bsky.social · 30/01/2026
FoldMason is out now in @science.org. It generates accurate multiple structure alignments for thousands of protein structures in seconds. Great work by Cameron L. M. Gilchrist and @milot.bsky.social. 📄 www.science.org/doi/10.1126/... 🌐 search.foldseek.com/foldmason 💾 github.com/steineggerla...
science.org
Multiple protein structure alignment at scale with FoldMason
Protein structure is conserved beyond sequence, making multiple structural alignment (MSTA) essential for analyzing distantly related proteins. Computational prediction methods have vastly extended ou...
4302147
Dongwook Kim @dongwookkim.bsky.social · 28/01/2026
Can ever-increasing sequence databases improve phylogenetic reconstruction of a gene family? Our new preprint introduces AmpliPhy, a pipeline that automates homolog enrichment to improve gene tree inference, built on a robust phylogenomic benchmark scheme. 🧵1/n 📃 doi.org/10.64898/2026.01.26.701724
doi.org
AmpliPhy improves gene trees by adding homologs without affecting alignments
In phylogenomics, gene tree reconstruction depends on multiple sequence alignment (MSA) and tree inference, and ongoing work continues to improve inference quality. Denser taxon sampling has been associated with improved gene tree inference, suggesting that adding homologs could be a practical route to higher accuracy as sequence databases continue to expand. However, adding sequences can influence multiple steps of typical inference pipelines, and little is known on its specific effect on the multiple sequence alignment, tree reconstruction, and rooting steps. We performed a large-scale empirical benchmark to quantify how homolog enrichment affects alignment and phylogenetic inference. Using an enrichment-impoverishment design and a measure of tree accuracy based on taxonomic congruence, we found that enrichment consistently improves tree inference quality, while effects on alignment quality are marginal. We show that this improvement is associated with accurate root placement on enriched trees when sensitive homolog search is accompanied. Notably, much of the benefit can be retained with relatively compact alignments produced by sequence addition. Building on these observations, we provide a tool, AmpliPhy, which efficiently improves phylogenetic reconstruction of protein families through homolog enrichment. The AmpliPhy open-source pipeline software is available at https://github.com/DessimozLab/ampliphy. ### Competing Interest Statement The authors have declared no competing interest. Swiss National Science Foundation, https://ror.org/00yjd3n13, 216623, 10005715
12615
Reposted by Dongwook Kim
Milot Mirdita @milot.bsky.social · 20/01/2026
My time in @martinsteinegger.bsky.social's group is ending, but I’m staying in Korea to build a lab at Sungkyunkwan University School of Medicine. If you or someone you know is interested in molecular machine learning and open-source bioinformatics, please reach out. I am hiring! mirdita.org
mirdita.org
Mirdita Lab - Laboratory for Computational Biology & Molecular Machine Learning
Mirdita Lab builds scalable bioinformatics methods.
710756
Reposted by Dongwook Kim
George Bouras @gbouras13.bsky.social · 08/08/2025
Stoked to finally have a preprint out for Phold, our tool that uses protein structural information to enhance phage genome annotation #phagesky 1/n www.biorxiv.org/content/10.1...
biorxiv.org
Protein Structure Informed Bacteriophage Genome Annotation with Phold
Bacteriophage (phage) genome annotation is essential for understanding their functional potential and suitability for use as therapeutic agents. Here we introduce Phold, an annotation framework utilis...
513766
Reposted by Dongwook Kim
Chan Yeong Kim @chanyeong-kim.bsky.social · 21/07/2025
Our new preprint is out! www.biorxiv.org/content/10.1... In this study, we present the largest systematic analysis of microbiome structure and function, integrating 85K uniformly processed metagenomes from diverse habitats worldwide. @podlesny.bsky.social @jonas-bio.bsky.social @borklab.bsky.social
biorxiv.org
Planetary microbiome structure and generalist-driven gene flow across disparate habitats
Microbes are ubiquitous on Earth, forming microbiomes that sustain macroscopic life and biogeochemical cycles. Microbial dispersion, driven by natural processes and human activities, interconnects mic...
12818
Reposted by Dongwook Kim
Laurie Belcher @lauriebelch.bsky.social · 16/07/2025
OrthoFinder just dropped a major update It’s faster, more accurate, and ready for thousands of genomes Let’s break it down (1/10) github.com/OrthoFinder/... www.biorxiv.org/content/10.1...
112673
Reposted by Dongwook Kim
Martin Steinegger 🇺🇦 @martinsteinegger.bsky.social · 07/07/2025
Folddisco finds similar (dis)continuous 3D motifs in large protein structure databases. Its efficient index enables fast uncharacterized active site annotation, protein conformational state analysis and PPI interface comparison. 1/9🧶🧬 📄 www.biorxiv.org/content/10.1... 🌐 search.foldseek.com/folddisco
816271
Reposted by Dongwook Kim
georghochberg.bsky.social @georghochberg.bsky.social · 11/06/2025
New paper from the lab from Sriram Garg in my group. We introduce a general substitution matrix for structural phylogenetics. I think this is a big deal, so read on below if you think deep history is important. academic.oup.com/mbe/advance-...
academic.oup.com
A general substitution matrix for structural phylogenetics.
Abstract. Sequence-based maximum likelihood (ML) phylogenetics is a widely used method for inferring evolutionary relationships, which has illuminated the
39752
Dongwook Kim @dongwookkim.bsky.social · 03/06/2025
Unicore is now published on GBE 🚀 Unicore rapidly identifies structural single-copy core genes from input species proteomes for phylogenetic analysis. Powered by Foldseek and ProstT5, Unicore enables linear-scale structure-based phylogeny of any given set of taxa. 🧵1/n 📃 doi.org/10.1093/gbe/evaf109
36831
Reposted by Dongwook Kim
Martin Steinegger 🇺🇦 @martinsteinegger.bsky.social · 27/04/2025
AFESM: a metagenomic guide through the protein structure universe! We clustered 821M structures (AFDB&ESMatlas) into 5.12M groups; revealing biome-specific groups, only 1 new fold even after AlphaFold2 re-prediction & many novel domain combos. 🧵 🌐 afesm.foldseek.com 📄 www.biorxiv.org/content/10.1...
414170
Reposted by Dongwook Kim
Martin Steinegger 🇺🇦 @martinsteinegger.bsky.social · 25/04/2025
Visit our posters at #RECOMB2025 for: Structural: MSAs, Virus DB, Core Genes, Motif Discovery, Multimer Clustering & Search, pLM Foldseek, Environmental analysis Metagenomics: Classification & Metabuli App GPU-based & RNA search, Proteome clustering, Novel Ribozyme discovery & get Marv stickers!
26419
Reposted by Dongwook Kim
Matthew Hahn @3rdreviewer.bsky.social · 10/04/2025
Not really my announcement to make--I am but a lesser co-author--but IQ-TREE 3 has just been released! (Most credit to Minh Bui and @roblanfear.bsky.social and their labs) ecoevorxiv.org/repository/v...
ecoevorxiv.org
IQ-TREE 3: Phylogenomic Inference Software using Complex Evolutionary Models
217896
Reposted by Dongwook Kim
EMBL-EBI @ebi.embl.org · 03/03/2025
🚀 #AlphaFold Database update AlphaFold DB now integrates The Encyclopedia of Domains (TED) – a resource designed to systematically identify & classify structural domains within AlphaFold-predicted protein structures. www.ebi.ac.uk/about/news/u... @pdbeurope.bsky.social
111844
Reposted by Dongwook Kim
Christophe Dessimoz @dessimoz.bsky.social · 26/02/2025
The PAN-GO paper is a remarkable milestone. It not only provides the most comprehensive picture of human gene function to date, but also carefully maps this knowledge across the tree of life! Congratulations @marcfeuermann.bsky.social, Pascale Gaudet & collaborators! www.sib.swiss/news/sib-hel...
sib.swiss
SIB helps create most complete, accurate resource for human gene functions
For the first time, biodata from humans have been integrated with that of other organisms to provide the most comprehensive picture of human gene function to date. The new ‘PAN-GO’ resource used e...
01612
Reposted by Dongwook Kim
Martin Steinegger 🇺🇦 @martinsteinegger.bsky.social · 22/02/2025
In our latest review, we explore 12 deep-learning tools for metagenomic analysis, covering their strengths, limitations, and key applications. We hope it serves as both a resource and inspiration for new ways to analyze metagenomic data. Great work by Eli Levy Karin! 📄 doi.org/10.1093/nsr/...
210644
Reposted by Dongwook Kim
Sina Majidian @sinamajidian.bsky.social · 03/01/2025
FastOMA is out now in Nature Methods 🎉: nature.com/articles/s41592-024-02552-8 A new orthology inference algorithm that scales linearly and is highly accurate. FastOMA can process all >2000 eukaryotic UniProt ref proteomes <24 hours 🚀. Try it out github.com/DessimozLab/fastoma @dessimoz.bsky.social
FastOMA retains OMA’s high precision accuracy and even improves upon it in terms of recall, positioning it on the Pareto frontier of orthology inference methods. 
FastOMA is not only fast but also accurate. a, QfO benchmar, agreement with SwissTree reference phylogeny covering manually curated gene trees. The error bars indicate 95% confidence intervals comparing FastOMA with EnsemblCompara, Domainoid, OrthoMCL, Ortholnspector, sonicparanoid, PANTHER, OrthoFinder, Hieranoid26 and the OMA family including OMA pairs, OMA groups and OMA GETHOGs (graph-based efficient technique for HOGs).

c) A computation time comparison of FastOMA and state-of-the-art alternatives.
https://www.nature.com/articles/s41592-024-02552-8
14118
Reposted by Dongwook Kim
Martin Steinegger 🇺🇦 @martinsteinegger.bsky.social · 23/12/2024
Unicore identifies single-copy protein structures across genomes using Foldseek, bypassing slow structure predictions by utilizing 3Di predictions from ProstT5, enabling rapid phylogenetic inference at the tree-of-life scale. 1/n 📄 www.biorxiv.org/content/10.1... 💾 github.com/steineggerla...
212256
Reposted by Dongwook Kim
bioRxiv Bioinfo @biorxiv-bioinfo.bsky.social · 23/12/2024
Unicore enables scalable and accurate phylogenetic reconstruction with structural core genes www.biorxiv.org/content/10.1101/202…
053
Reposted by Dongwook Kim
Bluesky @bsky.app · 03/12/2024
Scientists, academics, researchers: We’re excited to share that @altmetric.com is now tracking mentions of your research on Bluesky! 🧪
456294874995
Reposted by Dongwook Kim
Adam Schwarz @adamjschwarz.bsky.social · 03/12/2024
South Korean citizens helped lawmakers scale the National Assembly walls so they could bypass military barricades and vote against martial law.
81135103125
Reposted by Dongwook Kim
Richard Sever @richardsever.bsky.social · 10/11/2024
Reminder for newcomers that bioRxiv has Bluesky accounts in every subject category - great way to keep up (please re-skeet) connect.biorxiv.org/news/2023/09...
connect.biorxiv.org
bioRxiv expands on Mastodon and Bluesky
bioRxiv - the preprint server for biology, operated by Cold Spring Harbor Laboratory, a research and educational institution
6361286
Reposted by Dongwook Kim
Martin Steinegger 🇺🇦 @martinsteinegger.bsky.social · 28/11/2024
Interested in bioinformatics method development for proteins, structures or metagenomic analysis? Please check out my lab’s starter pack! 🔗 go.bsky.app/VJhXcSs
35611
Reposted by Dongwook Kim
Milot Mirdita @milot.bsky.social · 27/11/2024
MMseqs2 Release 16 Highlights: GPU-accelerated search📄, ORF or new 6-frame translated search modes, contig taxonomy always keeps the longest ORF, bug fixes (reduced memory and higher sensitivity) and relicensed as MIT 📄 biorxiv.org/content/10.1... 💾 mmseqs.com and 🐍Bioconda 🖥️🧬🧶
Promotional logo for MMseqs2 16 with the MMseqs2 Rocket mascot as a smart phone like App logo
011443
Reposted by Dongwook Kim
Roli Roberts @roliroberts.bsky.social · 25/11/2024
What did the Last Eukaryotic Common Ancestor (#LECA) look like? Consensus View in #PLOSBiology; massive authorship including @AncestralState, @lauraeme.bsky.social, John Archbald, @andrewjroger.bsky.social, @dackslabecb.bsky.social, Jeremy Wideman. plos.io/4g0alq4
6249101
Reposted by Dongwook Kim
Martin Steinegger 🇺🇦 @martinsteinegger.bsky.social · 23/11/2024
Our Big Fantastic Virus Database (BFVD) is now published NAR! It contains protein structure predictions of major viral clades, enhanced by petabase-scale homology search and it's explorable on the web. 🌐 bfvd.foldseek.com 💾 bfvd.steineggerlab.workers.dev 📄 academic.oup.com/nar/advance-...
6339126