Sign in

André Boler Barros, PhD

@asbarros.bsky.social
316 followers 157 following 108 posts

Data-driven individual, trying to get by in the uncertainty of statistics and life. Avid Rstats follower. Data Analyst at GIMM Institute | PhD in Molecular Biosciences

PostsRepliesMedia
André Boler Barros, PhD @asbarros.bsky.social · 13/07/2026
"Our results suggest that Schrödinger’s causal inference” (33)—where studies avoid stating (or even explicitly deny) an interest in estimating causal effects yet are otherwise embedded with causal intent (...) is common in the observational health literature." pmc.ncbi.nlm.nih.gov/articles/PMC...
pmc.ncbi.nlm.nih.gov
Checking your browser - reCAPTCHA
010
André Boler Barros, PhD @asbarros.bsky.social · 17/02/2026
When searching for power analysis (or just want to check the ideal number of cells) for scRNASeq studies, here are some resources that may be relevant: #bioinformatics
110
André Boler Barros, PhD @asbarros.bsky.social · 29/12/2025
"[commit 1 percent of profits for workers re-training] isn’t charity. (...) Helping retrain workers is common sense, and such a small ask that these companies would barely feel it, while the public benefits could be enormous." www.nytimes.com/2025/12/27/o...
nytimes.com
Opinion | A 1 Percent Solution to the Looming A.I. Job Apocalypse
120
André Boler Barros, PhD @asbarros.bsky.social · 26/11/2025
Today, I inaugurate a new Github repository: github.com/andrebolerba... I created this to store code that, although short and simple, may be useful for addressing specific needs and challenges.
github.com
GitHub - andrebolerbarros/BitsandPieces: This repository stores codes & notebooks done in response to small & simple challenges
This repository stores codes & notebooks done in response to small & simple challenges - andrebolerbarros/BitsandPieces
100
André Boler Barros, PhD @asbarros.bsky.social · 18/11/2025
I am very proud to announce that yesterday I have successfully defended my PhD. It was a long journey, filled with difficulties but with a dose of rewards - scientific, professional but most of all, personal. But, it definitely made me a better scientist, and a better person.
131
André Boler Barros, PhD @asbarros.bsky.social · 28/10/2025
"(...) the rise of generative AI in bioinformatics has not diminished my role, but redefined it. It has challenged me to become a better scientist. For good or ill, AI seems to be here to stay. I urge you to embrace the technology — not to replace your expertise, but to amplify it. #bioinfo
nature.com
‘Am I redundant?’: how AI changed my career in bioinformatics
A run-in with some artefact-laden AI-generated analyses convinced Lei Zhu that machine learning wasn’t making his role irrelevant, but more important than ever.
000
André Boler Barros, PhD @asbarros.bsky.social · 07/10/2025
Can't advise this lab more. If you'd like to work in a curious-driven and nurturing environment, with a high focus on robust data analysis, don't even think twice! #bioinfo #datascience
020
André Boler Barros, PhD @asbarros.bsky.social · 30/09/2025
Last week, I was fortunate enough to watch a talk from @tkorem.bsky.social , where he presented different things, from addressing inter- study variability on microbiome projects to the use of novel approaches on metagenomics alignment and processing. Interesting and very relevant! #bioinfo
030
André Boler Barros, PhD @asbarros.bsky.social · 29/09/2025
"GlucoStats demonstrates high efficiency in processing large-scale medical datasets in minimal time. Its modular design enables easy customization and extension, making it adaptable to diverse research and clinical needs" bmcbioinformatics.biomedcentral.com/articles/10.... #datascience #biostats
bmcbioinformatics.biomedcentral.com
Glucostats: an efficient Python library for glucose time series feature extraction and visual analysis - BMC Bioinformatics
Background The advancement of technology and continuous glucose monitoring (CGM) systems has introduced several computational and technical challenges for clinicians and researchers. The growing volume of CGM data necessitates the development of efficient computational tools capable of handling and processing this information effectively. This paper introduces GlucoStats, an open-source and multi-processing Python library designed for efficient computation and visualization of a comprehensive set of glucose metrics derived from CGM. It simplifies the traditionally time-consuming and error-prone process of manual CGM metrics calculation, making it a valuable tool for both clinical and research applications. Results Its modular design ensures easy integration into predefined workflows, while its user-friendly interface and extensive documentation make it accessible to a broad audience, including clinicians and researchers. GlucoStats offers several key features: (i) window-based time series analysis, enabling time series division into smaller ‘windows’ for detailed temporal analysis, particularly beneficial for CGM data; (ii) advanced visualization tools, providing intuitive, high-quality visualizations that facilitate pattern recognition, trend analysis, and anomaly detection in CGM data; (iii) parallelization, leveraging parallel computing to efficiently handle large CGM datasets by distributing computations across multiple processors; and (iv) scikit-learn compatibility, adhering to the standardized interface of scikit-learn to allow an easy integration into machine learning pipelines for end-to-end analysis. Conclusions GlucoStats demonstrates high efficiency in processing large-scale medical datasets in minimal time. Its modular design enables easy customization and extension, making it adaptable to diverse research and clinical needs. By offering precise CGM data analysis and user-friendly visualization tools, it serves both technical researchers and non-technical users, such as physicians and patients, with practical and research-driven applications.
010
André Boler Barros, PhD @asbarros.bsky.social · 22/09/2025
"Delphi-2M predicts the rates of more than 1,000 diseases (...), with accuracy comparable to that of existing single-disease models. Delphi-2M (...) also enables sampling of synthetic future health trajectories" www.nature.com/articles/s41...
nature.com
Learning the natural history of human disease with generative transformers - Nature
Delphi-2M forecasts a person’s future health, covering more than 1,000 diseases, provides insights into co-morbidity dynamics and generates synthetic data for the training of AI models that have never...
010
André Boler Barros, PhD @asbarros.bsky.social · 11/09/2025
Amazing repository with several references and resources for scRNASeq analysis github.com/crazyhottomm... #bioinfo #singlecell
000
André Boler Barros, PhD @asbarros.bsky.social · 11/09/2025
"Even if work is done by or with the help of experts (...), it is crucial that researchers understand how a method works, that they can assess data quality, and that they fundamentally understand what types of conclusions can and cannot be drawn from their data" www.nature.com/articles/s41...
nature.com
Push-button science - Nature Methods
Technological advances change not only what we can learn as scientists, but also how science is conducted. Here we explore how automation and outsourcing are affecting the act of doing science.
000
André Boler Barros, PhD @asbarros.bsky.social · 21/08/2025
Spreadsheets represent an everyday tool for most wet-lab scientists. So, why not use them at their highest potential, efficiently and ready for open science? This paper provides some recommendations for the use of spreadsheets: www.nature.com/articles/d41... #bioinfo #stats
nature.com
Six questions to ask before jumping into a spreadsheet
Spreadsheet software can be frustrating, but adopting some helpful habits can improve its effectiveness.
020
Reposted by André Boler Barros, PhD
The Transmitter @thetransmitter.bsky.social · 13/08/2025
FlyBase, a Drosophila database, will lose a third of its team in early October because the Harvard grant that covered the employees’ salaries was canceled. Scientists warn that losing FlyBase could devastate fly research. By @claudia-lopez.bsky.social www.thetransmitter.org/community/ha...
thetransmitter.org
Harvard University lays off fly database team
The layoffs jeopardize this resource, which has served more than 4,000 labs for about three decades.
3123129
André Boler Barros, PhD @asbarros.bsky.social · 12/08/2025
1/n Because, in bioinformatics, sharing is caring, let me share something I have recently started exploring - graph mapping and pangenome graphs. A pangenome graph encodes a reference genome built from many genomes in one structure, thus trying to encapsulate the known genetic variability.
110
André Boler Barros, PhD @asbarros.bsky.social · 26/06/2025
1/n Brief guide to statistical analysis of grouped data in preclinical research www.nature.com/articles/s42... In preclinical studies, clustering and nesting (C&N) scenarios, such as group-housed animals or cells on a single plate, are frequently found. This has important statistical implications
nature.com
A brief guide to statistical analysis of grouped data in preclinical research - Nature Metabolism
Clustering and nesting (C&N) arise in many preclinical studies, such as when animals are group-housed or share litters, or in cell culture. Ignoring C&N undermines the validity of analyses. He...
120
Reposted by André Boler Barros, PhD
Wolfgang Huber @wkhuber.bsky.social · 24/06/2025
Instead of painstakingly dissecting a set of primary data to find novel patterns, it can be more effective to fit an unsupervised cluster-factor-latent-spaces model and then painstakingly dissect the model parameters to find patterns imposed by the model inference.
1281
André Boler Barros, PhD @asbarros.bsky.social · 17/06/2025
Very interesting paper from Soares lab @gimmfoundation.bsky.social
010
André Boler Barros, PhD @asbarros.bsky.social · 28/05/2025
Reposting this a single time feels too short to show how much I agree with this
010
André Boler Barros, PhD @asbarros.bsky.social · 27/05/2025
gimm.pt/jobs/researc... Great job opportunity for any bioinformatician out there! @gimmfoundation.bsky.social #bioinfo
gimm.pt
Research Assistant – GIMMResearch Assistant – GIMMsearch
Thank you very much for your generous donation and for sharing your testimony.We hope to more people feel inspired to join this mission of discovery and innovation!
111
André Boler Barros, PhD @asbarros.bsky.social · 11/04/2025
Some thoughts about rarefaction adaorg.github.io/datamisfits/...
adaorg.github.io
Rarefaction - Is it a good idea? – The Data Misfits
Rarefaction in microbiome studies - Is it a good idea or just a bad practice? What are the alternatives?
110
André Boler Barros, PhD @asbarros.bsky.social · 09/04/2025
'One possible strategy is to organize consortium-led initiatives that integrate and curate reference datasets, establish criteria for dataset transparency, and mandate standardized preprocessing pipelines to reduce the variability in how competing models are evaluated' www.nature.com/articles/s41...
nature.com
A benchmarking crisis in biomedical machine learning - Nature Medicine
A lack of standardized benchmarks is hindering progress and patient benefits
000
Reposted by André Boler Barros, PhD
Marc Veldhoen @marcveld.bsky.social · 31/03/2025
Quando se corta na ciência When science is cut Instamos os políticos a reconhecerem que a ciência não é um interruptor que pode ser ligado e desligado sem consequências profundas e duradouras. www.publico.pt/2025/... 1/2
publico.pt
Quando se corta na ciência
Instamos os políticos a reconhecerem que a ciência não é um interruptor que pode ser ligado e desligado sem consequências profundas e duradouras.
154
Reposted by André Boler Barros, PhD
Darren Dahly @statsepi.bsky.social · 21/01/2025
Taking the opportunity to repost this since people hate it BECAUSE I'M RIGHT.
1110123
Reposted by André Boler Barros, PhD
GIMM Institute @gimminstitute.bsky.social · 16/01/2025
GIMM is on Bluesky! We're a recent research foundation created by the merger of 2 leading research institutes in Portugal, IGC and iMM, dedicated to answering fundamental questions of biology and human health & developing solutions to improve health and promote local and global equity. Visit gimm.pt
0147
Reposted by André Boler Barros, PhD
André Boler Barros, PhD @asbarros.bsky.social · 10/12/2024
Bluesky is indeed a social network on the rise. Another project with amazing potential is the new Portuguese institute - Gulbenkian Institute of Molecular Medicine (GIMM) - gimm.pt. I have created a Starter's pack for the GIMM Team, to link with each other and with the world go.bsky.app/RACwsRN
101
André Boler Barros, PhD @asbarros.bsky.social · 17/01/2025
Glad to see my institute @gimmfoundation.bsky.social on bluesky! Check up the account, and definitely follow them! Good things will be popping up frequently
030
Reposted by André Boler Barros, PhD
Yves Clément @yvesclement.bsky.social · 12/12/2024
Hi folks! I'm looking for a good and easy to use pipeline to annotate genomes from RNA-seq & protein data. Any suggestions? (I've already tried BRAKER3, I was wondering if people were using other tools out there...) 🧪
233
André Boler Barros, PhD @asbarros.bsky.social · 10/12/2024
Bluesky is indeed a social network on the rise. Another project with amazing potential is the new Portuguese institute - Gulbenkian Institute of Molecular Medicine (GIMM) - gimm.pt. I have created a Starter's pack for the GIMM Team, to link with each other and with the world go.bsky.app/RACwsRN
101
Reposted by André Boler Barros, PhD
Richard McElreath 🐈‍⬛ @rmcelreath.bsky.social · 02/12/2024
Working on a case study for survival analysis. These models are odd compared to typical GLM examples: missing data (censoring) that cannot be ignored, data model (likelihood) and data-generating model not same, how to visualize predictions, 2+ equivalent ways to program. Good real data wrestling.
Figure 21.2: Posterior predictive distributions of waiting times for the first adoption model (without
censoring). Each curve is a Kaplan-Meier plot for an individual prior simulation. Black curves
correspond to black cats. Orange curves correspond to all other cat colors. Left: Simulated curves
for samples of 1000 cats. Right: Simulated curves for samples of only 100 cats
1717
André Boler Barros, PhD @asbarros.bsky.social · 28/11/2024
In the advanced data analysis unit @GIMM, we have discussed sharing our experiences in checking, filtering and analysing data (omic and otherwise). So we decided to create a blog! adaorg.github.io/datamisfits/
adaorg.github.io
The Data Misfits
161
Reposted by André Boler Barros, PhD
Marcel Ribeiro-Dantas, Ph.D. @mribeirodantas.bsky.social · 25/11/2024
1/🧵Calling all container users! Whether you're a bioinformatician, developer, or a curious enthusiast, if you’re working with containers you need to check this out 🚀 Let me introduce you to Seqera Containers—a game-changer for building & hosting containers for free. 👇 buff.ly/3Oo8UG9
seqera.io
Containers | Seqera
Fetch Docker & Singularity containers with any combination of Conda / PyPI packages, for free.
262
Reposted by André Boler Barros, PhD
Seqera @seqera.io · 21/11/2024
💻Webinar: A new era for @nextflow.io! Join us on Dec 10 to explore the latest major version of the Nextflow VS Code extension, now powered by a new language server that elevates Nextflow into a first-class language in its own right! 👉 Register now: hubs.la/Q02Z0R0q0
1178
André Boler Barros, PhD @asbarros.bsky.social · 20/11/2024
p-value and null hypothesis testing: Where does it come from, where are we and where should we go from here: 1/n Despite being used throughout the 1700's, the p-value concept shows up in the 1900's by Karl Pearson and his Pearson's chi-square. But, the big boost came from Fisher
141
Reposted by André Boler Barros, PhD
Stephen Turner @stephenturner.us · 18/11/2024
A gentle introduction to pangenomics academic.oup.com/bib/article/25/6/b…
19626
André Boler Barros, PhD @asbarros.bsky.social · 19/11/2024
Forest plots as an interesting tool to reporting results instead of just p-values. towardsdatascience.com/unhappy-with... #stats
towardsdatascience.com
Unhappy with statistical significance (p-value)? Here is a simple solution
Why and how to use forest plots efficiently?
050
Reposted by André Boler Barros, PhD
Maarten van Smeden @maartenvsmeden.bsky.social · 18/11/2024
I made one for stats papers
A joke about statistics papers
15542148
André Boler Barros, PhD @asbarros.bsky.social · 18/11/2024
1/n Good alternatives for the p-values as major deciders of hypothesis validity are needed, either from a conceptual or from a methodological point of view. For the latter, I have been a great fan of this paper: www.nature.com/articles/s41...
nature.com
Use of the p-values as a size-dependent function to address practical differences when analyzing large datasets - Scientific Reports
Scientific Reports - Use of the p-values as a size-dependent function to address practical differences when analyzing large datasets
131
André Boler Barros, PhD @asbarros.bsky.social · 18/11/2024
"The ultimate goal of these approaches is to capture the complex interactions, dynamics and influences of microorganisms that can shed light into the mechanisms underpinning health and disease." academic.oup.com/bib/article/...
academic.oup.com
Statistical challenges in longitudinal microbiome data analysis
Abstract. The microbiome is a complex and dynamic community of microorganisms that co-exist interdependently within an ecosystem, and interact with its hos
000
André Boler Barros, PhD @asbarros.bsky.social · 17/11/2024
R lesson of the day (as most lessons, found completely by chance). When using ordered function in R instead of factor to created ordered factors, it will **definitely** have an impact in subsequent modelling you perform! 1/n
100
André Boler Barros, PhD @asbarros.bsky.social · 17/11/2024
www.nature.com/articles/s41... "A rinse-and-repeat iterative approach whereby experimentalists and modelers work closely together and learn each other’s language12 should help continue to shift focus from seeking correlations to exploring causation."
nature.com
Decoding immune kinetics: unveiling secrets using custom-built mathematical models - Nature Methods
Custom-built mathematical models make the immune response more predictable and offer mechanistic insights into fundamental immunology.
100
André Boler Barros, PhD @asbarros.bsky.social · 14/11/2024
Why you should "rehearse" your analysis before collecting the data towardsdatascience.com/why-hypothes... #Stats
towardsdatascience.com
Why Hypothesis Testing Should Take a Cue from Hamlet
To simulate or not to simulate, that is the question
011
André Boler Barros, PhD @asbarros.bsky.social · 14/11/2024
As a part of the Advanced Data Analysis Unit of the GIMM institute, I have worked with some transcriptomic datasets. Considering I have faced different data types, scenarios and analysis requirements, I have decided to create a github repository for RNASeq analysis. github.com/andrebolerba...
github.com
GitHub - andrebolerbarros/RNASeq: Repository with functions/ideas/considerations about RNASeq
Repository with functions/ideas/considerations about RNASeq - GitHub - andrebolerbarros/RNASeq: Repository with functions/ideas/considerations about RNASeq
120
André Boler Barros, PhD @asbarros.bsky.social · 14/11/2024
Have you ever faced a problem when working with dates in R, specifically when Excel decides your date is a number? Worry no more! We at the Advanced Data Analysis Unit in the GIMM Institute have developed a function to work around this problem: github.com/andrebolerba...
github.com
RFuncs/numeric_to_dates.R at main · andrebolerbarros/RFuncs
This repository is composed by R functions adapted/developed by me in order to analyze some datasets I look at on a daily basis - andrebolerbarros/RFuncs
110
André Boler Barros, PhD @asbarros.bsky.social · 23/10/2024
(Re)Introducing myself: I'm a data enthusiast within the realm of biology, play around with #bioinformatics in addition to #rstats! Currently working my way into #python!!
000
André Boler Barros, PhD @asbarros.bsky.social · 09/02/2024
A simple, yet super relevant explanation on some interesting statistical concepts towardsdatascience.com/a-visual-exp...
towardsdatascience.com
A Visual Explanation of Variance, Covariance, Correlation and Causation
Improve your data analysis skills by understanding basic statistical concepts
010
André Boler Barros, PhD @asbarros.bsky.social · 09/02/2024
"Failure to consider sex and gender in research and clinical care contributes to health inequities and, in particular, poor understanding of women’s health" www.nature.com/articles/s41...
nature.com
Consideration of sex differences is necessary to achieve health equity - Nature Reviews Nephrology
Improved understanding of the impact of sex and gender-related factors on human health and disease and the inclusion of people of all genders in research studies is necessary to reduce health inequities and enable a more personalized approach to patient care.
000
André Boler Barros, PhD @asbarros.bsky.social · 09/02/2024
On the importance of statistics - 1/n From a practical perspective, statistics encompasses experimental design, visualization, inference and prediction. Under a more philosophical lense, though, it presents the world as we perceive it, and quantifies uncertainty of said view.
100
André Boler Barros, PhD @asbarros.bsky.social · 09/02/2024
For anyone performing statistical analysis, in ecology or in any other field, this is incredibly relevant! besjournals.onlinelibrary.wiley.com/doi/10.1111/...
besjournals.onlinelibrary.wiley.com
Four principles for improved statistical ecology
<em>Methods in Ecology and Evolution</em> is an open access journal publishing papers across a wide range of subdisciplines, disseminating new methods in ecology and evolution.
243