Sign in

zx8754

@zx8754.bsky.social
151 followers 982 following 162 posts

Analysis - Data - Analysis - Data - Volleyball - Father - #rstats #STEM #ProstateCancer

PostsRepliesMedia
zx8754 @zx8754.bsky.social · 02/07/2026
I was today years old when I learned base #Rstats has a helper function `example()`. Better late than never, and it's only been half a three decades give or take ...
1142
zx8754 @zx8754.bsky.social · 04/11/2025
#rstats plot of the day - "Listogram" by r2evans stackoverflow.com/a/79808261/6...
listogram plot - Histogram of three-letter words, with the words themselves shown on the plot.
151
zx8754 @zx8754.bsky.social · 23/10/2025
"my default assumption is nothing is going to work" - @clauswilke.com
quote from the linked post: "my default assumption is nothing is going to work"

https://blog.genesmindsmachines.com/p/we-still-cant-predict-much-of-anything
010
zx8754 @zx8754.bsky.social · 23/10/2025
set.seed((\(x) sum(x %% 10) + length(x))(as.integer(charToRaw("Forty-two")))) #rstats
040
Reposted by zx8754
The Institute of Cancer Research @icr.ac.uk · 17/10/2025
Men with BRCA1 and BRCA2 mutations should get annual #ProstateCancer screening. The IMPACT trial finds BRCA1 carriers are 3 times more likely to have aggressive prostate cancers than non-carriers. Results from the international trial shared at @myesmo.bsky.social www.icr.ac.uk/about-us/icr...
033
zx8754 @zx8754.bsky.social · 08/10/2025
#rstats easier to compare the ratios
Barplots of population change ratios in "Most populous countries in Europe" 2000 vs 2025.

R code for the plot:

library(data.table)
library(ggplot2)
library(scales)

d <- fread("
Rank,Country,y2000,y2025
1,Russia,146000000,144000000
2,Germany,82800000,84080000
3,Britain,59510000,69550000
4,France,59330000,66660000
5,Italy,57640000,59150000
6,Ukraine,49430000,38980000
7,Spain,40000000,49320000
8,Poland,38650000,38140000")

d[, rate := y2025/y2000 ]
d[, Country := factor(Country, levels = d[ order(rate), Country]) ]
d[, lab := paste0(round(rate, 4), "\n", 
                    label_number(accuracy = 0.1, scale_cut = cut_short_scale())(y2025 - y2000)) ]
d[, col := as.factor(sign(y2025-y2000)) ]

ggplot(d, aes(Country, rate, label = lab, colour = col)) +
  geom_col(fill = "white", show.legend = FALSE, linewidth = 1.1) +
  geom_hline(yintercept = 1, colour = "blue", linetype = "dashed") +
  geom_text(vjust = 1.4, colour = "black") +
  scale_x_discrete(name = NULL) +
  scale_y_continuous(name = NULL, breaks = c(0, 1), limits = c(0, 1.3), expand = c(0, 0)) +
  scale_color_manual(values = c("-1"="red", "1" = "darkgreen")) +
  labs(title = "Most populous countries in Europe",
       subtitle = "2000 vs 2025",
       caption = "Source: @GEOMAPAS.GR | https://bsky.app/profile/simongerman600.bsky.social/post/3m2nz34kzyi27") +
  theme_minimal() +
  theme(panel.grid = element_blank())
ggsave("tmp.plot.jpeg", width = 8, height = 6)
010
zx8754 @zx8754.bsky.social · 02/09/2025
#rstats #TIL I knew about tibble::tribble, but just discovered data.table::rowwiseDT. "creates a data.table object by specifying a row-by-row layout. This is convenient and highly readable for small tables."
Create a data.table row-wise
Aliases: rowwiseDT

Keywords:

### ** Examples

rowwiseDT(
  A=,B=, C=,
  1, "a",2:3,
  2, "b",list(5)
)
       A      B      C
   <num> <char> <list>
1:     1      a    2,3
2:     2      b      5You can define a tibble row-by-row with tribble():

tribble(
  ~x, ~y,  ~z,
  "a", 2,  3.6,
  "b", 1,  8.5
)
#> # A tibble: 2 × 3
#>   x         y     z
#>   <chr> <dbl> <dbl>
#> 1 a         2   3.6
#> 2 b         1   8.5
040
zx8754 @zx8754.bsky.social · 01/09/2025
#rstats pipe help, where is this from? I am guessing it is "modify and re-assign". %<+% Google didn't help, Gemini says it is invalid R code or it is a regex. ChatGPT says it is from ggplot, searching GitHub didn't help. But found this function ggplot2::`%+replace%`(), are they related?
340
zx8754 @zx8754.bsky.social · 21/07/2025
As base::rbind also takes more than two dataframes as input, I expected at least the same from dplyr::bind_rows. And bind_rows handles rownames, factor levels, and mismatched column names better. #rstats
#list of dfs
set.seed(1); listOfDfs <- lapply(1:3, \(i){ mtcars[ sample(1:nrow(mtcars), 4), ]})

#input as list
do.call(rbind, listOfDfs)
bind_rows(listOfDfs)

#input as dataframe
rbind(listOfDfs[[ 1 ]], listOfDfs[[ 2 ]], listOfDfs[[ 3 ]])
bind_rows(listOfDfs[[ 1 ]], listOfDfs[[ 2 ]], listOfDfs[[ 3 ]])
030
zx8754 @zx8754.bsky.social · 20/06/2025
I love waffle charts, but this can be done with points. #rstats #ggplot
set.seed(2025)

d1 <- data.frame(expand.grid(1:50, 1:50))
d1$hot <- factor("no", levels = c("no", "yes"))
d1[ sample(seq(nrow(d1)), 1), "hot" ] <- "yes"
d1$time <- factor("before", levels = c("before", "today"))

d2 <- data.frame(expand.grid(1:50, 1:50))
d2$hot <- factor("no", levels = c("no", "yes"))
d2[ sample(seq(nrow(d2)), 100), "hot" ] <- "yes"
d2$time <- factor("today", levels = levels(d1$time))

library(ggplot2)

ggplot(rbind(d2, d1), aes(Var1, Var2, colour = hot)) +
  geom_point() +
  facet_wrap(vars(time)) +
  scale_colour_manual(values = c("no" = "grey90", "yes" = "red"), guide = "none") +
  theme_void()
040
zx8754 @zx8754.bsky.social · 29/05/2025
#rstats base ?`$`
library(dplyr)
library(purrr)
library(bench)

set.seed(42)

create_df <- function(rows, cols) {
  out <- replicate(cols, runif(rows, 1, 100), simplify = FALSE)
  out <- setNames(out, rep_len(letters, cols))
  as.data.frame(out) }

results <- press(
  rows = c(1000, 10000, 100000),
  cols = c(100,  1000,  10000),
  {
    d <- create_df(rows, cols)
    mark(
      d$f,
      d %>% .$f,
      d %>% pull(f),
      d %>% pluck("f"),
      min_iterations = 100
      )
    }
  )

ggplot2::autoplot(results)
bench::mark output plot
3162
zx8754 @zx8754.bsky.social · 27/05/2025
In case you are wondering about #rstats base: 1.0.0 = 1195 vs 4.5.0 = 1286
061
zx8754 @zx8754.bsky.social · 22/05/2025
#rstats base and purrr alternatives
df <- read.table(text = "stu_id	race1	race2	race3	race4	race5	race6
100	0	0	0	0	0	1
101	1	0	0	0	0	0
102	0	1	0	0	0	1", header = TRUE)

f <- function(...){
  ix <- which(c(...) == 1)
  if(length(ix) == 1) ix else 7L
  }

#base
apply(df[, -1], 1, f)
# [1] 6 1 7
#to create new column
#df$race <- apply(df[, -1], 1, f)

#purrr
pmap_int(df[, -1], f)
# [1] 6 1 7
#to create new column
# df <- df %>%
#   mutate(race = pmap_int(select(., -1), f))
041
zx8754 @zx8754.bsky.social · 05/05/2025
#rstats
meme:

Her: Babe please stop keeping 173 tabs open
Me: One of them has an R error from 3 weeks ago I'm still debugging
020
zx8754 @zx8754.bsky.social · 29/04/2025
Use a string in facet_grid to get a "title" facet_grid(dose ~ "len") #rstats #ggplot
ToothGrowth |> 
  ggplot(aes(x = dose, y = len)) + 
  geom_boxplot(aes(fill = supp)) +
  labs(y = NULL) +
  theme_bw() +
  facet_grid(dose~"len")

https://stackoverflow.com/a/79597025/680068
010
zx8754 @zx8754.bsky.social · 25/04/2025
#rstats base and tidy having a chat (this is not a dig to anyone, use whatever works)
# tidy: count 1s on subset of columns
# base: try this
colSums(mtcars[, c("vs", "am", "gear", "carb") ] == 1)
# tidy: but I don't want quotes...
# base: ok, cols by index
colSums(mtcars[, 8:11 ] == 1)
# tidy: but I want to see column names!
# base: ok, let's use expressions
colSums(mtcars[, as.character(substitute(list(vs, am, gear, carb)))[ -1 ] ] == 1)
# tidy: oh no, too many nested brackets, my head hurts!
# base: wrap into a fucntion
count_ones_base <- function(data, ...) {
  colSums(mtcars[, as.character(substitute(list(...)))[ -1 ] ] == 1)
}
count_ones_base(mtcars, vs, am, gear, carb)
# tidy: but I like pipes, weird function names like quos, 
#       and to be loud !! about it !!!
# base: ok.
library(dplyr)
library(tidyr)
170
zx8754 @zx8754.bsky.social · 23/04/2025
Love it! Borrowing Alex's idea, here is my attempt #rstats #crayon #inkgolf
library(crayon)
# 4 dark blue, 3 lighter blue, 2 pink, 1 red
cat(    
  sapply(c("blue4","blue3","blue2","blue1",
           "royalblue3","royalblue1","cadetblue1",
           "deeppink3","deeppink1","red3"),
         \(i) make_style(i)("."))
  )
000
zx8754 @zx8754.bsky.social · 23/04/2025
#rstats version
Google: rstats a day keeps stata away meaning

AI response: "RStats a day keeps Stata away" is a humorous way of saying that if you regularly use the R statistical programming language, you'll be less likely to rely on Stata, another statistical software package. It implies that consistent practice and knowledge of R can make it your primary tool, reducing the need for Stata.
0112
zx8754 @zx8754.bsky.social · 17/04/2025
my current state: import #squeakycleandata from #excel to #rstats youtu.be/vxdQov7gses
youtu.be
Please stop sending me your datasets.
YouTube video by Darren Dahly
341
zx8754 @zx8754.bsky.social · 08/04/2025
#rstats #ggplot version of this interesting approach data and codes are at gist.github.com/zx8754/85d0c...
R ggplot version of alternative to stacked bar plot. 

link to code for the plot:

https://gist.github.com/zx8754/85d0c669a3048eb58d7c58a11eaef9a2
030
zx8754 @zx8754.bsky.social · 11/11/2023
Hello world!
050