discourse.datamethods.org
Using robcov() to adjust for donor-level clustering in kidney transplant outcomes?
I am doing a study to assess the association between transplant factors and kidney transplant outcomes (specifically survival of the transplanted kidney) using US registry data. I’m using `cph()`.
A key issue is that when looking at outcomes, most recipients are not truly independent as they are clustered by donor. Specifically, in our analysis we have 57,153 donors where both kidneys were transplanted (into different recipients) and 39,389 donors where only one kidney was transplanted into a single recipient. So our analysis on recipient outcomes has n=153,695 recipients, and 96,542 donor clusters (each of which has either 1 or 2 recipients).
The focus of the research study is to assess if how the kidney was preserved impacts on kidney graft survival, and I adjust for several donor and recipient factors in the model.
This is how I am currently using robcov to get cluster robust SE (from which I am calculating 95% confidence intervals)
fit <- cph(Surv(time, event) ~ var1 + var2 + var3, data, x = TRUE, y = TRUE)
fit_robust <- robcov(fit, cluster = data$donor_id)
My understanding is that this uses the Huber-White method for updating the variance-covariance, which will increase the variance due to the clustering. I’ve seen other papers in the field use frailty terms (or other mixed effect models for non-survival outcomes). However, I think that robust SE using `robcov()` is probably a better solution here.
Is this a valid approach, given the large number of clusters, but small cluster size (either 2 or 1 per cluster)? Would the same approach be valid for logistic regression and linear (fit with `ols()`)?