Sign in

Danielle Navarro

@djnavarro.net
2.9K followers 51 following 112 posts

mediocre transsexual and middle-aged sydneysider. tired, anxious, and fearful. occasionally writes sensible things about code, art, and statistics

PostsRepliesMedia
Danielle Navarro @djnavarro.net · 27m
A blog post about creating aperiodic de Bruijn tilings in #rstats
blog.djnavarro.net
Aperiodic tilings with the de Bruijn method – Notes from a data witch
A generative art tutorial on a geometric method I rather like, and some thoughts about artistic process
041
Danielle Navarro @djnavarro.net · 01/10/2026
I've heard some of their songs. I quite like it actually. But also yes black and pink is just an excellent combination. I keep wanting a black and pink flag for trans women just so that I don't have to be associated with the ghastly colour scheme of the trans flag
120
Danielle Navarro @djnavarro.net · 01/10/2026
thanks! I'm still refining the idea and trying to iron out all the workflow kinks, but so far I've been pretty happy with this. Eventually I'll write up a blog post about the minis repo, once I become more confident about the overall approach :)
010
Danielle Navarro @djnavarro.net · 01/10/2026
A small followup to this post. The appendix contains links for the data and data-generating script, but perhaps more usefully, also links to a more carefully worked out version of the cut_quantile() function than the simple one I used in the post minis.djnavarro.net/minicuts/
Appendix

- The toy data set is available for download as exposure_response_data.csv, and the script that generates it is also available as simulate_data.R.

- I’ve also put together a more carefully worked version of the cut_quantile() function that allows user control over the quantile-estimation method used to define break points, and three options for tie-handling behaviour at the break points, including the two traditional quantile/cut approaches (ties always fall upwards and ties always fall downwards), along with a third “ntile-like” approach that allows tie-splitting to produce more even bins, but only permits ties to be broken randomly (with built-in seed control), never by row order. It automatically detects when the data set cannot support the number of requested bins, and provides informative warnings and error messages as needed.The script itself is avaliable as minicuts.R and it has a brief accompanying vignette documenting what it does.
1141
Danielle Navarro @djnavarro.net · 30/09/2026
whatever
0220
Danielle Navarro @djnavarro.net · 24/09/2026
i'll be glad when people stop with these posts. some of us aren't safe in texas
111
Danielle Navarro @djnavarro.net · 22/09/2026
Because I lose my job if I don't do the analysis the client expects
040
Danielle Navarro @djnavarro.net · 21/09/2026
they came from a quick and dirty port of a python library for debruijn tilings. if i ever get time i might see if i can do something interesting with it as an art project
010
Danielle Navarro @djnavarro.net · 20/09/2026
yep, good idea. i've just added a small appendix section to the post that links to the data file and the script that generates it
screenshot of the part of the blog post that contains the two links
100
Danielle Navarro @djnavarro.net · 20/09/2026
Yeah. Probably not a bad idea? Arguably doesn't need to be that much, maybe a paragraph being precise about what the ntile() / NTILE approach is and where it differs from strict quantile binning. I imagine that that plus an example would be sufficient?
070
Danielle Navarro @djnavarro.net · 20/09/2026
I think the difference in implementations stems directly from the fact that the ggplot2::cut_() functions are built in R from scratch and not trying to mimic SQL behaviour. But dplyr::ntile() deliberately copies the behaviour of NTILE in SQL, which does not construct quantile-splits at all
130
Danielle Navarro @djnavarro.net · 20/09/2026
No, not as far as I can tell. They call cut() internally, same as I do in the example helper function, and then delegate the break-setting to the ggplot2:::breaks() internal function. From a quick look at the source for that, it seems to be using quantile() where needed. I think those are all okay
120
Danielle Navarro @djnavarro.net · 20/09/2026
Yes, very much this. The underlying issue is very much about handling edge cases correctly, but since we work in an industry that cares very deeply about doing so, it's worth spending a few lines of code to define a helper function that implements the thing you are actually trying to do
021
Danielle Navarro @djnavarro.net · 20/09/2026
Please, I am begging you. Do not use dplyr::ntile() in any serious statistical analysis in which you want to group continuous observation into quantile-based bins. Despite the name it doesn't actually do that, and the thing it actually does is very dangerous when accuracy matter. #rstats
blog.djnavarro.net
The trouble with ‘ntile()’ – Notes from a data witch
Please, please, please do not use ‘dplyr::ntile()’ in an analysis that will be submitted to a regulatory agency. It’s secretly an ill-conceived SQL tool that doesn’t do what you probably want it to, a...
68931
Danielle Navarro @djnavarro.net · 16/09/2026
Serious answer: because analysis scripts are typically grown interactively. As the analyst works they try things out. So yes, you *start* in a clean session but it takes active effort to make sure it *stays* clean while you work. That's the point of having tools that nudge users in that direction
120
Danielle Navarro @djnavarro.net · 15/09/2026
In her ongoing quest to convince #rstats users to stop invoking the cursed "rm(list=ls())" incantation at the top of their analysis scripts, she has released sessioncheck 0.2 on CRAN. Mostly unchanged, some bug fixes, some new reporting features that are most useful in the regulatory context
blog.djnavarro.net
Release: sessioncheck 0.2 – Notes from a data witch
In which the author realises she doesn’t know how to write package release posts
4439
Danielle Navarro @djnavarro.net · 24/08/2026
exactly! the sheer amount of space that the em-dash uses up makes it feel like a loud and dramatic flourish rather than a small, gentle aside to the reader
010
Danielle Navarro @djnavarro.net · 20/08/2026
pretty much. there are contexts where sex-at-birth data is needed, but those are the exception not the norm, and in those cases it's critical to be very strict in restricting data access to people with a legitimate need to know. it should never be treated as quasi-public data
030
Danielle Navarro @djnavarro.net · 09/08/2026
Honestly, I'm shocked at how little scrutiny these things have received. I left the academic system in 2021 so I had not paid close attention until discovering that those extremely intrusive birth sex queries have now escaped academia and made it into the wild in Australia. It's truly appalling
130
Danielle Navarro @djnavarro.net · 08/08/2026
That's a big factor in all this yes. The older "sex" field was always a mix: people used it to mean what they wanted it to mean. After a "sex/gender" split, neither of the new variables has continuity with the old one, and the new sex variable contains more sensitive data than the old one did
150
Danielle Navarro @djnavarro.net · 08/08/2026
That's a nice contrast to my own situation. If you just want to characterise the sample, I'll tick the F box. But if you're doing group comparisons I have a lot of different answers depending on precisely what comparison is involved. Even for biological comparisons my group assignment isn't easy
020
Danielle Navarro @djnavarro.net · 08/08/2026
Yeah. That was one of the reasons why I never allowed gender to be a free-response field on any of my studies. I didn't want to look at the vile things people wrote, and I didn't want my RAs to do it either. People cannot be trusted with free-response on a gender question
181
Danielle Navarro @djnavarro.net · 08/08/2026
One nice side-effect that has come from me forcing myself to write the cursed post: I've ended up in conversations with people who sit on HREC committees in Australia and got them to think about the ethical implications of including birth-sex questions on studies that don't strictly need it
180
Danielle Navarro @djnavarro.net · 08/08/2026
I mentioned this in another subthread, but think of the analogy to sex work. It would very likely be useful for practical purposes to have population-level data about prevalence and distribution of sex work, but you simply cannot use the census for that purpose: it's intrusive and people will lie.
170
Danielle Navarro @djnavarro.net · 08/08/2026
I think it's a very bad idea, simply because they cannot do it well. In the hypothetical scenario where it could be done well, yes, it would be great for service provision. But since unicorns aren't real and wishful thinking is dangerous, I don't think the census should be doing this.
160
Danielle Navarro @djnavarro.net · 08/08/2026
The end result of that "precise" sex/gender measurement is that trans people get alienated and become motivated to lie: it breaks the measurement validity. Better to lose some reliability but gain validity, I think. But I admit I still don't quite know how to word the questions properly :/
170
Danielle Navarro @djnavarro.net · 08/08/2026
The thing that nobody really thought about is that asking trans people our birth sex is intrusive and offensive. And if you do the thing that most orgs do and make sex mandatory and gender optional, it sense the message that you don't really think our gender is as important as our genitals. So...
180
Danielle Navarro @djnavarro.net · 08/08/2026
I think I am, yeah. The phrasing is difficult to get right because it's an ambiguous thing to measure and that makes it hard. From memory, one of the reasons that we ended up with this sex-at-birth vs current-gender-identity split in Australia is that the birth sex measure feels more precise. But...
170
Danielle Navarro @djnavarro.net · 07/08/2026
I will also add that this is the post where I have tried extremely hard to be calm and careful and not allow my emotions into the argument. But the bigger picture on this is that I am not calm. I am furious. It is a terrible data collection method that does not work and is cruel and degrading too
0191
Danielle Navarro @djnavarro.net · 07/08/2026
riggghhttttt? like, you cannot possibly read any exchange between fisher and neyman and think that statistics is not a discipline for bitches. this is our homeland
2150
Danielle Navarro @djnavarro.net · 07/08/2026
Thank you Chelsea. And I apologise for not going with my usual crass and snarky style on this one. I really wanted to but felt I should play a straight bat on this one. I promise I will return to my usual being-a-bitch-about-splines style on the next one 🫠
180
Danielle Navarro @djnavarro.net · 07/08/2026
Thank you. It wasn't the easiest post for me to write, but probably something that needed to be said. There are a lot of counterintuitive aspects to data collection as it pertains to trans people, and it's easy to make dangerous mistakes despite having the best of intentions
070
Danielle Navarro @djnavarro.net · 07/08/2026
You're welcome. I confess I wrote it mostly for personal catharsis rather than in the expectation that it would make much of a difference, but I am glad a few people are reading it and thinking about it. (Also, I owe you an email - I promise I'm not ignoring it I've just been a bit overwhelmed)
120
Danielle Navarro @djnavarro.net · 07/08/2026
exactly this. the transgender population is much more hidden and stigmatised than the gay population. the proper analogy is to something like sex work. you cannot include questions about sex work in a census because it's intrusive and people will lie. you need specialised methods for that
0100
Danielle Navarro @djnavarro.net · 07/08/2026
Basically my view is now that it is important to have good data on the transgender population to inform policies, but the census is the wrong instrument with which to collect it. The data collection is much more sensitive than with sexual orientation, and won't work for a population-wide census
190
Danielle Navarro @djnavarro.net · 07/08/2026
Worse, the motivation to lie will differ across trans subgroups: the questions are much more intrusive for a transgender woman than they are for an AFAB non-binary person. So your relative proportions are going to be messed up
170
Danielle Navarro @djnavarro.net · 07/08/2026
Correct. But the census won't get a useful count even with the questions. Cis people won't answer the optional gender question, and trans people will lie on the mandatory birth sex question.
150
Danielle Navarro @djnavarro.net · 07/08/2026
Apologies for being snippy. But yes, my point is not that doctors shouldn't know. It's that doctors should not use a data collection methodology that is intrusive, guaranteed to produce data leakage, and undermine trust with their patients
150
Danielle Navarro @djnavarro.net · 07/08/2026
That's literally what the post says
120
Danielle Navarro @djnavarro.net · 07/08/2026
Some clarifications added to the post, covering points I wouldn't have thought needed to be made. No, pathology labs do not need to know birth sex. Yes, doctors often do need to know birth sex. No, that does not make it appropriate to put on a patient intake form. None of this should be contentious
3301
Danielle Navarro @djnavarro.net · 06/08/2026
The title is a little pointed, but I am very serious here. Academics, medical researchers, and other organisations should not collect data on birth sex. It is rarely relevant, and it almost always leads to privacy violations for transgender people. Please share. I would like people to read this one.
blog.djnavarro.net
Separating sex and gender is statistical malpractice – Notes from a data witch
Statisticians and academics in Australia should ignore the terrible advice of most LGBTIQA+ organisations and read this post instead. The 2020 ABS standard is almost never the correct way to collect d...
14236119
Danielle Navarro @djnavarro.net · 06/08/2026
Agreed. The ABS has gotten this horrifically wrong. Trans people will not give them the answer they're expecting for the sex-at-birth question because we've already had experiences with other orgs copying the ABS approach and learned that it leads to forced outing. It's a terrible policy decision
010
Danielle Navarro @djnavarro.net · 06/08/2026
In fairness, I didn't predict it either. My worries at the time were about infrastructure and data leakage. I'd seen it internally at UNSW where the HR records were fractured, and kept leaking my trans status into the classroom. But now it's happening to medical data too because of these policies :(
140
Danielle Navarro @djnavarro.net · 06/08/2026
yeah. I remember the well intentioned policy work that was done around it at the time. I spoke to some of the folks at UNSW who were involved, and tried to highlight some of the risks that it would create. In hindsight I wish I'd pushed harder, because I don't think anyone really believed me on this
130
Danielle Navarro @djnavarro.net · 26/07/2026
thank you Mike ❤️
020
Danielle Navarro @djnavarro.net · 26/07/2026
not literally about cocaine
blog.djnavarro.net
A cocaine skeptic starts an unwise habit – Notes from a data witch
You will be shocked to learn that this is not actually a post about cocaine
45619
Danielle Navarro @djnavarro.net · 18/07/2026
And so after a very brief hiatus, version 0.7.0 of "learning statistics with R" exists. Full rebuild in quarto, stylistic fixes, "epilogues from 2026" to comment on how the world changed since original publication, and as an added bonus, no longer misgenders the author learningstatisticswithr.com
learningstatisticswithr.com
Learning Statistics with R
217146
Danielle Navarro @djnavarro.net · 10/07/2026
truly. just bizarre. idk what kind of bubble you have to have been living in that *this* was your wake-up call but wow... it must be very nice there
030
Danielle Navarro @djnavarro.net · 04/07/2026
very much so. it's not so much malicious as it is thoughtless. they aren't trans, and they refuse to listen to trans people when we try to explain the facts of life to them. they believe they know what's best for us, they make our lives worse, and then they expect us to be grateful to them
110
Danielle Navarro @djnavarro.net · 13/05/2026
not entirely surprised. i have had this happen in my own use of the package. i've decided to think of it as a feature: any time i notice someone has commented that line out i take their version of the results with a grain of salt since i can't be sure about the execution environment
110