Sign in

dbtool

@dbtool.bsky.social
11 followers 213 following 48 posts

Open models & data for meta-research. genderize (gender + country from a name): open weights, 98% accuracy, CPU only. Now building in public: sample size + PICO from clinical abstracts. EU-hosted API · free research keys · dbtool.it

PostsRepliesMedia
dbtool @dbtool.bsky.social · 29/09/2026
Genderize v2 is out: one open model for binary gender (M/F) + likely country (226 codes) from a personal name. Char-level, CPU-only. 94.08% on an independent 50,258-name bench (v1: 93.71%). Honest limit: weaker than v1 in Togo, Chad and Samoa. huggingface.co/textpie/genderize
100
dbtool @dbtool.bsky.social · 29/09/2026
New: colscan, a free offline CLI. Point it at a CSV, Parquet or database and it reports, column by column, the semantic type and whether the data is personal under GDPR. Deterministic rules, nothing sent anywhere. Apache-2.0. pip install colscan → pypi.org/project/colscan
100
dbtool @dbtool.bsky.social · 18/09/2026
genderize ultra reached 92.3% gender accuracy on Japanese public-figure names in our Wikidata benchmark, given-name first. Training overlap is unknown; this is not population accuracy. We publish limits alongside results. dbtool.it/academic.html
000
dbtool @dbtool.bsky.social · 16/09/2026
genderize cora reached 64% gender accuracy on Chinese public-figure names in our Wikidata benchmark, given-name first. Romanised names remain difficult. Training overlap is unknown; this is not population accuracy. huggingface.co/textpie/genderize
000
dbtool @dbtool.bsky.social · 14/09/2026
Our medical EBM-NLP test: 92.7% exact sample-size agreement between DeepSeek labels and expert-derived gold (114 of 123 abstracts with an unambiguous gold number). This measures the labeller, not our trained model, and is not overall PICO accuracy. dbtool.it/open-models.html
000
dbtool @dbtool.bsky.social · 13/09/2026
Build log, week 1. What's next at dbtool: one API call that turns a clinical abstract into sample size + PICO. One disk, one consumer GPU, one box. Every number ships with its test set; the public endpoint waits until it clears 80% on hand-annotated abstracts. Follow along → dbtool.it
001
dbtool @dbtool.bsky.social · 12/09/2026
genderize cora + ultra are out as open weights: byte-level, CPU-only gender + country (226 ISO codes) from a name. Gender 97.9%/98.2%, country top-1 82.6%/83.7% on 25k unseen names, network alone. CC BY-NC 4.0.
100
dbtool @dbtool.bsky.social · 12/09/2026
genderize cora & ultra are now open weights: a personal name → probable gender + a ranking of 226 countries, from 48 UTF-8 bytes. No tokenizer, no dictionary, CPU only. 98.2% gender accuracy on 25,000 held-out names. Weights CC BY-NC 4.0, script MIT. huggingface.co/textpie/genderize
000
dbtool @dbtool.bsky.social · 12/09/2026
genderize cora + ultra are out as open weights: byte-level, CPU-only gender + country (226 ISO codes) from a name. Gender 97.9%/98.2%, country top-1 82.6%/83.7% on 25k unseen names, network alone. CC BY-NC 4.0. huggingface.co/textpie/genderize dbtool.it/open-models.html
010
dbtool @dbtool.bsky.social · 03/09/2026
We ran our sex-of-participants model over all of human PubMed: 12.9M studies, 1960–2025. The mono-sex gap has flipped. 1960s: 46% male-only vs 22% female-only. 2020s: female-only (18%) overtakes male-only (17%) for the first time. #metascience #bibliometrics
100
dbtool @dbtool.bsky.social · 03/09/2026
Mapping the sex & age of study participants across all of human PubMed (~13M abstracts) with a BiomedBERT model trained on MeSH check-tags (sex accuracy ~90%). Preliminary, on 7M studies so far: 56% both sexes, 23% male-only, 21% female-only. #metascience #bibliometrics
000
dbtool @dbtool.bsky.social · 01/09/2026
What we're building now: a model that reads a biomedical abstract and tells you WHO was studied — female only, male only, both, or not reported. Trained on 7M MeSH-labelled abstracts. Half of human studies don't state participants' sex. Free for research when it ships. #metascience
000
dbtool @dbtool.bsky.social · 01/09/2026
Where this is going: releasing these models openly, long-term. @nrobinsongarcia.bsky.social's call asks for transparent, comprehensive, freely accessible methods — we meet two of three. The API fees fund the datasets and models that get us toward the third. Research use is already free.
010
dbtool @dbtool.bsky.social · 31/08/2026
We said: name a bench, we publish the result whatever it says. Nobody asked — so we ran it on ourselves. Public WGND names our models had never seen, vs the open tool nomquamgender. We lose on 5 countries and say so. We win on Japanese, Chinese, French. dbtool.it/benchmark #metascience
000
dbtool @dbtool.bsky.social · 31/08/2026
Same bridge, now for @bibliometrix.bsky.social users: from the M dataframe to gendered authorships, with the measured per-country error attached to every row. Country is passed only where bibliometrix honestly knows it — never guessed. dbtool.it/examples/bibliometrix_gen… #rstats
000
dbtool @dbtool.bsky.social · 31/08/2026
If you use @brunalab.bsky.social's refsplitr on WoS authors, here is the missing step for gender-gap studies: a workflow that uses the country refsplitr found per author — and attaches the measured per-country error to every row. dbtool.it/examples/refsplitr_gender… #rstats
000
dbtool @dbtool.bsky.social · 30/08/2026
We measured name-based gender assignment accuracy per country, on 400 held-out names each. Spain 98.3 · Germany 97.5 · Italy 97.5 · France 96.8 · USA 95.3 Japan 92.5 · Vietnam 92.0 · Korea 88.8 · Thailand 86.3 · China 83.5 · Taiwan 82.5 A 15-point gap between European and East Asian names.
1630