dshkol @dshkol.bsky.social · 12/06/2026github.com/dshkol/pycan...github.comGitHub - dshkol/pycancensus: python port of cancensus R packagepython port of cancensus R package. Contribute to dshkol/pycancensus development by creating an account on GitHub. 000
dshkol @dshkol.bsky.social · 12/06/2026pycancensus 0.2 out on pypi now - imho the best way to work with canadian census data in python for humans and agents alike - the apis are still cancensus (R) flavoured but the python code is better and more performant now, bugfixes - better caching + functions for working w hierarchical vectors 100
dshkol @dshkol.bsky.social · 15/01/2026data, geo, viz, cities, causal inf and ml, generative design and art, llms, anyone doing anything interesting i suppose i've managed to survive picking through the remnants of the other place while wearing a thick mental hazard suit, but its getting a bit much now 100
dshkol @dshkol.bsky.social · 13/01/2026and 2. I've found that these agents work better in scripts than in notebooks. Like they iterate better, hallucinate less, and troubleshoot faster and catch issues on an R/py script than they do in a notebook. I suspect the notebook backend itself (esp ipynb) chews through working context. 000
dshkol @dshkol.bsky.social · 13/01/2026I think the single quarto notebook would probably work fine paired with the same set of SKILLS and would reduce the need for the rest of the harness. But: 1. I was interested in a solution for a system more elaborate than a single notebook 100
dshkol @dshkol.bsky.social · 12/01/2026So check it out! Would love some feedback. Disclaimer: I am not in anyway affiliated with STC and this is not in any way an official publication. Use with caution! 120
dshkol @dshkol.bsky.social · 12/01/2026And a follow up post on what turned out to be the most interesting part -- fighting model hallucination and errors. This one gets a bit weedy. dshkol.com/posts/the-da...dshkol.comBuilding a lie detector for The D-AI-LY | Dmitry ShkolnikBuilding a system using skills and defensive engineering to catch a model that's very good at lying convincingly to you. 120
dshkol @dshkol.bsky.social · 12/01/2026I wrote about the process and open-sourced the repos on my site dshkol.com/posts/the-da...dshkol.comThe D-AI-LY: An Autonomous Statistical Digest | Dmitry ShkolnikReplicating Statistics Canada's The Daily using Claude Code alongside dedicated tools and a skills-based harness. 130
dshkol @dshkol.bsky.social · 12/01/2026The idea was to simultaneously layer: 1. SKILL documents to carefully fine tune instructions, expectations, skepticism and reasoning 2. Giving CC access to and forcing reliance on specialized tools like the cansim R package 3. Over-engineered data provenance tracking to trace each data point 110
dshkol @dshkol.bsky.social · 12/01/2026Nothing in any given release is all that complicated and could be easily one-shotted by any current AI tool. The challenge is consistency of execution against previously unseen data and defending against the kind of hallucination and data mistakes that LLMs are prone to. 110
dshkol @dshkol.bsky.social · 12/01/2026CANSIM has 60k+ tables. StatCan covers a handful at a time in The Daily. What if the cost of coverage was (nearly) zero? The D-AI-LY (dshkol.com/thedaily/) checks for recently updated and neglected series and writes up releases for them with viz, links to source material, and reproducible code. 240
dshkol @dshkol.bsky.social · 27/10/2025the library has been tested extensively for equivalence with equivalent R code in cancensus, but hasn't been extensively tested in the wild, so please use with caution and please leave feedback and issues on the github page 020
dshkol @dshkol.bsky.social · 27/10/2025this relied heavily on agentic cli coding tools (mostly cc + sonnet 4.5). agentic coders work really well when given lots of examples and a clearly defined reward function, which in this case relied on the extensive unit testing @jensvb.bsky.social and I built for cancensus 110
dshkol @dshkol.bsky.social · 27/10/2025the cancensus R package has had 81k dls since 2018. It's the best way to interact with StatsCan data for R users. pycancensus is a full python port, with equivalent data access and manipulation grammar, output, and geospatial retrieval. check it out here: github.com/dshkol/pycan...github.comGitHub - dshkol/pycancensus: python port of cancensus R packagepython port of cancensus R package. Contribute to dshkol/pycancensus development by creating an account on GitHub. 1101