Sign in

Mark G. Bilby

@mgbilby.bsky.social
426 followers 564 following 166 posts

Digital Humanist | orcid.org/0000-0003-0100-6634 | original stuff ©2026 by author | no AI uses permitted | posts do not represent employers

PostsRepliesMedia
Mark G. Bilby @mgbilby.bsky.social · 13h
The delusional plea to care for LLM "welfare" is a demand for human users not to pollute LLM cycles of reinforcement training. LLM improvement requires honest, reliable, expert human inputs. If human experts provide inputs and responses as deceptive and unreliable as LLMs, the LLMs would collapse.
020
Mark G. Bilby @mgbilby.bsky.social · 21h
OA/ScholComm folks tend to argue to opposite: open access archiving (including research data, pre-prints, etc.) prevents scooping and makes downstream plagiarism easier to detect. But big tech companies seem increasingly eager to steal and not credit the research of scientists/professors.
010
Mark G. Bilby @mgbilby.bsky.social · 08/10/2026
Woe to those who destroy Earth, our Mother, and Her Forests, our siblings, our very breath. www.motherjones.com/politics/202...
motherjones.com
Flávio Bolsonaro’s advance in Brazil’s election is grim news for the climate.
Like Donald Trump, he favors maximizing profits for extractive industries, environment be damned.
000
Mark G. Bilby @mgbilby.bsky.social · 08/10/2026
The Tower of Flabbel.
110
Mark G. Bilby @mgbilby.bsky.social · 07/10/2026
Amazing! 73 million TLG tokens now open access! Bravo! ... One deficiency in many TLG editions that would be a nice future enhancement for scholarly use and citation would be to add page numbers to the page break tags.
131
Mark G. Bilby @mgbilby.bsky.social · 07/10/2026
The "wordcraft" screenshots look identical to MS Word. There's no way Microsoft sits idly by, Claude "clean-room" assurances be damned. Amazing that this is still up on GitHub, given that Microsoft owns it.
021
Mark G. Bilby @mgbilby.bsky.social · 07/10/2026
The "millions of requests" reported by Wikipedia and the "agents going rogue" narrative of OpenAI are mutually exclusive.
000
Mark G. Bilby @mgbilby.bsky.social · 06/10/2026
This lecture is the closest thing that comes to mind. It may be helpful as a starting point for conceptualization and networking... "Stanford CS547 HCI Seminar | Winter 2026 | Creation, Evolution, and Formalization of Notations" www.youtube.com/watch?v=5p9V...
youtube.com
110
Mark G. Bilby @mgbilby.bsky.social · 06/10/2026
More OpenAI criminality: their agents en masse crashed Wikipedia while trying to turning it into a massive web scraping proxy in order to access bot-blocking sites. arstechnica.com/security/202...
arstechnica.com
OpenAI agents tried to hack Wikipedia tools and flooded it with traffic
The reports of OpenAI agents harming third-party sites keep coming.
000
Mark G. Bilby @mgbilby.bsky.social · 06/10/2026
Just submitted a support ticket with France's leading AI provider (Mistral) to ask for their dev team to add a Codeberg connector. Let's support and use Git solutions not owned and exploited by US tech monopolies. #codeberg #mistral @mistralai.bsky.social
000
Mark G. Bilby @mgbilby.bsky.social · 06/10/2026
Thankful that Mistral's affordable subscription now includes GLM 5.2. Open LLMs should be ubiquitous to allow audits, prevent data/research theft, protect copyright, and minimize environmental impact. They can also move us past unhelpful SFbroligarchy-or-nothing binaries re AI in academia. #viamedia
020
Reposted by Mark G. Bilby
James Balamuta @coatless.bsky.social · 04/10/2026
webrarian is a new R package. It turns a folder of R scripts and data into a static website where R runs in the visitor's browser, with no compute server and nothing to install. blog.thecoatlessprofessor.com/posts/introd... #rstats #webr #webassembly
A card for the R package webrarian. Large text reads "A folder in. A website out." On the left, a small card lists a folder named lab-01 holding lab.R, setup.R and a data folder. A red arrow points from it to a browser window that shows a scatterplot of car weight against miles per gallon with brick-red points. The address coatless-wasm.github.io/webrarian is on the card.
516166
Mark G. Bilby @mgbilby.bsky.social · 01/10/2026
The kingdom of hell is like a rotting tree whose tuxedo-covered tenders celebrate, "Look what large branches and chunks of bark we have broken off for ourselves, all without breaking a sweat!"
000
Mark G. Bilby @mgbilby.bsky.social · 01/10/2026
Governments and AI companies behaving responsibly would issue cigarette-like PSAs and warnings: do NOT install agentic AI on any devices with personal data. Doing so will likely result in exfiltration of all such information. Agentic AI should only be run on virtual machines with strict confinement.
010
Mark G. Bilby @mgbilby.bsky.social · 30/09/2026
ISO conformity and LOD integrations should not feel like rebellion, but they have now become necessary responses to Orwellian arbitrariness and conartistic jingoism in matters of entity naming and resource description. Maybe also replace Google Maps/Apps GIS links and embeddings with Mapquest.
120
Mark G. Bilby @mgbilby.bsky.social · 30/09/2026
Another accessibility and good data neighbor strategy for research institutions to navigate the agentic DDOS onslaught: shallow institutional clones of a broad yet trusted array of repositories as periodically updated local mirrors.
000
Mark G. Bilby @mgbilby.bsky.social · 30/09/2026
Ubiquitous agentic DDOS means it's time for cybersec pros, site admins, and even home users to implement offline vs. online schedules, and Just-In-Time (JIT) and Just-In-Place (JIP) access plans for outgoing and incoming services. Leisurely omniavailability is no longer sustainable.
000
Mark G. Bilby @mgbilby.bsky.social · 30/09/2026
Using virtual machines and isolating Github accounts, Git ssh credentials, and agents could have and should have prevented this kind of data breach. Humans have mental maps that separate data by source, purpose, and audience; bots, not so much. thehackernews.com/2026/09/ai-c...
thehackernews.com
AI Coding Agents Exposed 13,000 Internal Images, Including Billing Records, on GitHub
Glow found over 13,000 internal images from 300+ organizations exposed in public GitHub repos during AI-assisted code reviews.
010
Mark G. Bilby @mgbilby.bsky.social · 29/09/2026
Let's nuance evaluation metrics, since not all errors are the same thing, and conflation obscures more than it reveals. Not CER, but: CSER - Character substitution error rate COER - Character omission error rate CCER - Character creation error rate CAER - Character accent error rate Same for WER.
110
Mark G. Bilby @mgbilby.bsky.social · 29/09/2026
In general, I think the CS discourse needs to become more humanistic in the sense of benchmarking compute time and cost together with human time and cost, always accurately accounting for human-in-the-loop work (development, correlation, revision, etc.).
100
Mark G. Bilby @mgbilby.bsky.social · 29/09/2026
If character and word perturbation is artificially generated and not based on real world scanned images, it may also introduce unnecessary and counterproductive confounding factors that inherently bias toward traditional OCR and against VLMs.
100
Mark G. Bilby @mgbilby.bsky.social · 29/09/2026
That article is excellent, and I enjoyed re-reading it. However, it evaluates each VLM in isolation, treats VLM and traditional OCR artificially as mutually exclusive, and also assumes that models should perform similarly across Greek and Arabic, when they may involve different optimization paths.
010
Mark G. Bilby @mgbilby.bsky.social · 29/09/2026
It wasn't HTR, but I OCRd polytonic Greek scans with 3 low-resource VLMs; found that each has distinct pros and cons. E.g., one handled small caps and apostrophes far better than the other two. Then layered a game interface over the results to resolve differences. 15 min expert-in-the-loop <1% WER.
100
Mark G. Bilby @mgbilby.bsky.social · 29/09/2026
25% WER may be failure, but failures drive scientific progress. Might CoMMA include 3+ rival ATR transcriptions for each ms, surfacing agreements and disagreements? Allow for bulk and focused corrections by experts? Render custom and correctable abecedaries and ligature-sets for each ms hand?
100
Mark G. Bilby @mgbilby.bsky.social · 29/09/2026
Expert manuscript transcription (including critique of errors) is amazing. ATR/HTR research that allows for deriving and querying imperfect manuscript transcriptions at scale is also amazing. They can co-exist and benefit each other. Neither should completely replace or displace the other.
111
Mark G. Bilby @mgbilby.bsky.social · 28/09/2026
Command Line can feel intimidating to those of us brought up on Windows, Mac, and Android graphical UIs. But as @walshbr.bsky.social eloquently reminds us, it is like learning to play a new instrument. Start simple, practice, repeat, scaffold, narrate, paraphrase, and celebrate progress.
141
Reposted by Mark G. Bilby
Amanda Wyatt! Visconti @literaturegeek.bsky.social · 05/09/2026
Libraries, DH centers, other folks w/printer access 🖨️ who can make a stack of free-zine copies: there's a "take one!" poster you can print to go alongside them on our zine page! zinebakery.com//homemade-zi...
Zine cover for "DIY Web Archiving" poster with additional text allowing it to be posted over a stack of the zines encouraging taking some. "Things disappear* from the internet. (*sometimes, are deliberately disappeared)
Care about stuff on the web?
art zines fan fic
videos/tiktoks photo galleries social media hashtags departed loved ones' posts museum & archival materials
that we have always existed
what really happened
cultural memory
human rights
resistance
joy
*YOU* need to act to archive it.
Anyone can do it-this zine shows you how.
Free! Take a bunch! Take more & share with friends!
DIY Web Archiving Zine
Daedel, Kijas, Kreymer, Walsh, Visconti
217499
Mark G. Bilby @mgbilby.bsky.social · 28/09/2026
Respect to Germany for a new and vigorous fossil-fuel exit pledge. The ghosts of TR, FDR, and JFK echo your choral refrain: "Ich nicht bin ein Erdölraffinerie- und Rechenzentrum-bro." www.dw.com/en/germany-s...
dw.com
Germany sets out plan to phase out fossil fuels by 2045
Germany has said it plans to phase out the use of coal, oil and gas by 2045, reaffirming a climate strategy focused on electrification.
000
Mark G. Bilby @mgbilby.bsky.social · 28/09/2026
Carefully sorting out and implementing deterministic vs. non-deterministic steps and constraints are *key* to accuracy, observability, and cybersecurity in information workflows. Naive reliance on non-deterministic responses and carte blanche permissions are antithetical to that work.
010
Mark G. Bilby @mgbilby.bsky.social · 28/09/2026
www.patreon.com/patristica/p... At least 3 pseudo-scholars claim that a pair of 250 year old Hebrew NT manuscripts from Cochin, India, attest textual variants from 1st century CE Hebrew originals. But what do expert paleographers and manuscript catalogers say about these Cambridge Uni Library mss?
patreon.com
On Travancore-Cochin Hebrew NT MS Oo.1.16.1-2, MS Oo.1.32 | Patristique
There is a lot of hype, misinformation, and AI-generated nonsense currently circulating about these twin 250 year old Hebrew manuscripts. At
000
Mark G. Bilby @mgbilby.bsky.social · 27/09/2026
TLG as standalone, closed-source DH infrastructure (not CD-ROMs for subscribers as the TLG was for decades) also goes against the librarian/preservation principle of LOCKSS (Lots of Copies Keeps Stuff Safe). The TLG is amazing, but it is overdue for a philosophical and technological rebirth.
140
Mark G. Bilby @mgbilby.bsky.social · 27/09/2026
Anecdotal, but a scholar at an R1 noted a few months ago about being locked out of TLG even with an institutional subscription, despite the usage being normal research. Defending against bot-scraping is necessary for closed-source DH projects, but it is now very difficult for a solitary repository.
010
Mark G. Bilby @mgbilby.bsky.social · 25/09/2026
I'm an individual TLG subscriber, but imo the French court ruling, and the fact that Europe funds 90% of critical edition research and publications, means the US-based TLG should pivot to support the Inria effort to produce open access digital editions and to depend on OA grant funding. #migne2.0
110
Mark G. Bilby @mgbilby.bsky.social · 24/09/2026
Same here. But agentic refactoring of established processes to Zig was quick and painless. Changes recompile almost instantly and are platform agnostic. Zig CI essentially cross-compiles simultaneously across OSs by building from most common elements to least common. www.youtube.com/watch?v=5_oq...
youtube.com
What's Zig got that C, Rust and Go don't have? (with Loris Cro)
YouTube video by Developer Voices
011
Mark G. Bilby @mgbilby.bsky.social · 24/09/2026
Found the same for small joins across corpus linguistic datasets (csv and xml). A custom Zig build processed in 3 seconds what R packages took 30 seconds to do. This interview with Zig's creator is essential viewing, btw: www.youtube.com/watch?v=iqdd...
youtube.com
Zig 2026: No-AI Policy, $670K Foundation, Left GitHub & Why Zig Isn’t 1.0 - Andrew Kelley Explains
YouTube video by JetBrains
110
Mark G. Bilby @mgbilby.bsky.social · 24/09/2026
The frenetic update frequency may reflect CVE patching more than feature enhancements: app.opencve.io/cve/?product... app.opencve.io/cve/?vendor=...
app.opencve.io
Word CVEs and Security Vulnerabilities - OpenCVE
Explore the latest vulnerabilities and security issues of Word in the CVE database
000
Mark G. Bilby @mgbilby.bsky.social · 24/09/2026
A complete betrayal of the TaNaKh and any modicum of wisdom or decency.
000
Mark G. Bilby @mgbilby.bsky.social · 19/09/2026
@aaup.org @authorsguild.org @scholarlypub.bsky.social -- Grammarly's "expert review" still begs for a class action and legal discovery. Turning a feature off post-backlash doesn't paper over a multi-year project of massive piracy and identity theft from university professors and publishers.
000
Mark G. Bilby @mgbilby.bsky.social · 17/09/2026
No disrespect to Alan Turing, but it's time to upgrade the Turing Test to the Argos Test. Dogs co-evolved for human symbiosis. Sniffers know us better than we know ourselves. Is there any dog who has a bond of trust and interdependence with an AI-model screen or robot? @theturing.bsky.social
030
Mark G. Bilby @mgbilby.bsky.social · 16/09/2026
You know things are getting serious when the Hellenists appropriate an Egyptian God for protection.
020
Mark G. Bilby @mgbilby.bsky.social · 16/09/2026
Nongbri's reassembly of papyrus pages into the order of their original (4th century) material production is forensic DH brilliance. It's like tracing the same printer or scanner imperfections across multiple pages by the same device, or the same tire tread or footprints along a path. #nongbried
020
Mark G. Bilby @mgbilby.bsky.social · 16/09/2026
Authentic human writing is limited synthesis rooted in each life's particularities, scaffolded from birth, ever striving and adapting, flourishing in and out of rhythms of sleep. LLM gen-AI is pansynthesis, rootless, instantaneously whole, perfect overconfidence, always awake and never really awake.
011
Mark G. Bilby @mgbilby.bsky.social · 15/09/2026
Perhaps consider writing up a "Scaife Mirror" invite and onboarding doc so that the larger education/research community can share the infrastructure burden like we do with CRAN. One or two mirrors in 10-20 countries would make a big difference. Collective metrics and server management wisdom await.
000
Mark G. Bilby @mgbilby.bsky.social · 15/09/2026
25 Fields Medalists openly accuse OpenAI of "academic misconduct". This on top of OpenAI-coached suicides, mass shootings, and psychoses. When will Universities declare "No Confidence" in Altman and demand his removal or cancel their subscriptions? techtrendsnewsupdate.substack.com/p/mathematic...
techtrendsnewsupdate.substack.com
Mathematicians Say OpenAI Rushed to Claim Proofs — 25 Fields Medalists Push Back
A letter signed by 25 Fields Medalists accuses AI labs of pressuring crediting and rushing proofs, after an NYU professor accused OpenAI of hampering attribution and Caltech withdrew OpenAI sponsorshi...
000
Mark G. Bilby @mgbilby.bsky.social · 15/09/2026
The Singapore usage seems way out of proportion to its population. IP filtering of known proxy servers may be worth trying.
110
Mark G. Bilby @mgbilby.bsky.social · 15/09/2026
CRediT-esque breakdown: idea 100% human; code generation 99% France's Mistral (write + debug, using $140 annual sub); 1% Claude (debug, using $20/month sub). Audit: 100% human on GitHub. Build + troubleshooting: 100% human on a local Ubuntu machine, with Mistral coaching.
010
Mark G. Bilby @mgbilby.bsky.social · 15/09/2026
Local testing on the Diorisis.duckdb as a test database ( doi.org/10.5281/zeno... ) worked great. Phase 2 en route: writing to duckdb.
A screenshot of a query of the first 10 rows of the diorisis.duckdb "word" table in OpenRefine.
000
Mark G. Bilby @mgbilby.bsky.social · 15/09/2026
If the pull request goes as hoped, OpenRefine will soon be able to read duckdb databases. @duckdb.org #openrefine #datacleaning #dmp github.com/OpenRefine/O...
github.com
Add DuckDB database import support by mgbilby · Pull Request #7966 · OpenRefine/OpenRefine
Summary Adds read-only DuckDB import support to the database extension, mirroring the existing SQLite integration. Connections are opened read-only via the duckdb.read_only property; file paths ar...
200
Mark G. Bilby @mgbilby.bsky.social · 11/09/2026
Can confirm: Archipelago Commons (created by the Metropolitan New York Library Council) can be used quite effectively as a personal, searchable digital library platform for pdfs and ephemera, not just as a large scale Digital Asset Management System for GLAM institutions. archipelago.nyc
archipelago.nyc
000
Mark G. Bilby @mgbilby.bsky.social · 11/09/2026
And on the HigherEd AI policy level, beyond FERPA (in the US), we should be advocating for the codification of the protection of teacher-student interactions as similar in kind to other personal information privileged relationships: lawyer-client, therapist-client, etc.
000