Sign in

Ian Maurer

@imaurer.bsky.social
156 followers 227 following 45 posts

Fighting cancer with code. Interests: Bioinformatics, Genomics, LLMs, Python, Rust, NLP. imaurer.com github.com/imaurer

PostsRepliesMedia
Ian Maurer @imaurer.bsky.social · 03/11/2025
I'm only here because @simonwillison.net posted about this on elon's plaything.
000
Ian Maurer @imaurer.bsky.social · 08/04/2025
Built on top of the powerful open access API services of the NIH: - ClinicalTrials.gov - MyVariant.info - PubMed / Pubtator3 www.ncbi.nlm.nih.gov/research/pub...
000
Ian Maurer @imaurer.bsky.social · 08/04/2025
Short video on YouTube: www.youtube.com/watch?v=bKxO...
youtube.com
Introduction to BioMCP
YouTube video by GenomOncology
100
Ian Maurer @imaurer.bsky.social · 08/04/2025
Github Repo (MIT License): github.com/genomoncolog...
github.com
GitHub - genomoncology/biomcp: BioMCP: Biomedical Model Context Protocol
BioMCP: Biomedical Model Context Protocol. Contribute to genomoncology/biomcp development by creating an account on GitHub.
100
Ian Maurer @imaurer.bsky.social · 08/04/2025
Proud to announce BioMCP, an open source Model Context Protocol (MCP) server for biomedical research AI assistants and agents. BioMCP searches and retrieves clinical trials, PubMed articles and genomic variants to provide up-to-date and relevant context to an LLM.
120
Ian Maurer @imaurer.bsky.social · 14/01/2025
slick!
010
Ian Maurer @imaurer.bsky.social · 14/01/2025
Listening now... came here to say congrats. :)
110
Ian Maurer @imaurer.bsky.social · 04/01/2025
I did finish it and listened to your podcast with @mkennedy.codes Good stuff. Subscribed on overcast.
010
Ian Maurer @imaurer.bsky.social · 03/01/2025
Great article this morning. Got to the blog scroll link and… now I am here. I hope I don’t get distracted too much longer. I want to finish it.
000
Ian Maurer @imaurer.bsky.social · 01/01/2025
Your line colors are messing with my mind.
050
Ian Maurer @imaurer.bsky.social · 22/12/2024
You can’t for-loop taste
020
Reposted by Ian Maurer
Ian Maurer @imaurer.bsky.social · 20/12/2024
Just everything else with AI, you can get good results with low effort. Then you can grind and keep eeking out gains. This has add’l benefits: - improvement on that task - greater vision of what else you can do - higher expectations on what good results are - better first drafts the next time.
031
Ian Maurer @imaurer.bsky.social · 20/12/2024
Just everything else with AI, you can get good results with low effort. Then you can grind and keep eeking out gains. This has add’l benefits: - improvement on that task - greater vision of what else you can do - higher expectations on what good results are - better first drafts the next time.
031
Ian Maurer @imaurer.bsky.social · 19/12/2024
How are they going to get the sun to shine like that in winter?
000
Ian Maurer @imaurer.bsky.social · 16/12/2024
1. No 2. NaN
000
Ian Maurer @imaurer.bsky.social · 14/12/2024
PDF version arxiv.org/pdf/2408.02442
arxiv.org
000
Ian Maurer @imaurer.bsky.social · 14/12/2024
Perplexity did find me this quote: “Upon inspection, we found that 100% of GPT 3.5 Turbo JSON-mode responses placed the "answer" key before the "reason" key, resulting in zero-shot direct answering instead of zero-shot chain-of-thought reasoning.” ar5iv.org/html/2408.02...
ar5iv.org
Let Me Speak Freely? A Study on the Impact of Format Restrictions on Performance of Large Language Models
Structured generation, the process of producing content in standardized formats like JSON and XML, is widely utilized in real-world applications to extract key output information from large language m...
100
Ian Maurer @imaurer.bsky.social · 14/12/2024
Any research on the impact of JSON key sort order on generating structured outputs? For instance, I do something like this: - analyze - draft - evaluate - final And then provide exemplars with errors to show self correction. But I haven’t done extensive testing. Just seems like a good idea.
100
Ian Maurer @imaurer.bsky.social · 13/12/2024
Congrats!
010
Ian Maurer @imaurer.bsky.social · 11/12/2024
Ugh. Sorry to hear this.
020
Ian Maurer @imaurer.bsky.social · 10/12/2024
Among Us lol.
000
Ian Maurer @imaurer.bsky.social · 10/12/2024
That's my favorite fun fact about SQLite. Seems to be a competitive advantage for them securing contracts from large organizations. Which, good on them, I think it's a more than fair way to do OSS.
2190
Ian Maurer @imaurer.bsky.social · 08/12/2024
That’d be sweet. Thank Claude for me.
010
Ian Maurer @imaurer.bsky.social · 08/12/2024
Bunny CDN still treating you alright?
100
Ian Maurer @imaurer.bsky.social · 08/12/2024
Nice looking script.
000
Ian Maurer @imaurer.bsky.social · 08/12/2024
You are right bytes not tokens.
110
Ian Maurer @imaurer.bsky.social · 08/12/2024
Git ingest: gitingest.com
gitingest.com
Git ingest
Replace 'hub' with 'ingest' in any Github Url for a prompt-friendly text
171
Ian Maurer @imaurer.bsky.social · 08/12/2024
Chat with any open source repo easily. Gitingest (free online tool) turns any GitHub repository into a single markdown file for pasting. Claude artifacts makes this 300k token output pretty easy to work with.
46412
Ian Maurer @imaurer.bsky.social · 05/12/2024
Makes sense. Thanks for the clarification. I was wondering if it was a copyright thing which didn’t make much sense given the technology.
000
Ian Maurer @imaurer.bsky.social · 05/12/2024
Great thoughts. AT protocol is interesting for the same reason podcasts are interesting. Love the open web. Can you explain the legalities of why you can’t have a play button for some podcasts that have a public mp3 url?
120
Ian Maurer @imaurer.bsky.social · 04/12/2024
What a great conversation. I really appreciate the historical perspective and the clear expression of tradeoffs and patterns that have evolved over time. Thanks to you both. Cheers.
000
Ian Maurer @imaurer.bsky.social · 04/12/2024
I have to say I am impressed that the Databricks team didn't leave out their own model (DBRX) from the results even though it doesn't exactly shine in this analysis.
010
Ian Maurer @imaurer.bsky.social · 04/12/2024
Nice paper from Databricks Mosaic Research that analyzes RAG performance with long-context LLMs. Lots of great data and insights about the leading models and how their performance improves or degrades as more content is added to the context window. ArXiv Paper: arxiv.org/pdf/2411.03538
191
Ian Maurer @imaurer.bsky.social · 02/12/2024
This is great. If you don't mind me asking, how did you collect the sources?
100
Ian Maurer @imaurer.bsky.social · 02/12/2024
This is great. I use instructor to handle some of the subtle api request differences. Otherwise every other library seems too much bloat for my taste. Did you dig into variations between providers? Also people interested in this might want to check out: github.com/genomoncolog...
github.com
GitHub - genomoncology/FuzzTypes: Pydantic extension for annotating autocorrecting fields.
Pydantic extension for annotating autocorrecting fields. - genomoncology/FuzzTypes
041
Ian Maurer @imaurer.bsky.social · 27/11/2024
Of course and thanks for the great SpaCy docs. Helped me out tremendously back in 2018 when I didn’t know how to spell NLP.
010
Ian Maurer @imaurer.bsky.social · 25/11/2024
Bluesky support for the open web makes me both very happy and bullish on bsky in general.
010
Ian Maurer @imaurer.bsky.social · 22/11/2024
I'm tracking the language models, libraries, and resources for generating structured outputs (JSON) on GitHub (2K stars). PRs welcome! Repo: github.com/imaurer/awes...
Awesome LLM JSON List
This awesome list is dedicated to resources for using Large Language Models (LLMs) to generate JSON or other structured outputs.

Table of Contents
Terminology
Hosted Models
Local Models
Python Libraries
Blog Articles
Videos
Jupyter Notebooks
Leaderboards
030
Ian Maurer @imaurer.bsky.social · 19/11/2024
Great presentation last night. Look forward to reading this to really understand the flow of the architecture. I have a lot of the same components just hooked up slightly differently.
110
Ian Maurer @imaurer.bsky.social · 19/11/2024
Yo dawg: www.postgresql.org/docs/current...
postgresql.org
F.36. postgres_fdw — access data stored in external PostgreSQL servers
F.36. postgres_fdw — access data stored in external PostgreSQL servers # F.36.1. FDW Options of postgres_fdw F.36.2. Functions F.36.3. Connection Management …
010
Ian Maurer @imaurer.bsky.social · 19/11/2024
GLiREL is a zero-shot Relation Extraction (RE) model capable of classifying unseen relations given the entities within a text. Builds upon GLiNER which does 0-shot Named Entity Recognition (NER). Repo: github.com/jackboyla/GL... Model: huggingface.co/jackboyla/gl... h/t @pacoid.bsky.social
1152
Ian Maurer @imaurer.bsky.social · 17/11/2024
Not seeing the benefit over using an LLM to 1) write your test then 2) write your code But maybe I am missing something.
000
Ian Maurer @imaurer.bsky.social · 17/11/2024
Found it on Overcast and downloaded. Look forward to listening. Good luck
000
Ian Maurer @imaurer.bsky.social · 30/10/2024
Thanks for sharing
010
Ian Maurer @imaurer.bsky.social · 30/10/2024
Bought a 1.5 TB microsd for $100/86 chf. Way slower but not terrible for datasets and local LLMs. I am pretty sure I am not getting a Mac the next time I buy my own laptop
000
Ian Maurer @imaurer.bsky.social · 30/10/2024
App feels smoother too. This is my first reply. Hopefully a few more key people move over and I can delete X.
010