Sign in

Tom Hipwell

@tomhipwell.co
52 followers 211 following 59 posts

VP Engineering at nory.ai. Past roles: Deel/Hofy, Bulb, JPMorgan. Learning. Shipping.

PostsRepliesMedia
Tom Hipwell @tomhipwell.co · 09/05/2025
I've been meaning to write a blog post a bit like this one for ages, but this is much better than I would have done. I think the author is right, private evals are very important and this post gives a good framework of how to design your own -> thundergolfer.com/blog/private...
thundergolfer.com
You should have private evals
Everybody should have a personal set of test prompts to try on LLMs.
000
Tom Hipwell @tomhipwell.co · 13/04/2025
Always really like Alex's post and I feel like this one nails it again: "the dynamic that matters is not a timeline of capacities but a timeline of accuracies" if you're not already maxed out with AI predictions then this one is worth a read
020
Tom Hipwell @tomhipwell.co · 11/04/2025
Firebase studio looks quite fun butttt: "To block the use of your prompts and responses for model training, do not use the App Prototyping agent, and do not use Gemini in Firebase within Firebase Studio. To block the use of your code for model training, turn off code completion and code indexing..."
000
Tom Hipwell @tomhipwell.co · 23/03/2025
Blog post to try and tie together a few themes I've been reading about this week -> tomhipwell.co/blog/cursor_...
tomhipwell.co
Cursor rules, prompt injections, voice to text and Diane | Tom Hipwell
Learning in the open | Tom Hipwell
011
Tom Hipwell @tomhipwell.co · 12/03/2025
Not sure if this is obvious but probably the most important evaluation criteria for any AI tool at the moment is the ability to choose your own model (and bring your own API key or OpenAI compatible API)
000
Tom Hipwell @tomhipwell.co · 02/03/2025
This is a great post, worth reading. I re-blogged it here to try and summarise the reasoning -> tomhipwell.co/blog/the_mod...
tomhipwell.co
The Model is the Product | Tom Hipwell
Learning in the open | Tom Hipwell
010
Reposted by Tom Hipwell
Sung Kim @sungkim.bsky.social · 24/02/2025
Claude 3.7 Sonnet
23614
Tom Hipwell @tomhipwell.co · 21/02/2025
Product teams talking too much about ICPs is a red flag. ICPs are for sales and marketing teams. They need narrow focus to max win rate. Product teams need to understand ICP for prioritisation, but think in PMF strength across segments. Peripheral vision. This is how you expand PMF and win more.
000
Tom Hipwell @tomhipwell.co · 17/02/2025
Another day, another AI dev flow. There’s some common patterns emerging now (using markdown files like spec.md etc.). This blog gives a step by step guide and prompts to borrow. The advice reduces to “spend a lot of time planning with reasoning models up front” -> harper.blog/2025/02/16/m...
harper.blog
My LLM codegen workflow atm
A detailed walkthrough of my current workflow for using LLms to build software, from brainstorming through planning and execution.
000
Reposted by Tom Hipwell
Gergely Orosz @gergely.pragmaticengineer.com · 07/02/2025
Will GenAI mean the end of software engineering? Really thoughtful take from @chiphuyen.bsky.social in The Pragmatic Engineer Podcast: Perhaps software engineering will change, like writing changed hundreds of years ago thanks to printing Full: www.youtube.com/watch?v=98o_...
8664
Tom Hipwell @tomhipwell.co · 29/01/2025
Really enjoyed Dario's blog post, I thought it had a bunch of interesting, verifiable predictions (I'm less interested in the export control stuff). It'll be interesting to see if they come good over the course of '25. Worth a read, cuts through the noise -> darioamodei.com/on-deepseek-...
darioamodei.com
Dario Amodei — On DeepSeek and Export Controls
On DeepSeek and Export Controls
000
Tom Hipwell @tomhipwell.co · 02/01/2025
Great read, useful survey of the field.
120
Tom Hipwell @tomhipwell.co · 02/01/2025
Good list, would be good to hear additions from #dataBS, I would add "what are embeddings" by @vickiboykis.com -> www.latent.space/p/2025-papers
latent.space
The 2025 AI Engineering Reading List
We picked 50 paper/models/blogs across 10 fields in AI Eng: LLMs, Benchmarks, Prompting, RAG, Agents, CodeGen, Vision, Voice, Diffusion, Finetuning. If you're starting from scratch, start here.
041
Tom Hipwell @tomhipwell.co · 31/12/2024
I like @simonwillison.net definition of slop, but we're missing a word to describe an algorithm that pushes low grade content. I'd propose gruel, e.g my LinkedIn feed is all gruel now. Netflix recommends gruel all the time. Spotify's playlists are stuffed full of gruel. You get the picture.
010
Tom Hipwell @tomhipwell.co · 24/12/2024
If you can solve this puzzle, then do not despair! You are still much smarter than o3
100
Tom Hipwell @tomhipwell.co · 18/12/2024
Great post, a starting point for something that has been confusing for a while
010
Tom Hipwell @tomhipwell.co · 17/12/2024
webdev arena is a _really_ fast way to get a feel for the coding abilities of the different models out there. Worth five minutes of your time to design a tricky prompt, then quickly assess each model generation as they get released -> web.lmarena.ai
web.lmarena.ai
WebDev Arena
WebDev Arena: AI Battle to build the best website
000
Tom Hipwell @tomhipwell.co · 14/12/2024
Interesting ideas
000
Tom Hipwell @tomhipwell.co · 10/12/2024
I've had this post sat in my drafts for well over 6 months now, but with the release of Sora to GA yesterday I thought I'd share it - Sora: An Idiot's Guide. tomhipwell.co/blog/sora/
tomhipwell.co
Sora: An idiot's guide | Tom Hipwell
Just what is latent space anyway?
200
Reposted by Tom Hipwell
Sung Kim @sungkim.bsky.social · 09/12/2024
Reinforcement Learning: An Overview This manuscript gives a big-picture, up-to-date overview of the field of (deep) reinforcement learning and sequential decision making, covering value-based RL, policy-gradient methods, model-based methods, and various other topics. arxiv.org/abs/2412.05265
0548
Tom Hipwell @tomhipwell.co · 08/12/2024
I had a play about with genini-1206 last night, trying some prompts into both o1 and 1206 in parallel, both models are crackers. Vibes for me are better with o1, feels a bit more succinct and gets to the point quicker. Early days of course.
100
Reposted by Tom Hipwell
Sung Kim @sungkim.bsky.social · 07/12/2024
Alibaba Qwen team just released the base models for Qwen2-VL. You can also wait for them to release Qwen2.5-VL, which should be sooner than later. 2B: huggingface.co/Qwen/Qwen2-V... 7B: huggingface.co/Qwen/Qwen2-V... 72B: huggingface.co/Qwen/Qwen2-V...
huggingface.co
Qwen/Qwen2-VL-2B · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
1132
Reposted by Tom Hipwell
Ethan Mollick @emollick.bsky.social · 07/12/2024
A test of how seriously your firm is taking AI: when o-1 (& the new Gemini model) came out this week, were there assigned folks who immediately ran the model through your internal, validated, firm-specific benchmarks to see how useful it as? Did you update any plans or goals as a result?
1119426
Tom Hipwell @tomhipwell.co · 06/12/2024
Some great ideas here, worth your time to explore
151
Reposted by Tom Hipwell
Chris @chris.blue · 05/12/2024
First new post in a couple of weeks! There's been a lot of activity around regattastorage.com this week, so I decided to write about the space. tl;dr It's pretty exciting!
materializedview.io
The Quest for a Distributed POSIX-Compatible Filesystem
Distributed POSIX filesystems have proven elusive, but we're getting closer. Perhaps that's all we need.
2209
Tom Hipwell @tomhipwell.co · 05/12/2024
This is nuts 🥜🐿️🐿️🤯
000
Tom Hipwell @tomhipwell.co · 04/12/2024
Insightful piece from @marcbrooker.bsky.social on Aurora DSQL. Aurora is a "distributed sql" db with SQL and ACID, global active-active, scalability both up and down (independent scaling of compute, reads, writes, and storage), and Postgres compatibility (psql works!)-> brooker.co.za/blog/2024/12...
brooker.co.za
DSQL Vignette: Aurora DSQL, and A Personal Story - Marc's Blog
000
Tom Hipwell @tomhipwell.co · 03/12/2024
Super handy index of publicly available LLM implementations / reference architectures, all searchable in one place, thanks to @strickvl.bsky.social -> zenml.io/llmops-datab...
zenml.io
ZenML - LLMOps Database
011
Tom Hipwell @tomhipwell.co · 02/12/2024
I wrote a fast guide on getting up and running quickly (an hour or so) and cheaply (free!) with semantic search for toy use cases (blogs etc.). I used chromadb, which is an Apache 2.0 licensed sqllite wrapper; all-MiniLM-L6-v2 is used for embedding generation -> tomhipwell.co/blog/baked_s...
tomhipwell.co
Baked Search: Building semantic search quickly for toy use cases | Tom Hipwell
End to end tutorial of building a free search backend.
000
Reposted by Tom Hipwell
Ted Underwood @tedunderwood.com · 01/12/2024
Did you know that attention across the whole input span was inspired by the time-negating alien language in Arrival? Crazy anecdote from the latest Hard Fork podcast (by @kevinroose.com and @caseynewton.bsky.social). HT nwbrownboi on Threads for the lead.
Transcript of Hard Fork ep 111: Yeah. And I could talk for an hour about transformers and why they are so important.
But I think it's important to say that they were inspired by the alien language in the film Arrival, which had just recently come out.
And a group of researchers at Google, one researcher in particular, who was part of that original team, was inspired by watching Arrival and seeing that the aliens in the movie had this language which represented entire sentences with a single symbol. And they thought, hey, what if we did that inside of a neural network? So rather than processing all of the inputs that you would give to one of these systems one word at a time, you could have this thing called an attention mechanism, which paid attention to all of it simultaneously.
That would allow you to process much more information much faster. And that insight sparked the creation of the transformer, which led to all the stuff we see in Al today.
1924653
Reposted by Tom Hipwell
vb @reach-vb.hf.co · 28/11/2024
To showcase how much you can do with just a 1.7B LLM, you pass free text, define a schema of parsing the text into a GitHub issue (title, description, categories, tags, etc) - Let MLC & XGrammar do the rest! That's it, the code is super readable, try it out today! 🤗 huggingface.co/spaces/reach...
huggingface.co
Github Issue Generator - a Hugging Face Space by reach-vb
Discover amazing ML apps made by the community
1172
Tom Hipwell @tomhipwell.co · 30/11/2024
This S1 analysis from Meritech went viral due to the (compounding!) IPO ratchet that ServiceTitan are subject to. About halfway down there are some handy benchmarks for median/top decile pre-IPO performance in vertical SaaS -> www.meritechcapital.com/blog/service...
meritechcapital.com
ServiceTitan S-1 Breakdown ‒ Meritech Capital
We help build market-leading companies in the markets that matter.
000
Tom Hipwell @tomhipwell.co · 30/11/2024
End to end tutorial of function calling with Llama-3.2-3B-Instruct, building gradually from string templating, to using Jinja, to implementing web search with Brave, via @sergiopaniego.bsky.social -> github.com/huggingface/...
github.com
huggingface-llama-recipes/tool_calling/tool_calling.ipynb at main · huggingface/huggingface-llama-recipes
Contribute to huggingface/huggingface-llama-recipes development by creating an account on GitHub.
000
Reposted by Tom Hipwell
Brett Adcock @adcock.bsky.social · 29/11/2024
Sharing a timelapse of a small fleet of Figure 02 humanoid robots doing work Our robots are reasoning fully autonomously Much of the core engineering infrastructure has been laid the last two years, 2025 is going to be 🤯
52514
Reposted by Tom Hipwell
Eugene Yan @eugeneyan.com · 27/11/2024
Feels good to be mentioned on HN for engineers learning AI 🥰 Helping others is a big reason I write. Here's a list on ML/AI: ## Building AI systems • Patterns for Building LLM-based Systems: eugeneyan.com/writing/llm-... • What We’ve Learned From A Year of Building with LLMs: applied-llms.org
Read through this making flashcards as you to: https://eugeneyan.com/writing/llm-patterns/
Then spin up a RAG-enhanced chatbot using pgvector on your favourite subject, and keep improving it when you learn about cool techniques

---

Lots of people can get impressive demos up and running, but if you want to run AI products in production, you're going to have to do system evals. System evals make sure your product is doing what it says on the box with unquantifiable qualities.
We wrote a zine on system evals without jargon: https://forestfriends.tech
Eugene Yan has written extensively on it https://eugeneyan.com/writing/evals/
Hamel has as well. https://hamel.dev/blog/posts/evals/
3709
Tom Hipwell @tomhipwell.co · 26/11/2024
uv cheat sheet, handy intro to the tool if, like me, you've not tried it yet-> docs.google.com/document/d/1... via @brandonrohrer.com
docs.google.com
uv cheatsheet
uv cheatsheet https://docs.astral.sh/uv/ Install on macOS and linux: curl -LsSf https://astral.sh/uv/install.sh | sh old way new way pyenv install 3.12 pyenv versions uv python install 3.12 uv pytho...
110
Tom Hipwell @tomhipwell.co · 26/11/2024
Bought a copy of "The Little Book" and I've been keeping it on my desk for quick reference, highly recommended!
010
Reposted by Tom Hipwell
Orion Weller @orionweller.bsky.social · 15/09/2023
Using LLMs for query or document expansion in retrieval (e.g. HyDE and Doc2Query) have scores going 📈 But do these approaches work for all IR models and for different types of distribution shifts? Turns out its actually more 📉 🚨 📝 (arxiv soon): orionweller.github.io/assets/pdf/L...
A plot: the x axis is baseline score of rankers, in ndcg@10. y axis is delta of model score after an expansion is applied.

There are three sets of results, one dataset for each shift type: TrecDL (no shift), FiQA (domain shift), ArguAna (query shift).  For each set of result, the chart shows a scatter plot with a trend line. We observe the same trend for all: as the baseline score increases, the delta when using expansion decreases. 

On TREC DL, worst models have a base score of ~40, and improve by 10 points w/expansion. the best models have a score of >70, and their performance decreases by -5 points w/expansion.

On FiQA, worse models have a base score of ~15, and improve by 5 points w/expansion. the best models have a score of ~45, and their performance decreases by -3 point w/expansion.

On ArguAna, worst models have a base score of ~25, and improve by >20 points w/expansion. the best models have a score of >55, and their performance decreases by -1 point w/expansion.
3436
Tom Hipwell @tomhipwell.co · 21/11/2024
The AWS Lambda PR/FAQ has lots of small, golden nuggets. PR/FAQs like this are a bit wordy but you can see how effectively the technique is used here to explain the impact of not just a new architecture but also a shift in AWS’s compute billing model - www.allthingsdistributed.com/2024/11/aws-...
allthingsdistributed.com
AWS Lambda turns 10: A rare look at the doc that started it
On AWS Lambda's 10th anniversary, I'm publishing the internal PR/FAQ that helped launch this groundbreaking service. This document provides insight into the customer problems we observed in the early ...
000