Sign in

Nicolay Gerold

@nicolaygerold.com
741 followers 145 following 643 posts

Daytime: building ai systems & platforms @ Aisbach Nighttime: hacking on generative ai Host of How AI Is Built

PostsRepliesMedia
Nicolay Gerold @nicolaygerold.com · 16/05/2025
OpenCode is really stepping it up. I love the sidebar stuff. Really plays well with nvim.
000
Nicolay Gerold @nicolaygerold.com · 25/04/2025
Thanks buddy
000
Nicolay Gerold @nicolaygerold.com · 13/02/2025
EU mobilizes 200 billion Euros for AI. Unless we have a massive political change that money will just go to waste. For innovation to happen, we don't need money first, but deregulation. We cannot work on breakthrough technologies, when we are at constant fear of being sued.
120
Nicolay Gerold @nicolaygerold.com · 13/02/2025
RAG is dead. Long live RAG. LLMs suck at long context. This paper shows what I have seen in most deployments. With longer contexts, performance degrades.
161
Nicolay Gerold @nicolaygerold.com · 31/01/2025
Dropping some new episodes on @howaiisbuilt.fm . Links below.
111
Nicolay Gerold @nicolaygerold.com · 25/01/2025
New podcast with Alex Garcia on search in sqlite
120
Nicolay Gerold @nicolaygerold.com · 18/01/2025
That's surprisingly on brand.
000
Nicolay Gerold @nicolaygerold.com · 09/01/2025
Developers treat search as a blackbox. Throw everything in a vector database and hope something good comes out. Throw all ranking signals into one big ML model and hope it makes something good out of it. You don’t want to create this witch’s cauldron. New episode on @howaiisbuilt.fm
141
Nicolay Gerold @nicolaygerold.com · 03/01/2025
The biggest lie in RAG is that semantic search is simple. The reality is that it's easy to build, it's easy to get up and running, but it's really hard to get right. And if you don't have a good setup, it's near impossible to debug. One of the reasons it's really hard is chunking.
121
Nicolay Gerold @nicolaygerold.com · 23/12/2024
@merve.bsky.social Can you show the Amazon people how to use a VLM to do the handwriting recognition. That's atrocious.
010
Nicolay Gerold @nicolaygerold.com · 23/12/2024
Getting ads for the remarkable pro while watching a review on the new Kindle Scribe. Someone got their targeting down.
000
Nicolay Gerold @nicolaygerold.com · 21/12/2024
You usually have supervisors and workers. And it is super easy to spin them up. And the workers run in "lightweight processes", which when they crash they can be spun up again super fast and since they are isolated, they don't bring down the entire system.
010
Nicolay Gerold @nicolaygerold.com · 19/12/2024
They use a large model (e.g. gpt-4o) to generate training data for a smaller one (gpt-4o-mini). This lets you build fast, cheap models that do one thing well or that are more capable because they have (nearly) identical capabilities distilled into a smaller number of parameters.
110
Nicolay Gerold @nicolaygerold.com · 17/12/2024
Didn't find the on demand pricing for it. Probably around 500/h if it becomes available for it :D
010
Nicolay Gerold @nicolaygerold.com · 17/12/2024
@chris.blue Getting reminded to buy myself a Christmas gift.
021
Nicolay Gerold @nicolaygerold.com · 16/12/2024
If you are lost in all the fuzz around the Byte Latent Transformer by Meta, read on. Meta has created BLT, a new AI model that works with raw bytes instead of tokens. Current AI models split text into tokens (fixed chunks of letters) before processing it. >>
110
Nicolay Gerold @nicolaygerold.com · 16/12/2024
Claude is surprisingly good at workout programming.
000
Nicolay Gerold @nicolaygerold.com · 15/12/2024
000
Nicolay Gerold @nicolaygerold.com · 14/12/2024
"Instead of being a one-way pipeline, agentic RAG allows you to check, 'Am I actually answering the user's question?'" Different questions need different approaches. ➡️ 𝗤𝘂𝗲𝗿𝘆-𝗕𝗮𝘀𝗲𝗱 𝗙𝗹𝗲𝘅𝗶𝗯𝗶𝗹𝗶𝘁𝘆: - Structured data? Use SQL - Context-rich query? Use vector search - Date-specific? Apply filters first
251
Nicolay Gerold @nicolaygerold.com · 14/12/2024
Damn....
110
Nicolay Gerold @nicolaygerold.com · 14/12/2024
Inequality joins in polars is massive.
130
Nicolay Gerold @nicolaygerold.com · 11/12/2024
Has been a while since, I have done gradient accumulation. I always tend to forget the last step (the check for hitting the length of the dataset) on the first implementation.
010
Nicolay Gerold @nicolaygerold.com · 10/12/2024
Coding a project > reading articles. Coding a project forces you to apply concepts directly. It’s a richer learning experience than just reading technical articles. You discover gaps, solve real problems, and solidify your understanding.
230
Nicolay Gerold @nicolaygerold.com · 10/12/2024
15Mio Ai Builders on Huggingface, we are still early.
000
Nicolay Gerold @nicolaygerold.com · 10/12/2024
Data Drift for Dummies. Data drift happens when the real world changes but your model doesn't. - Input drift: The data coming in changes (like cameras getting better resolution) - Label drift: What you're predicting changes (like what counts as "spam" evolving) >>
100
Nicolay Gerold @nicolaygerold.com · 08/12/2024
This might take a while
000
Nicolay Gerold @nicolaygerold.com · 05/12/2024
Many companies use ElasticSearch or OpenSearch and use 10% of the capacity. On top, they have to build ETL pipelines. Get data normalized. Worry about race conditions. All in all, when you want to do search on top of your existing database, you are forced to build distributed systems. #ai
122
Nicolay Gerold @nicolaygerold.com · 04/12/2024
I think it's an honor to be on the Huggingface list! 🤗
010
Nicolay Gerold @nicolaygerold.com · 26/11/2024
What's the architecture pattern that resolved most of your problems? For me: event sourcing.
Diagram for event sourcing: the client application writes events to a Kafka topic, whose events are consumed by a service, which performs the updates (e.g. insertions into primary databases or cache invalidations).
000
Nicolay Gerold @nicolaygerold.com · 25/11/2024
New Huggingface LLM Observability. Stored in Argilla or Datasets. I am missing some features, but seems better than OpenAI's alternative. github.com/cfahlgren1/o...
1101
Nicolay Gerold @nicolaygerold.com · 21/11/2024
Got the good stuff
020
Nicolay Gerold @nicolaygerold.com · 21/11/2024
With RAG these issues are amplified. We do not look at full documents anymore, but at bits and pieces. So we have to be extra careful. Today on @howaiisbuilt.fm we talk to Max Buckley. Max works at Google and has built a lot of interesting stuff with LLMs to improve knowledge bases for RAG. >>
101
Nicolay Gerold @nicolaygerold.com · 18/11/2024
Here we go again. Double trouble on IKEA. Last week on Founders today on Acquired.
010
Nicolay Gerold @nicolaygerold.com · 17/11/2024
Sunday hobby: slow cooking stuff to see how far salt, pepper and herbs can push it.
040
Nicolay Gerold @nicolaygerold.com · 15/11/2024
You probably guessed it, we are talking about the OG ranking function in search: BM25. Today we are back continuing our series on search on @howaiisbuilt.fm with @taidesu.bsky.social. We talk about BM25, how it works, what makes it great and how you can tailor it to your use-case.
121
Nicolay Gerold @nicolaygerold.com · 15/11/2024
Some query types might not work at all. It is very costly in terms of storage and compute. We have to keep our indexes in memory to achieve a low enough latency for search. What we are talking about today works for everything, works out of domain, and is one of the most efficient. >>
121
Nicolay Gerold @nicolaygerold.com · 15/11/2024
People implementing RAG jump straight into vector search. But vector search has a lot of downsides. Vector search is not robust out of domain. Different types of queries need different embedding models with different vector indices. >>
142
Nicolay Gerold @nicolaygerold.com · 12/11/2024
With what you know about me can you describe my personality? Not sure why, but I seem to like YAML over TOML.
010
Nicolay Gerold @nicolaygerold.com · 09/11/2024
"Sadly, it's a bit off a snake oil. These long context embedding models have tested basically all of them, not really working well. So it's [best length of chunks] something between like 500 and 1,000 tokens." Text embeddings are far from perfect. They struggle with long documents. >>
163
Nicolay Gerold @nicolaygerold.com · 09/11/2024
I need my 5 screens. Apparently I really love monorepos.
000
Nicolay Gerold @nicolaygerold.com · 07/11/2024
Today on @howaiisbuilt.fm we are talking to Charles Xie, the founder and CEO of Zilliz, the company behind Milvus. Charles previously worked at Oracle as a founding engineer of the 12c cloud database.
020
Nicolay Gerold @nicolaygerold.com · 05/11/2024
“There is no free lunch.” Every performance optimization comes with tradeoffs in either functionality, flexibility, or cost. When building search systems, there's a seductive idea that we can optimize everything: fast results, high relevancy, and low costs. But that’s not the reality.
Podcast Thumbnail with Title: Search Systems at Scale: Avoiding Local Maxima and Other Engineering Lessons
511
Nicolay Gerold @nicolaygerold.com · 02/11/2024
SLMs might be on the brink of breakthrough. SmolLM2 is crazy impressive.
000