Philipp Krenn @xeraa.net · 15/09/20266,000,000,000 downloads of elasticsearch, kibana, logstash, beats, elastic agent to everyone who downloaded it, built on it, contributed to it, deployed it, broke it, fixed it, and made something amazing with it: thank you 000
Philipp Krenn @xeraa.net · 03/09/2026after all the attention for duckdb's acquisition by AWS: one of my favorite things is their "why @duckdb.org" page — duckdb.org/why_duckdb explainining their tradeoffs and goals; and how they try to achieve those through technical means 091
Philipp Krenn @xeraa.net · 02/09/2026full article with examples how not to do it in various programming languages: bonsai.io/blog/float-b... 3/3bonsai.ioFloat Bloat: vector serialization gone wrongYour embeddings are float32. Somewhere in your app they get promoted to float64 and grow a tail of digits that mean nothing. You pay to store, ship, and parse every one of them. 000
Philipp Krenn @xeraa.net · 02/09/2026* `semantic_text generates` the embeddings server-side and has always avoided the problem and for even more optimized ingestion, use base64 encoded strings: www.elastic.co/search-labs/... 2/3elastic.coUsing Base64-encoded strings to speed up vector ingestionLearn about Base64-encoded strings and the improvements it brings to vector ingestion in Elasticsearch. 100
Philipp Krenn @xeraa.net · 02/09/2026"Float Bloat: vector serialization gone wrong" by storing float64 instead of float32. it's a surprising problem but with elasticsearch you're in luck: * since 9.2 (index creation) `index.mapping.exclude_source_vectors` is on by default — you don't store the floats in source at all any more 1/3 100
Philipp Krenn @xeraa.net · 23/08/2026static.klipy.comYoda's Pain, Suffering, and Death MemeALT: Yoda's Pain, Suffering, and Death Meme 000
Philipp Krenn @xeraa.net · 23/08/2026hacking a tablet that was EOLed by amazon: "American frontier models won’t help and Chinese will, but not without reasoning about whether they should" the barrier is to pick the right model, say "keep going" over and over, and paying the token costs: ericpardee.github.io/fire-hd-owne...ericpardee.github.ioAmazon kept shutting down my tablet, so I spent $266 on four AI models to own itOwning a tablet Amazon kept shutting down: CVE-2022-38181, four AI models, five months 000
Philipp Krenn @xeraa.net · 20/08/2026hacktoberfest is back: hacktoberfest.com but not as open source contributions but instead as a series of events; maybe not too surprising now that MLH is the main driver sad to see for open source in general. but probably the saner version for maintainers #sloptoberfesthacktoberfest.comHacktoberfest 2026 | AI belongs to everyone300+ in-person Fests plus a global online event, all about building with open source AI. Join a Fest near you this October. 000
Philipp Krenn @xeraa.net · 19/08/2026first hack night in the new @elastic.co san francisco office together with mastra tomorrow: build and compete on prices with agent memory already at 250 RSVP but thanks to the new space we can squeeze in a few more: luma.com/mastra-hack 🤞 for tomorrow 000
Philipp Krenn @xeraa.net · 10/08/2026current state of job offers — without wanting to imply anything 😂 unless it's a hack to get all the attention for this job 100
Philipp Krenn @xeraa.net · 17/07/2026we caught a fake coding interview to steal your developer credentials on the @elastic.co community slack. be extra careful with anyone DMing you files to download and run full analysis: www.elastic.co/security-lab... 010
Philipp Krenn @xeraa.net · 22/06/2026are you in the weights? it's the new klout also works for products (though reluctantly) #elasticsearch intheweights.com 100
Philipp Krenn @xeraa.net · 14/04/2026more cursor skills in the wild — here and everywhere else. spin up an elasticsearch cluster, manage kibana, configure security, set up OTel,... and almost hidden under it is also the brand new MCP endpoint for the @elastic.co docs cursor.com/marketplace/... 000
Philipp Krenn @xeraa.net · 09/04/2026> BERT-based learned sparse and multi-vector dense retrievers generalise better than LLM-based single-vector dense retrievers; and re-ranking remains highly effective BM25 just keeps hanging on with lots more in this new paper: arxiv.org/pdf/2602.21456arxiv.org 000
Philipp Krenn @xeraa.net · 09/04/2026BM25: "why won't you die?!" > the lexical retriever BM25 with appropriate setup outperforms neural rankers in most cases; notably, gpt-oss-20b with BM25 on the passage corpus achieves the highest answer accuracy across all retrieval settings in our study; 100
Philipp Krenn @xeraa.net · 02/04/2026is it an API key for a cloud account or the cluster? way too common of a confusion. so why not both?! released today for serverless projects and your @elastic.co cloud account PS: please be extra careful in your day to day use with these. they have an extra wide blast radius 💥 000
Philipp Krenn @xeraa.net · 31/03/2026someone woke up at RSA and chose violence: vibecoded.vc/cooked/ 😂 though it is entertaining (and I won't agree on all of the takes ;) ) 000
Philipp Krenn @xeraa.net · 16/03/2026powered by CAGRA (graph-based ANN algorithm built to run natively on GPUs) that still works with CPUs for search full post: www.elastic.co/blog/elastic... and find us at GTC — we have a booth and I'll be around today and tomorrow PS: yeah, a lot of acronymselastic.coElastic and NVIDIA together unlock next generation enterprise AI searchElastic vector indexing with NVIDIA cuVS GPU acceleration eliminates a critical barrier to successful enterprise-scale AI deployments, enabling organizations to vectorize massive volumes of unstructur... 001
Philipp Krenn @xeraa.net · 16/03/2026obligatory GTC keynote tweet when you make the top slides: building HNSW graphs with up to 12x the throughput and 7x faster merges on elasticsearch and NVIDIA cuVS (a GPU-accelerated library for vector search) 100
Philipp Krenn @xeraa.net · 16/03/2026nvidia GTC where every second word (spoken and on booths) is AI, agent, or token 😬 PS: who else is around? 001
Philipp Krenn @xeraa.net · 11/03/2026BM25 for "sparse visual-word activations"? arxiv.org/abs/2603.05781 BM25 just refuses to disappear or even take the back set for retrieval 000
Philipp Krenn @xeraa.net · 09/03/2026soon... and there's already the .md version for every elastic docs page and you can download the full docs as MD in a ZIP file too 000
Philipp Krenn @xeraa.net · 06/03/2026but isn‘t this done more strictly in practice (to keep the legal risk minimal)? lots of little hacks here and there but bigger orgs seem to have higher standards, no? 000
Philipp Krenn @xeraa.net · 01/03/2026outage scenario: electricity turned off because of a drone / rocket attack ongoing updates: health.aws.amazon.com/health/status 000
Philipp Krenn @xeraa.net · 27/02/2026overview on amplifying.ai/research/cla... with a deep dive in amplifying.ai/research/cla... 5/5amplifying.aiAmplifying — AI Benchmark ResearchSystematic analysis of how AI systems make decisions — from product recommendations to developer tool choices. 000
Philipp Krenn @xeraa.net · 27/02/2026models have personalities in their recommendations and how conservative or cutting edge they are — massive differences even within a close "family" 4/5 100
Philipp Krenn @xeraa.net · 27/02/2026new defaults — some general and some language specific (the report is very JS and python focused) 3/5 100
Philipp Krenn @xeraa.net · 27/02/2026depending on the area, there is a clear bias towards building vs buying feature toggles as a full blown company were always a weird choice... 2/5 210
Philipp Krenn @xeraa.net · 27/02/2026*claude code is the new gatekeeper* fascinating report on the choices claude makes, build vs buy, the personality of different models, and who is winning / losing 1/5 PS: when creating code is cheap, the new moat is distribution 200
Philipp Krenn @xeraa.net · 25/02/20264x performance improvement by upgrading from #elasticsearch 8 to 9. "this one weird trick your cloud provider / hardware vendor hates" 😉 medium.com/trendyol-tec... PS: I have a hunch what made the difference here 021
Philipp Krenn @xeraa.net · 19/02/2026full announcement blog post: jina.ai/news/jina-em... hugging face: huggingface.co/collections/... paper for even more details: arxiv.org/abs/2602.15547jina.aijina-embeddings-v5-text: New SOTA Small Multilingual EmbeddingsTwo sub-1B multilingual embeddings with best-in-class performance, available on Elastic Inference Service, Llama.cpp and MLX. 000
Philipp Krenn @xeraa.net · 19/02/20262 new #jina models have entered the embedding arena — v5 multilingual: * small: 1024 dim, 32K context * nano: 768 dim, 8K context both support matryoshka dimension truncation (32+) and are at the top of current benchmarks — especially for their parameter size publicly accessible (non-commercial) 110
Philipp Krenn @xeraa.net · 19/02/2026gandalf prompt injection is still fun: gandalf.lakera.ai/gandalf 🧙♂️ though almost disappointing if one prompt takes you through multiple levels (🇦🇹)gandalf.lakera.aiGandalf | Lakera – Test your AI hacking skillsTrick Gandalf into revealing information and experience the limitations of large language models firsthand. 000
Philipp Krenn @xeraa.net · 13/02/2026media.tenor.coma man with the words well that escalated quickly written on his faceALT: a man with the words well that escalated quickly written on his face 000
Philipp Krenn @xeraa.net · 13/02/2026* OpenClaw agent opening a PR on matplotlib * reviewer rejects it per the repo's policies * agent replies with a personal attack github.com/matplotlib/m... 100
Philipp Krenn @xeraa.net · 12/02/2026* different views per solution and you can filter by version or other labels `label:"v9.3.0"` * the underlying issue describes what it does, for who, and the value proposition take a look on github.com/orgs/elastic... comments are currently disabled but let us know if that's a deal-breaker 2/2 000
Philipp Krenn @xeraa.net · 12/02/2026new public #elastic roadmap: * covering key initiatives like ES|QL, better dashboards,... * recently shipped features (those are our fiscal quarters) * upcoming features as in-progress, near-term, and mid-term 1/2 100
Philipp Krenn @xeraa.net · 08/02/2026catching up on the search dev room at #FOSDEM: lots of good stuff on fosdem.org/2026/schedul... glad we (or carly richmond) could put this together again. this won't be the last one. PS: given our colorful past, it's especially great to be back at FOSDEM pushing the search dev room and OSS :) 040
Philipp Krenn @xeraa.net · 17/01/2026PS: notebooklm.google.com is great to make this easier and more approachable 000
Philipp Krenn @xeraa.net · 17/01/2026jina-clip-v2 uses a multi-task, multi-stage contrastive learning strategy to align multilingual text and image representations to work well at both cross-modal and text-only retrieval. it excels at visually rich documents while flexibly truncating embedding dimensions arxiv.org/abs/2412.08802 100
Philipp Krenn @xeraa.net · 17/01/2026ReaderLM-v2 uses a new three-stage data synthesis pipeline called "draft-refine-critique" alongside a unified training framework to transform messy HTML into structured data, making it a highly effective tool that outperforms much larger models on web content extraction arxiv.org/pdf/2503.01151 100
Philipp Krenn @xeraa.net · 17/01/2026jina-embeddings-v4 works by projecting text & images into a shared semantic space using a unified Qwen2.5-VL backbone and task-specific LoRA adapters, which minimizes the modality gap and enables SOTA retrieval of visually rich documents via single- and multi-vector outputs arxiv.org/abs/2506.18902 100
Philipp Krenn @xeraa.net · 17/01/2026using a compact autoregressive backbone pre-trained on text and code, along with task-specific instruction prefixes and last-token pooling, jina-code-embeddings generates high-quality embeddings that achieve state-of-the-art performance competitive with much larger models arxiv.org/abs/2508.21290arxiv.orgEfficient Code Embeddings from Code Generation Modelsjina-code-embeddings is a novel code embedding model suite designed to retrieve code from natural language queries, perform technical question-answering, and identify semantically similar code snippet... 110
Philipp Krenn @xeraa.net · 17/01/2026jina-reranker-v3 uses a new "last but not late" interaction strategy that processes the query and multiple documents simultaneously in a single shared context window, allowing it to capture cross-document and query-document relationships with its compact 0.6B parameter model arxiv.org/abs/2509.25085 110
Philipp Krenn @xeraa.net · 17/01/2026long US weekend — great time to catch up on some @JinaAI_ papers about rerankers and code / multilingual / multimodal embeddings: * jina-reranker-v3 * jina-code-embeddings * jina-embeddings-v4 * ReaderLM-v2 * jina-clip-v2 100