Sign in

Hyperplane

@hyperplane.bsky.social
36 followers 145 following 37 posts

Your weekly read. From POC to Production, at scale. 🫵 Follow our substack: thehyperplane.substack.com 👀 Our Ebook: hyperplane.gumroad.com/l/fine-tunin…

PostsRepliesMedia
Hyperplane @hyperplane.bsky.social · 06/05/2025
Updated link: hyperplane.gumroad.com/l/fine-tunin...
hyperplane.gumroad.com
Fine-Tuning STT Models for Edge Devices
This eBook explores the challenges of speech-to-text (STT) models failing to recognize children's voices and provides a practical solution. We cover data preparation, model fine-tuning, and optimizati...
010
Hyperplane @hyperplane.bsky.social · 02/04/2025
Finally made it! It’s been a long ride, but the first real STT guide for kids’ voices on edge devices is here. Check it out on Gumroad if you're curious! the link for the book mlvanguards.gumroad.com/l/fine-tunin... (it's free)
100
Hyperplane @hyperplane.bsky.social · 31/03/2025
Without observability, agents are just black boxes making guesses.
With the right tools, they become transparent, testable, improvable systems. Agentic systems are only as useful as their debug-ability.
000
Hyperplane @hyperplane.bsky.social · 28/03/2025
MLOps is more than just tools. Reproducibility, scalability, and observability are a must. Just setting up Kubeflow or MLflow doesn’t make you an MLOps expert.
000
Hyperplane @hyperplane.bsky.social · 27/03/2025
Launching tomorrow for all subscribers: open.substack.com/pub/mlvangua...
open.substack.com
$0.00. No Gimmicks. The first real STT guide for kids’ voices on edge devices.
Everyone else sells vague theory for $49. We’re giving away real engineering, code and code for free.
000
Hyperplane @hyperplane.bsky.social · 27/03/2025
Raw audio is unpredictable. Training is expensive. Inference is a balancing act. Deployment is… well, never as simple as pip install. Everyone’s hyped about LLMs and edge deployments, but few talk about what it actually takes to it to production. So we wrote the guide we wish we had.
100
Hyperplane @hyperplane.bsky.social · 26/03/2025
Real MLOps doesn’t happen in 1 day or 1 week. Knowing how to use different tools doesn’t make you an MLOps expert. It takes time to understand the whole process, and it takes time to learn how to have arguments to convince the management, the CEO, or other stakeholders.
010
Hyperplane @hyperplane.bsky.social · 26/03/2025
🚨 Alert: This article is dangerously true! mlvanguards.substack.com/p/the-malpra...
mlvanguards.substack.com
The malpractice of AI industry
The dangerous rise of false authority in AI, and how it’s quietly making us dumber.
000
Hyperplane @hyperplane.bsky.social · 26/03/2025
Still better than no boat at all! In all realness, code generation is a great assistant for an already great programmer 🤷
000
Reposted by Hyperplane
Neo4j @neo4j.com · 25/03/2025
How does GraphRAG work? What are its advantages over other RAG? Let's return to Michael Hunger's blog post in which he demonstrates its practical application using a hashtag#Neo4j example! Happy weekend reading! bit.ly/4f8wLVp hashtag#knowledgegraphs
bit.ly
What Is GraphRAG? - Graph Database & Analytics
GraphRAG is a powerful retrieval mechanism that improves Generative AI applications by taking advantage of the rich context in graph data structures.
021
Hyperplane @hyperplane.bsky.social · 26/03/2025
We show more in the upcoming eBook, free for all subscribers: mlvanguards.substack.com
mlvanguards.substack.com
ML Vanguards | Substack
Escaping PoC purgatory: Your Weekly Guide to production paradise. Click to read ML Vanguards, a Substack publication.
000
Hyperplane @hyperplane.bsky.social · 26/03/2025
- Normalize & clean transcripts (remove garbage text, repeated words, weird artifacts) - Filter out the junk - Split (70/15/15) & push to @hf.co for easy access during training 2/2
100
Hyperplane @hyperplane.bsky.social · 26/03/2025
In less than 24 hours, we turned 5K messy files (from @kaggle.com) into a clean dataset of ~3.9K audio+transcription pairs. Here’s how: - Resample & filter (standardize to 16kHz, cut long/empty clips) - Auto-transcribe with SpeechBrain (ran it on CPU — I'm GPU poor 😅) 1/2
130
Hyperplane @hyperplane.bsky.social · 25/03/2025
In the last 2 years, prompt engineering has been treated as an afterthought, a means to an end. But in reality, a prompt is the most crucial hyperparameter of any GenAI system. Its design can make or break the output quality, much like tuning a model's parameters determines its performance.
000
Hyperplane @hyperplane.bsky.social · 24/03/2025
It's kinda free for all newsletter subscribers: mlvanguards.substack.com
mlvanguards.substack.com
ML Vanguards | Substack
Escaping PoC purgatory: Your Weekly Guide to production paradise. Click to read ML Vanguards, a Substack publication.
000
Hyperplane @hyperplane.bsky.social · 24/03/2025
We're launching an ebook on 28th this week 😳 Regular speech-to-text tech struggles with kids’ voices since they’re higher-pitched and less predictable. We worked on that by creating a smaller, more accurate model that works well with children’s speech, even in noisy or low-power settings.
111
Hyperplane @hyperplane.bsky.social · 24/03/2025
Read more here: mlvanguards.substack.com/p/data-is-bo...
mlvanguards.substack.com
Data is boring
But broken search results are worse
000
Hyperplane @hyperplane.bsky.social · 24/03/2025
6. Vector Database A vector database like Vespa store sembeddings and enable allowing similarity searches. They also use metadata to improve relevance by associating vectors with key attributes like document type, page number, or detected visual features. 7/7
100
Hyperplane @hyperplane.bsky.social · 24/03/2025
5. Chunking Strategy Splits documents into manageable chunks for embedding: - Layout-based chunking is for visual embeddings. - Text density and structure for traditional embeddings. This preserving context without overloading the vector database 6/7
100
Hyperplane @hyperplane.bsky.social · 24/03/2025
4. Embedding Models For converting document content into vectors. - Traditional embeddings for documents with clean text extracted via OCR. - Vision Language Models (VLM) handle multimodal documents with complex visual structures like tables, charts, and diagrams. 5/7
100
Hyperplane @hyperplane.bsky.social · 24/03/2025
3. Decisional Algorithm The algorithm is centralized, making informed decisions based on input from the embedding decider. - Text-heavy documents are processed with OCR and text embedding models. - Documents with complex layouts use visual language models (eg ColPali) instead, skipping OCR. 4/7
100
Hyperplane @hyperplane.bsky.social · 24/03/2025
2. PDF Embedding Decider This decider analyzes the document's structure, using tools like a layout analyzer, visual element detector, or text density analyzer, to classify whether a traditional text embedding or a multimodal vision embedding is appropriate. 3/7
100
Hyperplane @hyperplane.bsky.social · 24/03/2025
1. PDF Reader The starting point of any pipeline is the PDF reader. Its job is to extract pages and pass them downstream. A high-quality reader ensures no lost information, whether the content is text-heavy, image-dense, or filled with tables and graphs. 2/7
100
Hyperplane @hyperplane.bsky.social · 24/03/2025
The principles of an indexing pipeline: 1/7
111
Hyperplane @hyperplane.bsky.social · 24/03/2025
Not all PDFs are created equally. Some PDFs are beautifully structured with clean text, while others are chaotic with dense layouts, tables, or images. Ignoring these differences means risking ineffective indexing and poor search retrieval.
000
Reposted by Hyperplane
Matt Pocock @mattpocock.com · 21/03/2025
It's the evals, stupid
3281
Hyperplane @hyperplane.bsky.social · 28/02/2025
AI can pretty much do anything, but it lacks that human creativity. Do you prefer quick tasks done or creativity? #AI #Question
000
Hyperplane @hyperplane.bsky.social · 26/02/2025
Dspy fixes this. It treats LLMs like actual programmable components instead of "hope this works" spells signatures, modules, optimizers, whatever, read the thing if you care we have a new article about it. with code mlvanguards.substack.com/p/prompts-ar...
mlvanguards.substack.com
Prompts are lying to you
Combining prompt engineering with DSPy for maximum control
020
Hyperplane @hyperplane.bsky.social · 26/02/2025
"prompt engineering" is just fancy copy-pasting at this point people tweaking prompts like they're adjusting a car mirror, thinking it'll make them drive better you’re optimizing nothing, you’re just guessing
132
Hyperplane @hyperplane.bsky.social · 24/02/2025
It's easy to over use DRY which slooows you down as a dev. Some code bits can be repeated because the end goal is to have maintainable and readable code, and not just unique lines of code
000
Hyperplane @hyperplane.bsky.social · 21/02/2025
Here’s the full article if you want to nerd out open.substack.com/pub/mlvangua...
open.substack.com
The mind’s keyboard: How Brain2Qwerty is transforming thoughts into text
A deep dive into the most exciting research yet, turning brainwaves into text
020
Hyperplane @hyperplane.bsky.social · 21/02/2025
Hot take: Meta’s working on a brain-reading tech that turns thoughts into text. No surgery, just a headset. It’s got typos (32% errors) but skilled typists did way better. It's wild to imagine this helping folks who can’t speak or type. Would you trust a computer with your thoughts?
120
Hyperplane @hyperplane.bsky.social · 21/02/2025
🙈 and we happened to write an article on it with Marvelous MLOps on substack: mlvanguards.substack.com/p/unicorns-a...
mlvanguards.substack.com
Unicorns and Rainbows: The Reality of Implementing AI in a Corporate
The AI Bubble: Hype, herd mentality, and harsh realities
000
Hyperplane @hyperplane.bsky.social · 21/02/2025
Agreed. Plenty of times implementing AI in the enterprise is backfiring and provides little to no value. However, Zalando found a good use case for it that actually works corporate.zalando.com/en/newsroom/...
corporate.zalando.com
100
Hyperplane @hyperplane.bsky.social · 20/02/2025
Or just bad!
000
Hyperplane @hyperplane.bsky.social · 19/02/2025
Trend keeps moving with all these new tools making it tough to keep track. The question is ChatGPT or Perplexity for better searching?
020
Hyperplane @hyperplane.bsky.social · 18/02/2025
here is the full article if you are curious: mlvanguards.substack.com/p/data-is-bo...
mlvanguards.substack.com
Data is boring
But broken search results are worse
020
Hyperplane @hyperplane.bsky.social · 18/02/2025
Indexing messy PDFs? Try these: LayoutAnalyzer for messy formatting. TextDensityAnalyzer to tell text from images. VisualElementAnalyzer for charts/diagrams. TableDetector to grab tables. Convert them to images and use Vision Models like ColPali/ColQwen2.
121
Hyperplane @hyperplane.bsky.social · 29/01/2025
Are you stuck in the whirlpool of AI trends? Building production-grade AI isn’t about chasing buzzwords — it’s about combining engineering knowledge with practical AI to deliver systems that actually work, scale, and drive real-world impact. We got you covered at mlvanguards.substack.com
mlvanguards.substack.com
ML Vanguards | Substack
Escaping PoC purgatory: Your Weekly Guide to production paradise. Click to read ML Vanguards, a Substack publication.
020