Sign in

Tensorlake

@tensorlake.ai
43 followers 7 following 56 posts
PostsRepliesMedia
Tensorlake @tensorlake.ai · 06/11/2025
Amazing!!! We always want you to try it out with the "hardest" documents - So glad to hear it worked! The goal of an effective document parser SHOULD be to help with the hardest problems. a low-quality scan of a Norwegian document with statistical tables from 1926 is the perfect test doc 🔥
010
Reposted by Tensorlake
Henning @radbrt.eurosky.social · 05/11/2025
The tensorlake playground was, unlike AWS textract and every other tool I have tried, able to parse my angled, low-quality scan of Norwegian pay statistics from 1926. Not that 1926 Norwegian statistical tables is a generally useful benchmark…
121
Tensorlake @tensorlake.ai · 05/11/2025
We published: ✓ Full methodology ✓ Corrected OCRBench v2 ground truth ✓ Comparative analysis across all major providers Read the full benchmark: tlake.link/benchmarks Stop benchmarking vanity metrics. Start measuring what breaks.
tlake.link
Benchmarking the Most Reliable Document Parsing API | Tensorlake
Learn how Tensorlake built the most reliable document parsing API by measuring what actually matters: structural preservation, reading order accuracy, and downstream usability. See benchmark results c...
000
Tensorlake @tensorlake.ai · 05/11/2025
The results were clear: Tensorlake: 86.8% TEDS, 91.7% F1 AWS Textract: 80.7% TEDS, 88.4% F1 Azure: 78.1% TEDS, 88.1% F1 Docling: 63.8% TEDS, 68.9% F1 The gap? 670 fewer manual reviews per 10k documents.
100
Tensorlake @tensorlake.ai · 05/11/2025
We evaluated on OCRBench v2, OmniDocBench, and 100 real enterprise docs using two metrics that predict production success: TEDS (Tree Edit Dist): Measures if tables stay tables JSON F1: Measures if downstream systems can use the output Not "is the text similar?" but "can automation actually work?"
100
Tensorlake @tensorlake.ai · 05/11/2025
Traditional benchmarks test on clean PDFs and measure text accuracy. But your production failures come from: - Collapsed tables - Jumbled reading order - Missing visual content - Hallucinated extractions None of this shows up in your scores.
100
Tensorlake @tensorlake.ai · 05/11/2025
Document parsing benchmarks have been measuring the wrong thing. We tested every major parser on real enterprise documents. The results will change how you think about OCR accuracy 🧵
Two dense document pages flank a skeptical person’s sticker-style portrait against a green gradient, link text centered below.
122
Tensorlake @tensorlake.ai · 23/10/2025
Want to build scalable data lakes w/ Tensorlake + @qdrant.bsky.social? In the free Qdrant Essentials Course, learn how to: - Architect vector-powered data lakes - Optimize ETL pipelines - Create knowledge graphs - Integrate @langchain.bsky.social agents for natural language queries t.co/OoPZswrL7z
Promotional banner for the Qdrant Essentials Course featuring Tensorlake. Text reads: ‘Improve collection querying with knowledge graphs.’ On the left is the Qdrant logo and course title; on the right is a smiling woman with long curly brown hair wearing a cream-colored top, set against a purple gradient grid background.
012
Tensorlake @tensorlake.ai · 16/10/2025
Try it yourself with our SEC filing analysis notebook: tlake.link/notebooks/vl... Shows how to extract cryptocurrency metrics from 10-Ks and 10-Qs using page classification Full changelog: tlake.link/changelog/vlm What would you build with this?
tlake.link
New: Vision Language Models for Document Processing
Tensorlake now uses Vision Language Models (VLMs) across multiple features including page classification, figure/table summarization, and structured extraction, enabling faster and more intelligent do...
000
Tensorlake @tensorlake.ai · 16/10/2025
Where we leverage VLM support: 📄 Page Classification: Large docs, specific sections needed 📊 Table/Figure Summarization: Visual data in reports ⚡ skip_ocr=True: When reading order is complex and for diagrams and scanned docs Text extraction still uses OCR for best quality
100
Tensorlake @tensorlake.ai · 16/10/2025
Real results from analyzing 8 SEC filings: - 1,500+ total pages - 427 relevant pages identified by VLM - Processing time: 5 minutes → 45 seconds per document All without sacrificing accuracy
100
Tensorlake @tensorlake.ai · 16/10/2025
Our solution: VLMs understand document structure visually Example: Extracting crypto holdings from SEC filings 1. VLM classifies which pages contain financial data (~50 out of 200 pages) 2. Extract only from relevant pages 3. Skip 70% of processing Result: 80-90% faster ⚡
100
Tensorlake @tensorlake.ai · 16/10/2025
The problem: Processing 200-page documents when you only need specific information is slow and expensive Traditional approach: OCR everything → Convert to text → Search → Extract This wastes time processing irrelevant pages
100
Tensorlake @tensorlake.ai · 16/10/2025
New: Vision Language Models now power key document processing features We're using VLMs for: - Page classification in large documents - Table/figure summarization - Fast structured extraction (skip_ocr mode) Here's what this means for document processing 🧵
100
Reposted by Tensorlake
David Calavera @calavera.dev · 14/10/2025
The company I work for, @tensorlake.ai, is hiring a couple of roles remote within the US: tensorlake.ai/careers You might be a great fit if you like working with Rust, Python, K8s, me?, and you enjoy building products for developers.
tensorlake.ai
Tensorlake
Transform Data Into Knowledege
082
Tensorlake @tensorlake.ai · 10/10/2025
Build approval workflows that trigger on specific feedback. Extract complete edit history for regulatory compliance. Route documents based on flagged sections, all programmatically. Live now in our API, SDK, and Cloud.
000
Tensorlake @tensorlake.ai · 10/10/2025
Now you can parse .docx files with tracked changes preserved as clean, structured HTML: - <del> tags for deletions - <ins> tags for insertions - <span class="comment"> for reviewer notes
100
Tensorlake @tensorlake.ai · 10/10/2025
Most parsers strip all tracked changes when you extract the text. That means: ❌ Lost audit trails ❌ Manual review of revision history ❌ No programmatic access to reviewer comments ❌ Workflows that can't route based on specific edits
Tensorlake interface showing parsed Word document with tracked changes preserved as HTML tags, displaying an insurance claim report
100
Tensorlake @tensorlake.ai · 02/10/2025
Perfect for: → RAG pipelines (better chunking) → Knowledge graphs (accurate trees) → Document navigation → Table of contents generation Changelog: tlake.link/changelog/he... Try it: tlake.link/notebooks/he...
tlake.link
Try it in Colab
No description
000
Tensorlake @tensorlake.ai · 02/10/2025
Every section header now returns: - level: 0 for #, 1 for ##, 2 for ###, etc - content: clean text - proper nesting for up to 6 levels Enable with: cross_page_header_detection=True That's it.
100
Tensorlake @tensorlake.ai · 02/10/2025
Tensorlake analyzes numbering patterns (1, 1.1, 1.2) and visual structure across the ENTIRE document. Then corrects misidentified header levels automatically. Works even when headers span page breaks.
100
Tensorlake @tensorlake.ai · 02/10/2025
OCR engines constantly mess up document hierarchy. Section 2.2 becomes a top-level header (##) instead of nested (###). We just shipped automatic header correction. 🧵 How it works:
Comparison of document header detection. Left side "Just OCR" shows incorrect hierarchy with section 2.2 at wrong indent level. Right side "Header Correction" shows proper nesting where 2.2 is correctly indented under section 2. Bottom shows Python code: doc_ai.parse_and_wait() with cross_page_header_detection=True parameter. Green gradient background with Tensorlake logo.
110
Tensorlake @tensorlake.ai · 19/09/2025
There's no reason your applications should not be citation-ready. Dive deeper and try out the Colab notebook linked at the bottom of the blog
tlake.link
Citation-Aware RAG: How to add Fine Grained Citations in Retrieval and Response Synthesis | Tensorlake
Learn how to build citation-aware RAG systems that link AI responses back to exact source locations in documents. This technical guide covers document parsing with spatial metadata, chunking strategie...
000
Tensorlake @tensorlake.ai · 19/09/2025
Step 3: Generating AI responses with verifiable citations Once your chunks carry anchors, retrieval doesn’t change. You can use the dense, hybrid, or reranker setup you already have. Consider hiding the anchors in prose, while keeping them in output and making IDs clickable.
RAG citation workflow diagram on dark green background showing document processing pipeline: Document (PDF/Image) → Tensorlake Document AI → Parsed Elements (Text, Tables, Figures, and Bounding Box) → merge and insert anchors → Chunks and Anchors (Clean text and citation IDs) → splits to Citation Metadata (page, bounding box, citation IDs) and Vector DB (embeddings, text, and metadata). URL: https://tlake.link/blog/rag-citations
100
Tensorlake @tensorlake.ai · 19/09/2025
Step 2: Create contextualized chunks Iterate through page fragment objects and create appropriately sized chunks by combining them. As you create the chunks, you can create contextualized metadata to help during retrieval.
Before and after comparison of document chunking on dark green background. Top panel "Without Contextualized Chunking" shows plain text: "SMOTE creates a broader decision region for the minority class...". Bottom panel "With Contextualized Chunking" shows same text with citation anchor "<c>2.1</c>" and metadata: {"2.1": {"page": 23, "bbox": {...}}}. URL: https://tlake.link/blog/rag-citations
100
Tensorlake @tensorlake.ai · 19/09/2025
Step 1: Parse docs with bounding boxes Using our Document AI API you get a full document layout. For each page fragment you have access to the page number, fragment type, content, and bounding box. Making it easy to add metadata and anchor points to chunks before embedding.
Tensorlake Document AI interface showing document layout analysis with JSON output on left displaying fragment types, content, and bounding box coordinates, and PDF preview on right with highlighted text regions and yellow bounding boxes overlaid on research paper content
100
Tensorlake @tensorlake.ai · 19/09/2025
Citations. When users ask "where did this come from?" your system should point to the exact page fragment...not just "file_name.pdf". Built citation-aware RAG with spatial metadata has: → Parse docs with bounding boxes → Embed citation anchors in chunks → Return page numbers + coordinates A 🧵
100
Reposted by Tensorlake
David Calavera @calavera.dev · 15/09/2025
Job update: a couple of weeks ago, I joined @tensorlake.ai full time. I’m having a lot of fun building the product with @diptanu.bsky.social and the rest of this wonderful team. We have a few open positions if you’d like to work with us: www.linkedin.com/jobs/search/...
194
Tensorlake @tensorlake.ai · 11/09/2025
Trust in retrieval comes from evidence. Tensorlake ties every lookup back to the original table cell. Read the blog and try it out for yourself 👇
tlake.link
Parse and Retrieve Dense Tables Accurately with Tensorlake | Tensorlake
Learn how Tensorlake preserves structure in dense, multi-page tables—returning DataFrames with summaries and bounding boxes for accurate, explainable retrieval.
000
Tensorlake @tensorlake.ai · 11/09/2025
In finance, clinical trials, or performance benchmarks, dense tables contain mission-critical data. But flatten that data like most parsers do and trust is lost. Tensorlake restores trust by preserving structure, generating summaries for effective embeddings, and attaching evidence via b-boxes.
Side-by-side comparison of a dense healthcare data table in PDF format and its structured DataFrame output. A green background with the Tensorlake logo shows an arrow pointing from the PDF to the DataFrame. The caption reads “Parse Dense Tables Reliably” with the link “tlake.link/blog/dense-tables” at the bottom.
111
Tensorlake @tensorlake.ai · 11/09/2025
“Because the AI said so” isn’t good enough. Every answer should come with receipts (citations + context). Learn how to make your AI correct and verifiable in this month’s Document Digest newsletter 👇
tlake.link
The Document Digest by Tensorlake
Product updates and dev insights from the Tensorlake team.
000
Tensorlake @tensorlake.ai · 05/09/2025
You can now login into Tensorlake using Microsoft and Azure SSO credentials. This is the beginning of better integration with Microsoft Azure and Tensorlake. If you are using Azure, and need better Document Ingestion and ETL for unstructured data reach out to us!
020
Tensorlake @tensorlake.ai · 05/09/2025
To build trustworthy AI, your data needs proof. Get citations for every field extracted with Tensorlake. Read the blog and try our citations with the example notebooks: tlake.link/blog/citations
000
Tensorlake @tensorlake.ai · 21/08/2025
Now it's time to dive in deeper, check out the blog post we wrote about advanced RAG (and try it out yourself with the Colab notebooks we included). Fact‑check Tesla headlines w/ SEC filings using our advanced RAG support. Code + Colabs in the blog: tlake.link/advanced-rag
tlake.link
Accelerate Advanced RAG with Tensorlake | Tensorlake
Advanced RAG that survives production: keep context fresh, preserve structure, and plan retrieval using Tensorlake to turn messy PDFs into traceable answers. We demonstrate it by fact-checking Tesla n...
000
Tensorlake @tensorlake.ai · 21/08/2025
Step 4: Test your context-aware agents This is the fun part, use Tensorlake to extract key claims from news articles, then use your @langchain.bsky.social agent to query ChromaDB and determine whether the claims are rooted in fact.
Code snippet styled in a terminal window with green background. The Python code defines a Pydantic model NewsArticle with two fields: article_key_points as a list of strings describing key points of the article, and article_summary as a string summarizing the article. It then defines structured_extraction_options with a StructuredExtractionOptions object using the NewsArticle schema. Finally, it calls doc_ai.parse_and_wait with file=article and the structured extraction options, assigning the result to article_result.Code-styled summary window with green background. Three claims about Tesla filings are listed:  Claim: Record deliveries/deployments in Q4 2024 — Supported by Filings? YES — Notes: Clearly stated in filings.  Claim: Deliveries/deployments indicate quarterly financials/profits — Supported by Filings? NO — Notes: Explicitly contradicted by filings.  Claim: Tesla’s profits or net income figures for Q4 2024 — Supported by Filings? NO — Notes: Not yet released; filings only preview data.
100
Tensorlake @tensorlake.ai · 21/08/2025
Step 3: Contextualize Queries Don't rely on users to make queries that are specific for your data. Instead, make sure you contextualize your query so that during hybrid search you're finding the most relevant and accurate chunks. In our example, we used @langchain.bsky.social to help
Code snippet styled in a terminal window with green background. The Python function create_query(state: State) uses llm.invoke to generate a natural language query for searching a vector database of Tesla SEC filings. It prints the query content and returns a dictionary containing a messages list with an AIMessage holding the query, along with the query string itself.
100
Tensorlake @tensorlake.ai · 21/08/2025
TIP: Keep it fresh! Don’t just rebuild indexes monthly or annually, watch for new/changed filings and re-ingest continuously. The only rule: idempotency keyed on the SEC accession number.
100
Tensorlake @tensorlake.ai · 21/08/2025
Step 2: Chunk, Store, and Retrieve Data With clean, structured, and accurate data you can chunk and embed your documents in a way that is effectively and accurately retrievable by agents. In our example, we used @chonkieai.bsky.social and ChromaDB.
Code snippet styled in a terminal window with green background. The Python dictionary chunk_data is being built, containing keys such as id, pdf_url, chunk_index, text, start_index, end_index, and metadata. The metadata includes nested fields like source_type set to 'tesla_sec_filing', pdf_url, chunk_id, total_chunks, chunk_index, filing_date pulled from structured_data, key_points, and page_classifications.
100
Tensorlake @tensorlake.ai · 21/08/2025
Step 1: Ingest and pre-process your documents With a single API call, you can turn messy PDFs into page-aware, table-preserving, structured context. Simply: 1. Define page classes for different documents 2. Define schemas to extract relevant data 3. Parse
Code snippet displayed in a terminal-style window with green background. The Python code shows a call to doc_ai.parse_and_wait passing arguments: file=filing_url, structured_extraction_options=structured_extraction_options, and page_classifications=page_classifications, with the result assigned to a variable named result.
100
Tensorlake @tensorlake.ai · 21/08/2025
“RAG is dead” is lazy. What’s dead is cosine‑N without a retrieval plan. We ship advanced RAG...out of the box: • Classify pages → target sections • Extract structured fields → filter by form_type, fiscal_period • Verify data; cite page/bbox Want to know how? 🧵👇
Screenshot of a Tensorlake blog post titled Accelerate Advanced RAG with Tensorlake by Dr. Sarah Guthals, dated August 19, 2025. The header image reads ‘RAG isn’t dead, undisciplined retrieval is’ with a green wave design and Tensorlake logo. The page shows the TL;DR summary emphasizing the need for a retrieval plan over naive Top-N cosine RAG, and a table of contents with sections on the Freshness Principle, Accelerate Advanced RAG, a real-world application fact-checking Tesla SEC filings, and treating context as a hard requirement.
110
Tensorlake @tensorlake.ai · 29/07/2025
Follow this quick tutorial and build a *smart* real estate agent...agent: tlake.link/nlp-agent
tlake.link
Real Estate Agent with LangGraph (using CLI) - Tensorlake
Build a real estate agent using LangGraph to interact with purchase agreements and answer agent queries.
000
Tensorlake @tensorlake.ai · 29/07/2025
Build a smart real estate agent (no license required). 🧠 LangGraph (by @langchain.bsky.social) + 📝 Tensorlake Contextual Signature Detection = ✅ Knows who signed ✅ When they signed ✅ If it’s ready to close Full tutorial + code linked below 👇
100
Tensorlake @tensorlake.ai · 14/07/2025
Most "unstructured" parses fail on when layout gets tricky: multiple columns, fragmented text blocks, mixed reading order Tensorlake doesn't. ✅ Authors parsed as one clean chunk ✅ Abstract follows, exactly as it should Unstructured ≠ unordered Preserve reading order. Parse with Tensorlake.
010
Tensorlake @tensorlake.ai · 12/06/2025
Try out the tool with this simple Colab Notebook
tlake.link
Google Colab
000
Tensorlake @tensorlake.ai · 12/06/2025
Want to see the tool in action? Check out this quick demo or try it out in the Colab Notebook (linked in the comments)
youtube.com
LangChain + Tensorlake: Unlocking Document Understanding for Agents
YouTube video by Tensorlake
100
Tensorlake @tensorlake.ai · 11/06/2025
Just published 🐦‍⬛ langchain-tensorlake 💚 A new @langchain.bsky.social tool to parse real-world documents (PDFs, scans, forms) with Tensorlake & feed structured data right into your agents. Built for devs wrangling docs in legal, finance, healthcare & more. Learn more: tlake.link/langchain-tool
tensorlake.ai
Tensorlake
Transform Data Into Knowledge
010
Tensorlake @tensorlake.ai · 05/06/2025
Follow this quick tutorial: tlake.link/nlp-agent
tlake.link
Real Estate Agent with LangGraph (using CLI) - Tensorlake
Build a real estate agent using LangGraph to interact with purchase agreements and answer agent queries.
000
Tensorlake @tensorlake.ai · 05/06/2025
Build a smart real estate agent (no license required). 🧠 LangGraph (by @langchain.bsky.social) + 📝 Tensorlake Contextual Signature Detection = ✅ Knows who signed ✅ When they signed ✅ If it’s ready to close Full tutorial + code linked below 👇
120
Tensorlake @tensorlake.ai · 28/05/2025
That one missing signature? It might delay a claim or void a contract Now in Tensorlake: Contextual Signature Detection → Detect handwritten, typed, or image-based → Trigger routing, alerts, or human review → API + SDK + Playground Read the full blog post tlake.link/signature-de...
tlake.link
Tensorlake
Transform Data Into Knowledege
000