Sign in

typedef

@typedef.ai
11 followers 0 following 88 posts

We are here to eat bamba and revolutionize the world of query engines. The Spark is gone, let's rethink data processing with a pinch of AI

PostsRepliesMedia
typedef @typedef.ai · 22/10/2025
Note: auto-routing is being explored; today you keep full control. check the repo for more: github.com/typedef-ai/f...
github.com
GitHub - typedef-ai/fenic: Build reliable AI and agentic applications with DataFrames
Build reliable AI and agentic applications with DataFrames - typedef-ai/fenic
000
typedef @typedef.ai · 22/10/2025
Mix providers (OpenAI, Anthropic) with simple aliases Use defaults for simple ops; override model_alias for complex ones Balance cost/latency/quality without extra orchestration
100
typedef @typedef.ai · 22/10/2025
Teams often wire a single model and pay in either cost or quality. With Fenic, you register multiple models once and select them per call.
100
typedef @typedef.ai · 22/10/2025
fenic's Multiple Model Configuration & Selection lets you pick the right model for each step, cheap where you can, powerful where you must. Think of it as a per-operator model dial across your pipeline.
100
typedef @typedef.ai · 21/10/2025
Thanks to @danielvanstrien.bsky.social and @lhoestq.hf.co for the collaboration and feedback that made this possible and to David Youngworth you built and maintains the integration!
000
typedef @typedef.ai · 21/10/2025
A few things you can do with this new integration. 1. Rehydrate the same agent context anywhere (local → prod) 2. Versioned, auditable datasets for experiments & benchmarks
huggingface.co
fenic
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
100
typedef @typedef.ai · 21/10/2025
Fenic ❤️ Hugging Face Datasets! You can now turn any fenic snapshot into a shareable, versioned dataset on @hf.co perfect for reproducible agent contexts and data sandboxes. Docs: huggingface.co/docs/hub/dat...
huggingface.co
fenic
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
120
typedef @typedef.ai · 21/10/2025
"AI confidence is high — but production results still lag." Our cofounder, Yoni Michael, shares why in CIO. Read it here 👉 www.cio.com/article/4069... #CIO #AIinEnterprise #Typedef
cio.com
CIOs’ AI confidence yet to match results
While a large percentage of IT and business leaders believe their AI efforts will meet or exceed expectations, only a small number have successfully deployed projects thus far.
000
typedef @typedef.ai · 24/09/2025
Common patterns: multi-step enrichment, RAG prep, nightly jobs with partial recomputes. for more check the Github repo: github.com/typedef-ai/f...
000
typedef @typedef.ai · 24/09/2025
With fenic, it’s explicit and simple: call .cache() where it matters. Protect pricey semantic ops (classify/extract) from re-execution Reuse cached results across multiple downstream analyses Recover from mid-pipeline failures without starting over
100
typedef @typedef.ai · 24/09/2025
Think of it as checkpointing for LLM workloads: cache after costly ops, restart from there if something fails. Without caching, teams re-pay tokens and time on retries: flaky APIs, disk hiccups, long recomputes.
100
typedef @typedef.ai · 24/09/2025
fenic's Local Data Caching & Persistence keeps expensive AI steps from rerunning and your pipelines resilient.
100
typedef @typedef.ai · 23/09/2025
Mix providers (OpenAI, Anthropic) with simple aliases Use defaults for simple ops; override model_alias for complex ones Balance cost/latency/quality without extra orchestration
000
typedef @typedef.ai · 23/09/2025
Teams often wire a single model and pay in either cost or quality. With Fenic, you register multiple models once and select them per call.
100
typedef @typedef.ai · 23/09/2025
fenic's Multiple Model Configuration & Selection lets you pick the right model for each step, cheap where you can, powerful where you must. Think of it as a per-operator model dial across your pipeline.
100
typedef @typedef.ai · 21/09/2025
Why do most AI projects stall? Because going from prototype → production is HARD. On Data Exchange, we share how Typedef makes inference-first pipelines actually work at scale. 👉 thedataexchange.media/typedef-fenic/
thedataexchange.media
The Fenic Approach to Production-Ready Data Processing
Kostas Pardalis on Inference-First Data Frames, Markdown as Structure, Semantic Query Operations, and Production AI Debugging.
000
typedef @typedef.ai · 20/09/2025
We’re honored to be featured in AI World Today! 🚀 Our co-founder Yoni Michael shares how Typedef is closing the gap between AI prototypes and production, making inference a first-class data operation. 👉 Read the full interview: www.aiworldtoday.net/p/interview-...
aiworldtoday.net
Bridging the AI Gap: How Yoni Iny's Typedef is Revolutionizing Data Processing
Yoni Michael, tech veteran and Typedef co-founder, transforms AI-powered data analytics with an innovative serverless platform for LLM workflows.
000
typedef @typedef.ai · 20/09/2025
We’re building the AI-native, inference-first infrastructure that powers scalable, production-ready LLM pipelines—no infrastructure headaches, just reliable results. Read more in AIM about how we’re overcoming pilot paralysis: aimmediahouse.com/ai-startups/...
aimmediahouse.com
For AI to Scale, Infrastructure Has to Change-Typedef Gets It
Typedef, a new AI infrastructure startup that officially launched on June 18, 2025, raised $5.5 million in seed funding, led by Pear VC.
000
typedef @typedef.ai · 20/09/2025
Fenic brings the reliability of DataFrame pipelines to AI workloads—semantic joins, markdown parsing, transcripts, and more—now strengthened with the 0.3.0 update. Dive into the latest improvements. → www.techzine.eu/blogs/data-m...
techzine.eu
Typedef project Fenic: A ‘dataframe’ for LLMs
Typedef provides purpose-built AI data infrastructure services for cloud workloads that need to handle LLM-powered pipelines, unstructured data Typedef  is Helping AI and Data Teams Build Faster,…
000
typedef @typedef.ai · 20/09/2025
AI fatigue is everywhere. But it’s not inevitable. In AI Journal, Typedef co-founder Yoni Michael shares how teams can escape “pilot paralysis” and move AI from prototype to production with confidence. 👉 Read the article: aijourn.com/ai-fatigue-i...
aijourn.com
AI Fatigue Is Real, But It's Fixable | The AI Journal
Enterprises have embraced generative AI with high expectations – new business insights, automated agents, real-time decision-making. What many got instead are
000
typedef @typedef.ai · 20/09/2025
Common patterns: review mining, invoice parsing, lead enrichment, spec extraction. for more, check the GitHub repo: github.com/typedef-ai/f...
000
typedef @typedef.ai · 20/09/2025
Define a Pydantic schema; get type-checked structs (ints, bools, lists, Optionals) Auto-prompting via function calling / structured outputs (OpenAI, Anthropic) Use unnest() and explode() to work with the data—no manual JSON wrangling
100
typedef @typedef.ai · 20/09/2025
Most teams hand-roll JSON parsing, brittle regex, and post-hoc validators. That’s slow and error-prone. With fenic, you keep it declarative.
100
typedef @typedef.ai · 20/09/2025
fenic's Structured Output Extraction turns LLM text into validated tables, directly in your DataFrame. Think of it as schema-first parsing: you define a Pydantic model; Fenic enforces it and returns structured columns.
100
typedef @typedef.ai · 18/09/2025
Common patterns: doc mining, content ingestion, RAG prep, taxonomy extraction. for more, including examples and documentation, check: github.com/typedef-ai/f...
010
typedef @typedef.ai · 18/09/2025
Type safety: Embedding/Markdown/JSON columns prevent incompatible ops Built-ins that matter: normalize, similarity, jq queries for JSON and many more Less glue: query structured + unstructured togethe,; mix dataframes + SQL + AI in one plan
100
typedef @typedef.ai · 18/09/2025
Most teams treat AI artifacts as loose strings/arrays: schema drift, brittle casting, ad-hoc JSON parsing, and inconsistent similarity math. In fenic, these are first-class.
110
typedef @typedef.ai · 18/09/2025
fenic's First-Class AI Data Types make embeddings, markdown, and JSON real, typed columns, with the right operations built in. Think of it as strong types for meaning and structure: safer pipelines, richer queries.
100
typedef @typedef.ai · 16/09/2025
for examples and more information, check: github.com/typedef-ai/f...
000
typedef @typedef.ai · 16/09/2025
Most teams fight drift: regex stacks, ad-hoc prompts, inconsistent tags. With fenic you can: Define classes once; get schema-clean, consistent labels Zero-shot or few-shot with real examples (not just descriptions) Batching, caching, retries built-in all testable in the same plan.
100
typedef @typedef.ai · 16/09/2025
Think of it as CASE/WHEN for unstructured text: you define the classes and fenic constrains outputs to exactly those labels.
100
typedef @typedef.ai · 16/09/2025
fenic's Semantic Classification turns free-text into clean enums right inside your DataFrame.
100
typedef @typedef.ai · 12/09/2025
Common patterns: job↔candidate matching, RAG candidate retrieval, fuzzy dedupe/linking. for examples, demos and more information: github.com/typedef-ai/f...
000
typedef @typedef.ai · 12/09/2025
With fenic, it’s one call in your pipeline. Get top-K matches with similarity scores Batch ANN under the hood; models & rate limits live in your Session Compose with selects/filters/LIMIT to control cost and spend
100
typedef @typedef.ai · 12/09/2025
Think of it as a k-NN join for meaning: for each row in A, fetch the top-K most similar rows in B (cosine), with scores. Most teams roll their own: separate embedding jobs, vector index wiring, threshold tuning, brittle Python loops.
100
typedef @typedef.ai · 12/09/2025
fenic's Semantic Similarity Join (Vector Join) finds nearest neighbors across tables using embeddings, right inside your DataFrame.
110
typedef @typedef.ai · 09/09/2025
No breaking changes from 0.3.x. Full write-up, examples, and release links: www.typedef.ai/blog/fenic-0...
typedef.ai
Fenic 0.4.0 Released: Declarative Tools, MCP, and HuggingFace — plus major DX & reliability gains
Fenic 0.4.0 adds declarative tools, an MCP server, GPT-5/Claude 4.1, HuggingFace connector, local metrics, and DX & performance improvements. Upgrade now.
010
typedef @typedef.ai · 09/09/2025
HuggingFace connector: hf://… URIs; also directory loaders to turn folders into DataFrames. Built-in metrics: track latency, tokens, and cost per pipeline locally. DX & stability: clearer errors (e.g., union()), handy null()/empty(), safer S3 auth, smarter retries. 
100
typedef @typedef.ai · 09/09/2025
MCP server out-of-the-box: plug Fenic tools into Claude Code, Gemini CLI, Cursor & friends. Latest model support: GPT-5 + Claude Opus 4.1 with fail-fast provider key validation.
100
typedef @typedef.ai · 09/09/2025
fenic 0.4.0 is live: declarative tools for agents, a production-ready MCP server, and direct reads from HuggingFace plus big DX & reliability gains.  Highlights: Declarative tools: define function-calling tools as data (type-safe, reviewable, reusable). 
151
typedef @typedef.ai · 08/09/2025
You get reliable, table-ready outputs instead of free‑form text.
000
typedef @typedef.ai · 08/09/2025
pass a Pydantic model to semantic.extract or semantic.map and fenic will prompt the model, parse the output, and return typed/struct columns (first‑class support for OpenAI/Anthropic structured responses).
100
typedef @typedef.ai · 08/09/2025
fenic ensures LLM outputs conform to schemas using Pydantic models.
110
typedef @typedef.ai · 06/09/2025
You can clean/filter/LIMIT before invoking expensive AI calls to reduce cost and provide a better context to the model.
000
typedef @typedef.ai · 06/09/2025
Create from pandas, select/filter/group_by/join, use a Session (like SparkSession) and mix SQL + LLM operations,.
100
typedef @typedef.ai · 06/09/2025
fenic offers standard DataFrame operations with a familiar Spark/Pandas-like API.
100
typedef @typedef.ai · 05/09/2025
Because UDFs are plannable, you can still filter/limit/group with standard DataFrame ops to reduce AI cost, keep pipelines auditable, and only push to the model the enriched, relevant data it needs.
010
typedef @typedef.ai · 05/09/2025
Use them when native column ops aren’t enough, e.g. complex normalization, domain enrichment (spaCy, custom tokenizers, external feature libs), or bespoke pre/post-processing before embeddings/LLM calls.
100
typedef @typedef.ai · 05/09/2025
fenic UDFs allow you to inject arbitrary Python (including external libraries) directly into the fenic execution plan while preserving lazy planning, metrics and reproducibility.
110
typedef @typedef.ai · 03/09/2025
Sessions include local storage for ephemeral tables/caching, making it possible to checkpoint & resume pipelines by saving tables/views and reloading them to continue work.
010