Sign in

Vivek Kalyan

@vivekkalyan.com
74 followers 169 following 59 posts

Applied AI. Building cartograph.app. Previously: Head of AI @ handshakes.ai 🇸🇬 I write (sparsely) at: vivekkalyan.com

PostsRepliesMedia
Vivek Kalyan @vivekkalyan.com · 20/01/2025
Inspired by my experience consulting for companies and the rise of inference time compute for LLMs (as well as the acceptance of users waiting while the LLM *thinks*), I suggest many teams working on RAG should rethink if you can improve your accuracy by relaxing latency.
010
Vivek Kalyan @vivekkalyan.com · 20/01/2025
Most guides on RAG systems are over-optimizing for latency. If you are using RAG to automate any type of knowledge work, accuracy is much more important. I explore the idea of spending more compute for RAG systems to significantly improve performance. www.vivekkalyan.com/writing/scal...
vivekkalyan.com
Scaling Compute for RAG
Unlock higher RAG accuracy by strategically spending compute.
110
Vivek Kalyan @vivekkalyan.com · 11/01/2025
This really resonates with me. I spent 6 years building the AI products at my previous company and felt first hand the consequences of decisions (both good and bad). I've much more opinions on how to do certain things because of it.
020
Vivek Kalyan @vivekkalyan.com · 11/01/2025
Try using a local UI (I'm using msty.app), and call Claude via API. Fits my bursty nature of usage, and monthly cost is much cheaper generally. You do lose their system prompt + Artifacts but I don't miss it much.
msty.app
Msty - Using AI Models made Simple and Easy
AI beyond just plain chat. Private, Offline, Split chats, Branching, Concurrent chats, Web Search, RAG, Prompts Library, Vapor Mode, and more. Perfect LM Studio, Jan AI, and Perplexity alternative. Us...
120
Vivek Kalyan @vivekkalyan.com · 19/12/2024
This is a super high impact project. There are tons of production models in the real world still running BERT/RoBERTa models from the 2018-2019 era, I'm sincerely hoping these models are easy to finetune, just the 8k context length will be good enough reason to upgrade.
010
Vivek Kalyan @vivekkalyan.com · 12/12/2024
They are the victims of their own media success, they raised at ridiculous valuation and they need to show that they disrupting the software engineering market beyond something like copilot/cursor.
000
Vivek Kalyan @vivekkalyan.com · 08/12/2024
Do you have the link to the public repo? Java is highly requested, we just have been focusing on validating the usefulness of the current docs before spending time adding new languages.
100
Vivek Kalyan @vivekkalyan.com · 08/12/2024
Cartograph solves this by automatically creating and updating both written documentation and visual representations that stay synchronized with code changes. We are very early, but there's much more to ship. We would love to get some feedback on our product.
010
Vivek Kalyan @vivekkalyan.com · 08/12/2024
In complex systems, we spend most of our time understanding code - but most AI tools are focused on helping you write more code. Engineering teams struggle with outdated documentation and inconsistent system diagrams. This slows down development and makes it difficult to understand complex systems.
110
Vivek Kalyan @vivekkalyan.com · 08/12/2024
We're launching early access for Cartograph (cartograph.app). It takes your codebase, and automatically generates architecture diagrams and documentation for them. See some demos of open-source repos here (no sign-up required): cartograph.app/demo (Reply here if you'd like to see others added)
220
Vivek Kalyan @vivekkalyan.com · 07/12/2024
Spent more than 2 hours yesterday night trying to figure out why my solution for day 6 part 2 was working on the examples but not on the test input. Slept on it, and figured it out in 5mins when I woke up today morning. 🤦‍♂️
000
Vivek Kalyan @vivekkalyan.com · 04/12/2024
Congrats on the move! And your first 🦋 post.
000
Vivek Kalyan @vivekkalyan.com · 04/12/2024
Yeah, hot mess is the right description of it 😂. Would be interested to see if there is a nicer solution.
010
Vivek Kalyan @vivekkalyan.com · 03/12/2024
Yeah, I've been writing rust for a few months now for parsing codebase into a graph (cartograph.app). But, it's also a conscious decision to force myself to write "idiomatic Rust", i.e. more functional.
cartograph.app
Cartograph | AI knowledge base for your code
110
Vivek Kalyan @vivekkalyan.com · 01/12/2024
Oh look, it's that time of the year. I will doing Advent of Code 2024 in Rust, hoping to get more experience using Rust to solve a wide range of problems. github.com/vivekkalyan/...
github.com
GitHub - vivekkalyan/advent-of-code-2024: Solving Advent of Code in Rust
Solving Advent of Code in Rust. Contribute to vivekkalyan/advent-of-code-2024 development by creating an account on GitHub.
340
Vivek Kalyan @vivekkalyan.com · 29/11/2024
Do you have sources for the budget claims? I have not seen any numbers comparing models so would be interested to know.
100
Vivek Kalyan @vivekkalyan.com · 27/11/2024
@eugeneyan.bsky.social's blog is a gold mine if you are doing ML/AI in the industry. Writing design docs before ML projects start is an important process that I introduced at my prev org, and the post on design docs was one of the references I used to create our template.
161
Vivek Kalyan @vivekkalyan.com · 27/11/2024
Methodology paper detailing how to train small, efficient off-topic classifiers using synthetic data from LLMs.
010
Vivek Kalyan @vivekkalyan.com · 26/11/2024
Yeah, your feed is the training data for your brain. Curate it for what you want to see more of
010
Vivek Kalyan @vivekkalyan.com · 26/11/2024
How does using an assert statement work in practice? Are you using something like pytest as your benchmark runner?
100
Vivek Kalyan @vivekkalyan.com · 26/11/2024
Very disappointed your bio doesn't say learning machine.
110
Vivek Kalyan @vivekkalyan.com · 26/11/2024
The Gemini team really delivered with gemini-exp-1121. It slaps for writing tasks, it's outputs avoid the typical AI feel that other models (like ChatGPT) have, something previously only Claude achieved.
030
Vivek Kalyan @vivekkalyan.com · 26/11/2024
Good guidelines. One thing I want to add is if you are benchmarking your system's real world performance - you probably want to keep a version of the benchmark without the easy examples removed. Especially if you need to report to non-technical management how your system is performing.
010
Vivek Kalyan @vivekkalyan.com · 25/11/2024
The most on-pulse demo would be on the bluesky firehose data.
020
Vivek Kalyan @vivekkalyan.com · 25/11/2024
Models will follow the order of the schema you specify! So just putting the COT field first works well in practice. Structured outputs also ensure that the model response can be parsed as json. Use Pydantic/Zod here.
110
Vivek Kalyan @vivekkalyan.com · 25/11/2024
I'm building cartograph.app, an AI powered platform that helps software teams save time with documentation and architecture diagrams on codebases. Just recently added the ability to extract and link HTTP routes to find dependencies across services. Supports Python/JS/TS/Rust now.
cartograph.app
Cartograph | AI knowledge base for your code
020
Vivek Kalyan @vivekkalyan.com · 25/11/2024
I've never been this excited about social media
010
Vivek Kalyan @vivekkalyan.com · 25/11/2024
I'm really tempted to update my blog...
000
Vivek Kalyan @vivekkalyan.com · 25/11/2024
"The bar after a conference" is a great description for what many of us are craving for.
000
Vivek Kalyan @vivekkalyan.com · 25/11/2024
Unless you absolutely need Claude Artifacts, you can try an open source UI and calling the API directly. It comes to be way cheaper than $20/month and doesn't have rate limits that even Claude pro tier has (had?).
120
Vivek Kalyan @vivekkalyan.com · 24/11/2024
Tagging the authors: @natolambert.bsky.social @jacobcares.bsky.social @valentinapy.bsky.social @hamishivi.bsky.social @alisawuffles.bsky.social @nouhadziri.bsky.social @soldaini.net @nlpnoah.bsky.social @yizhongw.bsky.social @pdasigi.bsky.social
010
Vivek Kalyan @vivekkalyan.com · 24/11/2024
This took me more time than expected, hope it's useful for other people! I really appreciate the open science work being done by allenai, and this was a fun weekend evening read.
100
Vivek Kalyan @vivekkalyan.com · 24/11/2024
And I really like this section: what didn't work. Both online DPO (training a reward model is hard) and rejection sampling (sampling n responses, choosing the best and using the rest as rejected) didn't improve performance.
100
Vivek Kalyan @vivekkalyan.com · 24/11/2024
They also introduce more test datasets which use LLM-as-a-Judge - really seeming to be a common theme here. HREF uses human reference as a guide for LLM-as-a-Judge, allowing for a test set that can be used for open ended tasks such as summarization and rewriting.
100
Vivek Kalyan @vivekkalyan.com · 24/11/2024
OLMES - they release the Open Language Model Evaluation System which is an evaluation suite for datasets similar to lm-evaluation-harness, but with more support for datasets, configurations, detailed output data for analysis. github.com/allenai/olmes
github.com
GitHub - allenai/olmes: Reproducible, flexible LLM evaluations
Reproducible, flexible LLM evaluations. Contribute to allenai/olmes development by creating an account on GitHub.
100
Vivek Kalyan @vivekkalyan.com · 24/11/2024
PPO - replaces the standard reward model with a verification function when the answer can be verified correct (e.g., in mathematics). The model is trained using PPO, only getting a reward if the answer is correct. This works really well to improve domains which have verifiable answers.
100
Vivek Kalyan @vivekkalyan.com · 24/11/2024
DPO ablations - reducing memory requirements by caching log probs across the dataset instead of computing on the fly with a reference model and separating forward passes for chosen and rejected sequences.
100
Vivek Kalyan @vivekkalyan.com · 24/11/2024
DPO ablations - doing PPO is hard, primarily because of the reward model, which is hard to evaluate. They conclude that **for the budget** DPO works well enough for preference tuning.
100
Vivek Kalyan @vivekkalyan.com · 24/11/2024
DPO ablations - there is much focus on ensuring model follows instruction, mostly leveraging synthetic data methods to create {chosen, rejected} pairs. They also measure that you can take an existing dataset, regenerate it using their synthetic pipeline and it improves performance. Really cool.
110
Vivek Kalyan @vivekkalyan.com · 24/11/2024
DPO ablations - on-policy data (generations from the base SFT model) improve model performance. Although, it's interesting to note that they don't seem to train the model for multiple rounds like Llama 3.1 (maybe cost concerns?)
110
Vivek Kalyan @vivekkalyan.com · 24/11/2024
DPO ablations - the number of unique prompts in the preference dataset matter. More unique prompts result in better downstream performance, while more generations for the same prompts does not. Using different prompts from SFT data also improve performance.
110
Vivek Kalyan @vivekkalyan.com · 24/11/2024
It's interesting how well LLM-as-a-judge worked for them here, it's completely viable now to evaluate your data using LLMs (even for seemingly fuzzy tasks like preference). related: aligned.substack.com/p/crisp-and-...
aligned.substack.com
Crisp and fuzzy tasks
Why fuzzy tasks matter and how to align models on them
110
Vivek Kalyan @vivekkalyan.com · 24/11/2024
DPO - preference data is made of a on-policy model pool (Tulu 3 7B/70B) and off-policy model pool (other models). GPT-4o is used to rate generations from 4 random models from the model pool from 1-5. The highest is taken as the chosen response and the rejected response is sampled from the rest.
110
Vivek Kalyan @vivekkalyan.com · 24/11/2024
SFT - performance varies greatly depending on the random seed. They tried model soups (averaging weights of multiple models using github.com/arcee-ai/mer...), but decided to choose a single best model instead.
100
Vivek Kalyan @vivekkalyan.com · 24/11/2024
SFT - using a better pretrained model (Qwen 2.5) results in better scores for GSM8K and MATH. Kinda surprising since the final model is still Llama 3.1, wonder if they evaluated Qwen 2.5 on other datasets as well and if the performance was not good enough to be the base model?
100
Vivek Kalyan @vivekkalyan.com · 24/11/2024
SFT - a couple of data ablations: - Adding safety data didn't negatively affect other skills. - Measured the removal of math data on GSM8K and MATH, confirming that the scores drop. - Measured the the effectiveness of WildChat (diverse chat data) and Persona (skill based data)
100
Vivek Kalyan @vivekkalyan.com · 24/11/2024
SFT - for the data, train models on skill specific data mixtures, keep the mixtures that lead to the best performance on those skills. Later, combine the mixtures and do decontamination/downsampling of larger datasets.
100
Vivek Kalyan @vivekkalyan.com · 24/11/2024
Data decontamination - n-gram matching works better than embedding methods in terms of precision - 8-gram matching only on user-turns (since completions are LLM generated) - remove training sets that have more than 2% overlap with any evaluation set
100
Vivek Kalyan @vivekkalyan.com · 24/11/2024
Data sources - they combat mode collapse using persona-driven synthetic data generation, where unique personas (e.g., "ML researcher focused on neural networks") guide LLMs to create diverse training data for specific skills like precise instruction following, math, and coding.
100
Vivek Kalyan @vivekkalyan.com · 24/11/2024
Data sources - they curate from public datasets based on three criteria - diversity (for generalization), targeting their identified core skills and clear licensing/provenance for proper usage. Datasets they considered: docs.google.com/spreadsheets...
docs.google.com
Big Instruction Compilation
100