elvis @eos.bsky.social · 24/01/2025Hallucinations are generally bad for real-world LLM applications. Folks like Karpathy have suggested that hallucination is an LLM's greatest feature. Is there any evidence for the latter? 221
elvis @eos.bsky.social · 06/01/2025Google recently published this great whitepaper on Agents. 2025 is a huge year for AI Agents. Here's what's included: - Introduction to AI Agents - The role of tools in Agents - Enhancing model performance - Quick start to Agents with LangChain - Production applications with Vertex AI Agents 110
elvis @eos.bsky.social · 06/01/2025Dive into Time-Series Anomaly Detection: A Decade Review Provides a survey of time-series anomaly detection solutions. 100
elvis @eos.bsky.social · 02/01/2025Proposes a self-training strategy to mitigate overthinking in o1-like LLMs. This approach can reduce token output by 48.6% while maintaining accuracy on the widely-used MATH500 test set as applied to QwQ-32B-Preview. There are three parts to this study: - analysis of the overthinking issue 120
elvis @eos.bsky.social · 02/01/2025Just came across this interesting project. RWKV combines the best of RNN and transformers. There is code for training your own model, fine-tuning, GUI, API, fast WebGPU inference, and more. Cool project! 110
elvis @eos.bsky.social · 31/12/2024Agentarium is a new Python framework for managing and orchestrating AI agents. Features include (from the repo): • 🤖 Advanced Agent Management: Create and orchestrate multiple AI agents with different roles and capabilities 110
elvis @eos.bsky.social · 20/12/2024A Survey of Mathematical Reasoning in the Era of Multimodal LLMs arxiv.org/abs/2412.11936 000
elvis @eos.bsky.social · 13/12/2024Very exciting guide on working with LLMs. A few chapters have been made available already like structured output and evals. More to come soon! 110
elvis @eos.bsky.social · 11/12/2024IBM open-sources Granite Guardian, a suite of safeguards for risk detection in LLMs. 151
elvis @eos.bsky.social · 10/12/2024A Survey on LLMs-as-Judges Presents a comprehensive survey of the LLMs-as-judges paradigm from five key perspectives: Functionality, Methodology, Applications, Meta-evaluation, and Limitations. 110
elvis @eos.bsky.social · 10/12/2024A Multi-Agent Framework for Synthetic Data Generation Presents MAG-V, a multi-agent framework that first generates a dataset of questions that mimic customer queries. It then reverse engineer alternate questions from responses to verify agent trajectories. arxiv.org/abs/2412.04494 110
elvis @eos.bsky.social · 09/12/2024Reinforcement Learning: An Overview This gem just dropped on arXiv. An up-to-date overview of reinforcement learning and sequential decision-making. arxiv.org/abs/2412.05265 020
elvis @eos.bsky.social · 08/12/2024It was a huge week of AI and LLM papers. Here are the top ML Papers of the Week (Dec 2-8): - Genie 2 - GenCast - OpenAI o1 - Auto-RAG - Reverse Thinking - Retrieval-Augmented Reasoning for LLMs Read on for more: nlp.elvissaravia.com/p/top-ml-pap...nlp.elvissaravia.com🥇Top ML Papers of the WeekThe Top ML Papers of the Week (December 2 - 8) 060
elvis @eos.bsky.social · 05/12/2024RARE: Retrieval-Augmented Reasoning for LLMs Extends the rStar reasoning framework to enhance reasoning accuracy and factual reliability of LLMs. 110
elvis @eos.bsky.social · 05/12/2024Nice set of tips for mitigating AI hallucinations. www.kapa.ai/blog/ai-hall... 030
elvis @eos.bsky.social · 04/12/2024DataLab: A Unified Platform for LLM-Powered Business Intelligence Introduces DataLab, a unified BI platform that integrates an LLM-based agent framework with an augmented computational notebook interface. 140
elvis @eos.bsky.social · 02/12/2024Steel is an open-source browser API for AI agents and apps. Useful to build AI apps and agents that interact with the web. Features include: - Full browser control - Session management - Proxy support - Extension support - Debugging tools - Anti-detection - Resource management 140
elvis @eos.bsky.social · 01/12/2024It was a huge week of AI and LLM papers. So we summarized the top trending papers covering a new efficient attention mechanism, LLM-as-a-Judge survey, improving reasoning, GUI Agents, and more. Here are the top papers of the week and key insights: nlp.elvissaravia.com/p/top-ml-pap...nlp.elvissaravia.com🥇Top ML Papers of the WeekThe Top ML Papers of the Week (November 25 - December 1) 040
elvis @eos.bsky.social · 28/11/2024LLM-brained GUI Agents Great survey on LLM-brained GUI agents, including techniques and applications. arxiv.org/abs/2411.18279 041
elvis @eos.bsky.social · 28/11/2024This new paper extends in-context learning through high-level automated reasoning. It achieves state-of-the-art accuracy (79.6%) on the MATH benchmark with Qwen2.5-7B-Instruct, surpassing GPT-4o (76.6%) and Claude 3.5 (71.1%). 122
elvis @eos.bsky.social · 27/11/2024LLMs surpass human experts in predicting neuroscience results Scientific discovery is the next big goal for AI. We are seeing a huge number of research studies tackling AI-powered scientific discovery from different angles and for different problems. 140
elvis @eos.bsky.social · 26/11/2024o1 Replication Journey - Part 2 Shows that combining simple distillation from O1's API with supervised fine-tuning significantly boosts performance on complex math reasoning tasks. 130
elvis @eos.bsky.social · 25/11/2024Pushing Frontiers in Open Language Model Post-Training This is probably one of the important open-source efforts in post-training of LLMs. 120
elvis @eos.bsky.social · 25/11/2024Measuring Bullshit in the Language Games played by ChatGPT Proposes that LLM-based chatbots play the ‘language game of bullshit’ 100
elvis @eos.bsky.social · 24/11/2024Cheaper and larger context-length text embedding models. Voyage AI releases voyage-3 and voyage-3-lite embedding models. 151
elvis @eos.bsky.social · 24/11/2024👋Hi! Anyone here? I am starting to post here and looking to connect with the AI community. 040
elvis @eos.bsky.social · 24/11/2024Top ML Papers of the Week (Nov 18 - 24): - FinRobot - Bi-Mamba - AlphaQubit - BABY-AIGS - The Dawn of GUI Agent - Agents for Automated Bug Fixing Read on for more: nlp.elvissaravia.com/p/top-ml-pap...nlp.elvissaravia.com🥇Top ML Papers of the WeekThe Top ML Papers of the Week (November 18 - 24) 051
elvis @eos.bsky.social · 17/05/2023Reasoning over Structured Data with LLMs Proposes StructGPT to improve the zero-shot reasoning ability of LLMs over structured data. Effective for solving question answering tasks based on structured data. paper: arxiv.org/abs/2305.09645 020
elvis @eos.bsky.social · 14/05/2023Top ML Papers of the Week (May 8 - 14): - PaLM 2 - ImageBind - InstructBLIP - StarCoder - MultiModal-GPT - LLM explains neurons in LLMs ... open.substack.com/pub/nlpnews/p/top… 021
elvis @eos.bsky.social · 10/05/2023Wow! I just found this open-source implementation of PaLM models. Models are trained with 8K context length on all of C4. Different PaLM model sizes available: 150m, 410m, and 1B. Apparently, a 2B parameter model and instruction-tuned models are on the way! github.com/conceptofmind/PaLM 010
elvis @eos.bsky.social · 08/05/2023Can LLMs transform Computational Social Science? This new paper provides an analysis and road map for using LLMs as computational social science tools, including prompting best practices and an evaluation pipeline. arxiv.org/abs/2305.03514 020
elvis @eos.bsky.social · 07/05/2023Top ML Papers of the Week (May 1 - 7): - scGPT - GPTutor - PMC-LLaMA - Unlimiformer - Distilling Step-by-Step - Are Emergent Abilities of LLMs a Mirage? ... open.substack.com/pub/nlpnews/p/top…open.substack.com🥇Top ML Papers of the WeekThe top ML Papers of the Week (May 1 - May 7) 031
elvis @eos.bsky.social · 02/05/2023A Review of ChatGPT Applications A review of ChatGPT applications in education, marketing, software engineering, and healthcare. Includes discussion around benefits, drawbacks, and research directions. arxiv.org/abs/2305.00237 031
elvis @eos.bsky.social · 02/05/2023Who else is building and experimenting with LLMs here? Say hello 👋 020
elvis @eos.bsky.social · 01/05/2023Finetuning LLaMA on Medical Papers - fine-tunes LLaMA on 4.8m biomedical papers - enhances capabilities in the medical domain - the proposed model, PMC-LLaMA, achieves high performance on biomedical QA benchmarks paper: arxiv.org/abs/2304.14454 code: github.com/chaoyi-wu/PMC-LLaMA 031
elvis @eos.bsky.social · 30/04/2023Top ML Papers of the Week (April 24 - 30): - AudioGPT - Track Anything - Agents Learn Soccer Skills - Harnessing the Power of LLMs - Scaling Transformer to 1M tokens - A Cookbook of Self-Supervised Learning ...open.substack.com🥇Top ML Papers of the WeekThe top ML Papers of the Week (April 24 - 30) 042