Ksenia Se / Turing Post @turingpost.bsky.social · 11/12/2024"Densing Law of LLMs" paper: arxiv.org/abs/2412.04315 011
Ksenia Se / Turing Post @turingpost.bsky.social · 11/12/2024Here are the key findings from the study: • Costs to run models are dropping as they are becoming more efficient. • The release of ChatGPT sped up the growth of efficiency of new models up to 50%! • Techniques like pruning and distillation don’t necessarily make models more efficient. 100
Ksenia Se / Turing Post @turingpost.bsky.social · 11/12/2024Reading about scaling laws recently I came by the interesting point: Focus on a balance between models' size and performance is more important that aiming for larger models Tsinghua University and ModelBest Inc propose the idea of “capacity density” to measure how efficiently a model uses its size 124
Ksenia Se / Turing Post @turingpost.bsky.social · 10/12/20241. GoogleDeepMind's Genie 2 Generates 3D environments with object interactions, animations, and physical effects from one image or text prompt. You can interact with them in real-time using a keyboard and mouse. Paper: deepmind.google/discover/blo... Our example: www.youtube.com/watch?v=YjO6... 100
Ksenia Se / Turing Post @turingpost.bsky.social · 10/12/2024An incredible shift is happening in spatial intelligence! Here are 2 latest revolutional World Models, which create interactive 3D environments: 1. GoogleDeepMind's Genie 2 2. AI system from World Labs, co-founded by Fei-Fei Li Explore more below 👇 120
Ksenia Se / Turing Post @turingpost.bsky.social · 05/12/2024What is Flow Matching? Flow Matching (FM) is used in top generative models, like Flux, F5-TTS, E2-TTS, and MovieGen with state-pf-the-art results. Some experts even say that FM might surpass diffusion models👇 100
Ksenia Se / Turing Post @turingpost.bsky.social · 05/12/2024INTELLECT-1 by Prime Intellect INTELLECT-1 is a 10B open-source LLM trained over 42 days on 1T tokens across 14 global nodes, leverages the PRIME framework for exceptional efficiency (400× bandwidth reduction). github.com/PrimeIntelle... 100
Ksenia Se / Turing Post @turingpost.bsky.social · 05/12/2024MultiFoley by Adobe Research MultiFoley is an AI model generating high-quality sound effects from text, audio, and video inputs. Cool demos highlight its creative potential. arxiv.org/abs/2411.17698 100
Ksenia Se / Turing Post @turingpost.bsky.social · 05/12/2024ShowUI by Show Lab, NUS, Microsoft ShowUI is a 2B vision-language-action model tailored for GUI tasks: - features UI-guided token selection (33% fewer tokens) - interleaved streaming for multi-turn tasks - 256K dataset - achieves 75.1% zero-shot grounding accuracy arxiv.org/abs/2411.17465 100
Ksenia Se / Turing Post @turingpost.bsky.social · 05/12/2024OLMo 2 by Allen AI OLMo 2, a family of fully open LMs with 7B and 13B parameter, is trained on 5 trillion tokens. allenai.org/blog/olmo2 110
Ksenia Se / Turing Post @turingpost.bsky.social · 05/12/2024Amazing models of the week: • Alibaba’s QwQ-32B • OLMo 2 by Allen AI • ShowUI by Show Lab, NUS, Microsoft • Adobe's MultiFoley • INTELLECT-1 by Prime Intellect 🧵 100
Ksenia Se / Turing Post @turingpost.bsky.social · 02/12/2024Boundless Socratic Learning with Language Games, Google DeepMind This framework leverages recursive language-based "games" for self-improvement, focusing of feedback, coverage, and scalability. It suggests a roadmap for scalable AI via autonomous data gen and feedback loops arxiv.org/abs/2411.16905 110
Ksenia Se / Turing Post @turingpost.bsky.social · 02/12/2024MH-MoE: Multi-Head Mixture-of-Experts @msftresearch.bsky.social’s MH-MoE improves sparse MoE by adding multi-head attention, reducing perplexity without increasing FLOPs, and demonstrating robust performance under quantization. arxiv.org/abs/2411.16205 100
Ksenia Se / Turing Post @turingpost.bsky.social · 02/12/2024LLM-as-a-Judge: Presents a taxonomy of methodologies and applications of LLMs for judgment tasks, highlighting bias, vulnerabilities, and self-judgment, with future directions in human-LLM collaboration and bias mitigation arxiv.org/abs/2411.16594 100
Ksenia Se / Turing Post @turingpost.bsky.social · 02/12/2024Star Attention: NVIDIA introduced a block-sparse attention mechanism for Transformer-based LLMs. It uses local/global attention phases to achieve up to 11x inference speedup on sequences up to 1M tokens, retaining 95-100% accuracy. arxiv.org/abs/2411.17116 Code: github.com/NVIDIA/Star-... 100
Ksenia Se / Turing Post @turingpost.bsky.social · 02/12/2024Top 5 researches of the week: • Natural Language Reinforcement Learning • Star Attention, NVIDIA • Opportunities and Challenges of LLM-as-a-judge • MH-MoE: Multi-Head Mixture-of-Experts, @msftresearch.bsky.social • Boundless Socratic Learning with Language Games, Google DeepMind 🧵 120
Ksenia Se / Turing Post @turingpost.bsky.social · 02/12/2024Paper: Large Language Model-Brained GUI Agents: A Survey arxiv.org/abs/2411.18279 100
Ksenia Se / Turing Post @turingpost.bsky.social · 02/12/20242. Prompt engineering or creating a plan: After "looking" at the screen, the agent prepares a prompt for the AI model, which includes the user’s instructions, the visual data (like screenshots or buttons layout), and other context the agent needs to understand the task. 100
Ksenia Se / Turing Post @turingpost.bsky.social · 02/12/20241. Understanding the environment: Firstly, the agent needs to "see" the software it’s working with. This is done through methods that capture the layout of the app or website, such as screenshots, lists of buttons and menus called widget trees. 100
Ksenia Se / Turing Post @turingpost.bsky.social · 02/12/2024LLM-brained GUI agents are a way to interact with GUIs in a much more flexible and human-like way. They blend LLMs' capabilities with software interaction to work with websites, mobile apps, and desktop software, simplifying complex tasks. A new survey on LLM-brained GUI agents was published👇 111
Ksenia Se / Turing Post @turingpost.bsky.social · 01/12/2024Our TuringPost Twitter Library provides useful resources, such as lists of tools and models, papers and courses on various popular aspects of AI and machine learning for you every week. Check it out. www.turingpost.com/t/Twitter-Li... 100
Ksenia Se / Turing Post @turingpost.bsky.social · 01/12/2024Top 10 GitHub Repositories to master ML, AI and Data Science: • 100 Days of ML Code • Data Science For Beginners • Awesome Data Science • Data Science Masters • Homemade Machine Learning • 500+ AI Projects List with Code • Awesome Artificial Intelligence ... Check out for more👇 121
Ksenia Se / Turing Post @turingpost.bsky.social · 30/11/2024Benefits of HiAR-ICL: - Thanks to thought cards, HiAR-ICL requires less human input to guide it and adapts better to different types of problems -Faster by focusing only on actions from thought card - Accuracy: It achieved 79.6% accuracy on a math benchmark, compared to GPT-4o’s 76.6%. 100
Ksenia Se / Turing Post @turingpost.bsky.social · 30/11/2024Thought Cards with Monte Carlo Tree Search (MCTS): Thought cards are like reusable problem-solving templates. After MCTS generates multiple possible solution paths for each problem, HiAR-ICL picks the best one by weighing the paths using Value of Computation (VOC). 110
Ksenia Se / Turing Post @turingpost.bsky.social · 30/11/2024This new approach to In-Context Learning (ICL) is very interesting: The idea of HiAR-ICL (High-level Automated Reasoning in ICL) is to teaches the model abstract thinking patterns, using atomic reasoning actions and thought cards. It doesn't rely only on examples and prompts like in ICL Details👇 120
Ksenia Se / Turing Post @turingpost.bsky.social · 29/11/2024What is Self-Supervised Learning? Whether you are a professional or a beginner in Machine Learning, we made our flashcards to be easy to digest and help everyone refresh key ML concepts. Here's Self-Supervised Learning, a technique used to train ML models👇 100
Ksenia Se / Turing Post @turingpost.bsky.social · 29/11/2024In NLRL, LLMs take on different roles like: • They critique and plan using prompts (for example, explain their choices). • Learn to evaluate actions using language (act as a validator). • They are "actors" (decision-makers) and "critics" (evaluators) in a feedback loop 110
Ksenia Se / Turing Post @turingpost.bsky.social · 29/11/20242. Language Temporal-Difference Estimate: Uses the language Bellman equation to breaks down the value of a state into two parts: the immediate reward and the value of the next state. 100
Ksenia Se / Turing Post @turingpost.bsky.social · 29/11/20241. Language Monte-Carlo Estimate: Uses a language aggregator tool to combine the information into a summary value. 100
Ksenia Se / Turing Post @turingpost.bsky.social · 29/11/2024Natural Language Reinforcement Learning (NLRL) redefines Reinforcement Learning (RL). NLRL's main idea: The core parts of RL like goals, strategies, and evaluation methods are reimagined using natural language instead of rigid math. Let's explore this approach more precisely🧵 121
Ksenia Se / Turing Post @turingpost.bsky.social · 29/11/2024A quick reminder of what AI agents are built of: 1. Profiling: Keeps the agent aligned with its purpose. 2. Knowledge: Provides domain-specific expertise. Find other components below👇 P.S.: As an example, here's LangChain’s Harrison Chase's agent framework. 100
Ksenia Se / Turing Post @turingpost.bsky.social · 28/11/2024AI should be for everyone, right? :) Here are 6 FREE AI courses for beginners: • Introduction to Artificial Intelligence • Artificial Intelligence for Beginners • AI For Everyone • Machine Learning for Beginners • Introduction to Data Science Specialization • Data Science for Beginners Links👇 140
Ksenia Se / Turing Post @turingpost.bsky.social · 28/11/2024Does LLaVA-o1 VLM challenge OpenAI's o1 model? Here's what enhances LLaVA-o1's complex multimodal reasoning: - Reasoning in 4 stages - Effective inference-time scaling - Stage-level beam search - Special dataset with step-by-step reasoning examples 👇 100
Ksenia Se / Turing Post @turingpost.bsky.social · 28/11/2024Hymba's performance: Hymba-1.5B achieves 1.32% higher accuracy than Llama-3.2-3B on commonsense reasoning tasks. It reduces memory cache size by 11.67x and achieves 3.49x faster throughput compared to Llama-3.2-3B. 110
Ksenia Se / Turing Post @turingpost.bsky.social · 28/11/2024Hymba's attention map combines 3 key elements: - Meta tokens: Offloads attention from the "beginning-of-sequence" (BOS) token, and the model focuses more on real input tokens. - Sliding window attention - SSMs: Summarizes the overall context while focusing on current tokens. 100
Ksenia Se / Turing Post @turingpost.bsky.social · 28/11/2024Meta tokens: They are special, learnable tokens added at the input's start to help Hymba focus on key information. They create a modified sequence where the tokens act as a guide for the attention mechanism. During inference, the meta tokens are fixed, and their computations can be preprocessed. 100
Ksenia Se / Turing Post @turingpost.bsky.social · 28/11/2024Hybrid-head architecture: Hymba combines attention and SSM layers in parallel within the same layer, unlike existing hybrid models that use separate layers for them. This makes Hymba flexible and effective across different tasks. 100
Ksenia Se / Turing Post @turingpost.bsky.social · 28/11/2024NVIDIA's Hymba small model is a great combo of 2 concepts: - Transformer attention to help the model remember details. - State Space Models (SSMs) to efficiently summarize context. This model is also a treasure trove of interesting features. Here are the details: 200
Ksenia Se / Turing Post @turingpost.bsky.social · 28/11/2024A Flexible Large Language Models Guardrail Development Methodology Applied to Off-Topic Prompt Detection Proposes a guardrail methodology for detecting off-topic prompts, utilizing synthetic datasets and fine-tuning for robust performance. arxiv.org/abs/2411.12946 huggingface.co/collections/... 100
Ksenia Se / Turing Post @turingpost.bsky.social · 28/11/2024Do I Know This Entity? Knowledge Awareness and Hallucinations in Language Models Identifies directions in model representations to mitigate hallucinations and refine entity recognition. arxiv.org/abs/2411.14257 100
Ksenia Se / Turing Post @turingpost.bsky.social · 28/11/2024Building Trust: Foundations of Security, Safety, and Transparency in AI Establishes safety and security frameworks for responsible AI development and standardized risk management. arxiv.org/abs/2411.12275 100
Ksenia Se / Turing Post @turingpost.bsky.social · 28/11/2024Search, Verify and Feedback: Towards Next Generation Post-training Paradigm of Foundation Models via Verifier Engineering Proposes a post-training paradigm for foundation models, enhancing scalability and alignment with verification techniques. arxiv.org/abs/2411.11504 github.com/icip-cas/Ver... 100
Ksenia Se / Turing Post @turingpost.bsky.social · 28/11/2024One to Rule Them All: Natural Language to Bind Communication, Perception, and Action Introduces a robotic architecture integrating LLMs for task execution and dynamic environmental adaptation. arxiv.org/abs/2411.15033 110
Ksenia Se / Turing Post @turingpost.bsky.social · 28/11/2024Loss-to-Loss Prediction: Scaling Laws for All Datasets Develops predictive scaling laws for model performance across tasks and datasets, enhancing efficiency and planning. arxiv.org/abs/2411.12925 Models: huggingface.co/KempnerInsti... 100
Ksenia Se / Turing Post @turingpost.bsky.social · 28/11/2024When Precision Meets Position: BFloat16 Breaks Down RoPE in Long-Context Training Optimizes long-context processing by resolving precision issues, accelerating training while preserving performance. arxiv.org/abs/2411.13476 Code: github.com/haonan3/Anch... 100
Ksenia Se / Turing Post @turingpost.bsky.social · 28/11/2024Ultra-Sparse Memory Network Introduces ultra-sparse architectures that reduce latency and improve memory efficiency, rivaling larger dense models. arxiv.org/abs/2411.12364 100
Ksenia Se / Turing Post @turingpost.bsky.social · 28/11/2024Generative World Explorer Develops Genex for imaginative exploration in 3D environments, enabling decision-making and belief revision without physical movement. arxiv.org/abs/2411.11844 100
Ksenia Se / Turing Post @turingpost.bsky.social · 28/11/2024Drowning in Documents: Consequences of Scaling Reranker Inference Analyzes challenges of scaling rerankers for large datasets, showing limitations and proposing robust listwise reranking alternatives. arxiv.org/abs/2411.11767 100
Ksenia Se / Turing Post @turingpost.bsky.social · 28/11/2024Openscholar: Synthesizing Scientific Literature With Retrieval-Augmented LMs Builds a retrieval-augmented LLM, outperforming GPT-4o in synthesizing scientific queries from a large literature dataset. arxiv.org/abs/2411.14199 Code: github.com/AkariAsai/Op... 100