Adina Yakup @adinayakup.bsky.social · 21/09/2026XiaomiMiMo just released 2 SoTA models. One might be the new BEST open model yet🔥 huggingface.co/collections/... Both are: - Sparse MoE + 1M context + MIT licensed - Native omni: text/image/video/audio 1181
Adina Yakup @adinayakup.bsky.social · 19/09/2026HyperFlow⚡New MiniMax-H3 variant from Video Rebirth huggingface.co/videorebirth... - Open weight 8 step LoRA - 3× faster with data free self distillation - Video + stereo audio intact 061
Adina Yakup @adinayakup.bsky.social · 10/09/2026DeepSeek v4.1 Flash is just another level 🤯 huggingface.co/deepseek-ai/... - Asymmetric Causal-Encoder-Decoder: 550B MoE, input 8B / output 16B - Native vision merged into one endpoint - KV cache crushed: ~1/4 the HBM vs last one, 437× smaller than their first modelhuggingface.codeepseek-ai/DeepSeek-V4.1-Flash · Hugging FaceWe’re on a journey to advance and democratize artificial intelligence through open source and open science. 2754
Adina Yakup @adinayakup.bsky.social · 09/09/2026Ling-3.0-flash-VL just dropped from Ant Group huggingface.co/inclusionAI/... - Native image + video: understand > reason > act > verify - 124B/5.5B active - 1M context - MIT licensedhuggingface.coinclusionAI/Ling-3.0-flash-VL · Hugging FaceWe’re on a journey to advance and democratize artificial intelligence through open source and open science. 0121
Adina Yakup @adinayakup.bsky.social · 08/09/20263 new open agentic models from Nex-AGI just dropped on @hf.co 🔥 Nex-N2.5: mini / Pro / Max - mini: 35B, tool-calling on 2×H100 - Pro: 397B hybrid-attention MoE, single 8×H100 node ( weights coming soon ) - Max: 1.6T, MoE - All Apache 2.0 huggingface.co/collections/...huggingface.coNex-N2.5 - a nex-agi CollectionWe’re on a journey to advance and democratize artificial intelligence through open source and open science. 020
Adina Yakup @adinayakup.bsky.social · 07/09/2026When everyone is releasing big models, OpenBMB keeps shipping small but strong one. Here is the latest one: MiniCPM5-2B 🔥 huggingface.co/openbmb/Mini... - Dense 2B - Full data pipeline open as well: web, code, agent, RL - 131 context - Apache 2.0 3240
Adina Yakup @adinayakup.bsky.social · 04/09/2026Ling-3.0-flash-Fin 💰 A finance tuned build of Ling-3.0-flash from Ant Group huggingface.co/collections/... - 124B/ 5.1B active - 256K context - MIT license - Finance native agent workflows - Already being quantized for llama.cpp 050
Adina Yakup @adinayakup.bsky.social · 10/07/2026LingBot-Video 🎬 MoE video model built for embodied AI from Ant group huggingface.co/collections/... - 30B/3B - Apache 2.0 - Trained on web videos + 70K hours of embodied data - Tops RBench: ahead of Cosmos3/Veo 3/Seedance 1.5 pro 071
Adina Yakup @adinayakup.bsky.social · 07/07/2026Agents-A1 🤖🔬 New agentic model from Shanghai AI Lab, InternScience team huggingface.co/collections/... - 35B MoE (built on Qwen3.5-35B-A3B) - Apache 2.0 - 256K context - Trained for long-horizon agent work - Includes quantized variantshuggingface.coAgents-A1 - a InternScience CollectionWe’re on a journey to advance and democratize artificial intelligence through open source and open science. 2112
Adina Yakup @adinayakup.bsky.social · 07/07/2026LingBot Vision 👀🤖 A self-supervised vision backbone family for dense spatial perception from Ant Group huggingface.co/collections/... 141
Adina Yakup @adinayakup.bsky.social · 06/07/2026Tencent just released HY3 🔥 huggingface.co/collections/... - 295B / 21B MoE , 256K context - Apache 2.0 - FP8 version included 👀 - Switchable reasoning: no think / low / high - Hallucinates half as often as before - Stable across agent frameworks 0202
Adina Yakup @adinayakup.bsky.social · 01/07/2026BAAI just released the Orca paper 🔥 ( weights coming soon ) huggingface.co/papers/2606.... A Multimodal Latent World Model: it learns the world itself first, and text/images/actions are just different ways to read it out 💡 171
Adina Yakup @adinayakup.bsky.social · 23/06/2026Unlimited-OCR 🔥New OCR from Baidu huggingface.co/baidu/Unlimi... It can parse hundreds of pages in a single pass while maintaining stable speed. The key is R-SWA (Reference Sliding Window Attention), which keeps KV cache constant during decoding. 🏆 93% on OmniDocBench 📈 +6% over DeepSeek-OCR 1917
Adina Yakup @adinayakup.bsky.social · 17/06/2026Really cool to see the GLM 5.2 blog on Hugging Face 🔥 huggingface.co/blog/zai-org...huggingface.coGLM-5.2: Built for Long-Horizon TasksA Blog post by Z.ai on Hugging Face 0262
Adina Yakup @adinayakup.bsky.social · 16/06/2026GLM 5.2 is here 🔥 huggingface.co/collections/... ✨ 753B - 1M context ✨ MIT license ✨ GLM IndexShare: reuses the indexer across layers, 2.9x fewer FLOPs/token at 1M ✨ AIME 2026: 99.2 (beats GPT-5.5, Claude Opus 4.8) ✨ vLLM / SGLang / Transformers supported 0362
Adina Yakup @adinayakup.bsky.social · 12/06/2026MiniMax-M3 just dropped 🔥 huggingface.co/MiniMaxAI/Mi... ✨ 428B / 23B active ✨ 1M context ✨ MiniMax Sparse Attention (MSA) And it’s not just weights! - paper: huggingface.co/papers/2606.... - kernel: huggingface.co/kernels/Mini... - Transformers support Love how this was released❤️ 1365
Adina Yakup @adinayakup.bsky.social · 11/06/2026PP-OCRv6 just released by Baidu huggingface.co/collections/... ✨ tiny 1.5M / small 7.7M / medium 34.5M ✨ 48+ languages ✨ Supports handwritten/printed/industrial/screen and card text ✨ Edge friendly deployment 1375
Adina Yakup @adinayakup.bsky.social · 08/06/2026Macaron-V1-Preview-749B 👀 a Mixture-of-LoRA personal agent model from MindLab ✨ 744B base + 5 specialist LoRAs ✨ Generative UI as a core skill ✨ Personal agent focused ✨ 202K context ✨ MIT license huggingface.co/collections/...huggingface.coMacaron-V1 - a mindlab-research CollectionWe’re on a journey to advance and democratize artificial intelligence through open source and open science. 070
Adina Yakup @adinayakup.bsky.social · 08/06/2026dots.tts 🔊 New TTS from Xiaohongshu (RedNote) huggingface.co/collections/... ✨ 2B - Apache 2.0 ✨ Fully continuous architecture (no codec tokens) ✨ 48kHz synthesis ✨ Zero-shot voice cloning 0250
Adina Yakup @adinayakup.bsky.social · 29/05/2026Step-3.7-Flash 🔥 New VL model from StepFun_ai huggingface.co/collections/... ✨ 198B / 11B active - MoE ✨ 256K context ✨ 3 reasoning level ✨ Up to 400 tokens/sec 🤯huggingface.coStep-3.7-Flash - a stepfun-ai CollectionWe’re on a journey to advance and democratize artificial intelligence through open source and open science. 0162
Adina Yakup @adinayakup.bsky.social · 28/05/2026Qwen just dropped a new Text to Image benchmark + a judge model huggingface.co/collections/... ✨ 56 fine-grained evaluation facets ✨ Measures creativity beyond prompt alignment ✨ Covers storytelling/typography/design & physical logic ✨ Human aligned judge model (ρ = 0.92)huggingface.coQwen-Image-Bench - a Qwen CollectionWe’re on a journey to advance and democratize artificial intelligence through open source and open science. 1100
Adina Yakup @adinayakup.bsky.social · 25/05/2026MiniCPM5-1B is an impressive release in the 1B class! huggingface.co/collections/... ✨ 1B - Apache 2.0 ✨ Hybrid reasoning with Think / No-Think modes ✨ 128K context ✨ Runs on CPU/Apple Silicon/GPU ✨ Strong eval result in the same size class 0150
Adina Yakup @adinayakup.bsky.social · 22/05/2026BitCPM4-CANN 🔥Native 1.58-bit LLM training system on Ascend NPUs huggingface.co/collections/... ✨ 0.5B/1B/3B/8B - Apache 2.0 ✨ 6× less memory at inference ✨ Only 4.5% training throughput overheadhuggingface.coBitCPM4-CANN - a openbmb CollectionFull-pipeline ternary quantized model trained on CANN. 050
Adina Yakup @adinayakup.bsky.social · 22/05/2026LongCat-Video-Avatar 1.5🐱 an audio driven avatar video generation framework from Meituan huggingface.co/meituan-long... ✨ Multi-character + multi-audio support ✨ Drive video from audio alone or audio + image + text ✨ 8-step inference ✨ Whisper-Large powered lip sync ✨ MIT license 2163
Adina Yakup @adinayakup.bsky.social · 21/05/2026Hy-MT2 🔥 New translation model family from Tencent Hunyuan ✨ 1.8B / 7B / 30B-A3B MoE ✨ Supports 33 languages ✨ 1.8B > 440MB with 1.25-bit quantization ✨ Runs on device with faster inference ✨ 1.8B outperforms some commercial APIs 2222
Adina Yakup @adinayakup.bsky.social · 19/05/2026HiDream-O1-Image is getting a lot of attention🔥 A few things that make it different: ✨ Interesting architecture: no VAE, no disjoint encoders, just raw pixels and text in one shared token space ✨ 8B + MIT license ✨ Native 2048×2048 ✨ Built in reasoning agent 1102
Adina Yakup @adinayakup.bsky.social · 19/05/2026ByteDance dropped Lance👀 huggingface.co/bytedance-re... This 3B model can generate images + edit images + generate videos + edit videos, and understand both images/videos. It's trained from scratch on only 128 A100s, and beats several 7B+ models on GenEval and VBench! 0485
Adina Yakup @adinayakup.bsky.social · 19/05/2026✨Big update from Baidu The PaddleOCR now supports Transformers as an inference backend 🔥 Really cool to see it becoming easier to use within the @hf.co ecosystem! huggingface.co/blog/PaddleP...huggingface.coPaddleOCR 3.5: Running OCR and Document Parsing Tasks with a Transformers BackendA Blog post by PaddlePaddle on Hugging Face 020
Adina Yakup @adinayakup.bsky.social · 15/05/2026Intern S2 preview 🔥 A scientific multimodal model from Shanghai AI Lab huggingface.co/internlm/Int...huggingface.cointernlm/Intern-S2-Preview · Hugging FaceWe’re on a journey to advance and democratize artificial intelligence through open source and open science. 151
Adina Yakup @adinayakup.bsky.social · 14/05/2026Ant group just dropped Ring-2.6-1T 🔥 1T reasoning model, built for real world agent workflows. ✨ MIT license ✨ 128K >> 256K context (YaRN) ✨ Async RL + IcePop training architecture ✨ Dual reasoning : "high" for fast agent loops, "xhigh" for deep reasoning = Better cost/performance tradeoff 👀 140
Adina Yakup @adinayakup.bsky.social · 14/05/2026Ant group just dropped Ring-2.6-1T 🔥 1T reasoning model, built for real world agent workflows. ✨ MIT license ✨ 128K >> 256K context (YaRN) ✨ Async RL + IcePop training architecture ✨ Dual reasoning: "high" for fast agent loops, "xhigh" for deep reasoning = Better cost/performance tradeoff 👀huggingface.coinclusionAI/Ring-2.6-1T · Hugging FaceWe’re on a journey to advance and democratize artificial intelligence through open source and open science. 040
Adina Yakup @adinayakup.bsky.social · 14/05/2026Ant group just dropped Ring-2.6-1T 🔥 1T reasoning model, built for real world agent workflows. ✨ MIT license ✨ 128K >> 256K context (YaRN) ✨ Async RL + IcePop training architecture ✨ Dual reasoning: "high" for fast agent loops, "xhigh" for deep reasoning = Better cost/performance tradeoff 000
Adina Yakup @adinayakup.bsky.social · 13/05/2026MemPrivacy ㊙️ a lightweight privacy preserving model for edge cloud AI agents from MemTensor. huggingface.co/collections/... ✨ 1.7B/4B - RL/SFT ✨ High precision privacy extraction ✨ 4 level privacy taxonomy (PL1–PL4) ✨ Semantic preserving typed placeholdershuggingface.coMemPrivacy - a IAAR-Shanghai CollectionMemPrivacy provides specialized privacy extraction models for edge-cloud memory systems in LLM-powered agents. 021
Adina Yakup @adinayakup.bsky.social · 11/05/2026MiniCPM V4.6 🔥 a 1B MLLM that actually runs on your phone, just released by OpenBMB huggingface.co/openbmb/Mini... ✨ 1B - Apache2.0 ✨ Runs on iOS, Android, HarmonyOS ✨ ~1.5× faster throughput than Qwen3.5 0.8B ✨ Mixed 4x/16x visual token compression 0101
Adina Yakup @adinayakup.bsky.social · 11/05/2026Qwen released WebWorld 🌍 an open world model series for web agents ✨ 8B/14B/32B+Dataset ✨Apache2.0 ✨+9.9% MiniWob++, +10.9% WebArena ✨ Matches Claude Opus 4.1 & Gemini 3 Pro on factuality,beats GPT-5 as world model ✨Unified action space, 30+ step simulation, 5 state formats 140
Adina Yakup @adinayakup.bsky.social · 29/04/2026Ling-2.6-1T just dropped by Antgroup , one day after Ling 2.6 Flash. Both optimized for the same goal: usable intelligence at the lowest token cost 💰 ✨ MLA + Linear Attention hybrid for long context efficiency ✨ Answers without verbose CoT, lower token cost ✨ 1T / 63B active ✨ MIT license 1111
Adina Yakup @adinayakup.bsky.social · 28/04/2026SenseTime is back to open source with SenseNova U1 🔥 A unified multimodal model in collab with MMLab NTU huggingface.co/collections/... ✨ no visual encoder & VAE ✨ NEO Unify end to end ✨ End-to-end pixel–word modeling 170
Adina Yakup @adinayakup.bsky.social · 27/04/2026MiMo-V2.5 🔥 a native omni modal MoE built for agents released by XiaomiMiMo huggingface.co/collections/... ✨Base (310B total / 15B active) & Pro (1T total / 42B active) ✨MIT license ✨Text + image + video + audio ALL IN ONE ✨Up to 1M context ✨Agentic RL + MOPD post-traininghuggingface.coMiMo-V2.5 - a XiaomiMiMo CollectionWe’re on a journey to advance and democratize artificial intelligence through open source and open science. 0120
Adina Yakup @adinayakup.bsky.social · 24/04/2026DeepSeek V4 🐳🔥 makes 1M token context & SOTA Agentic/coding more accessible!! huggingface.co/collections/... ✨Pro (1.6T/49B active) & Flash (284B/13B active) ✨1M token context ✨LiveCodeBench 93.5 🤯 beats Gemini 3.1 Pro & Opus 4.6 ✨MIT licensed 0252
Adina Yakup @adinayakup.bsky.social · 23/04/2026LLaDA2.0-Uni from Antgroup is an interesting direction 👀 huggingface.co/inclusionAI/... ✨ 16B MoE/ 1B active per token ✨ Unifies generation + understanding + reasoning in a single diffusion LLM ✨ MoE backbone + diffusion decoder ✨ Apache 2.0 151
Adina Yakup @adinayakup.bsky.social · 23/04/2026MiMo-V2.5-ASR 🔥 new ASR model from Xiaomi Model: huggingface.co/XiaomiMiMo/M... Demo: huggingface.co/spaces/Xiaom... ✨ MIT license ✨ Support Mandarin, English & Chinese dialects ✨ Code-switching (no language tags needed) ✨ Multi-speaker / meetings ✨ Noisy & far-field audio 000
Adina Yakup @adinayakup.bsky.social · 23/04/2026Tencent just released HY3 preview on @hf.co First drop from TencentHunyuan rebuilt infra👀 Model: huggingface.co/collections/... Demo: huggingface.co/spaces/tence... ✨ 295B MoE /21B active ✨ 256k context ✨ Hybrid fast/slow thinking ✨ Solid BrowseComp & WideSearch (search agents) 082
Adina Yakup @adinayakup.bsky.social · 22/04/2026Qwen3.6-27B just dropped on @huggingface 🔥 huggingface.co/Qwen/Qwen3.6... huggingface.co/Qwen/Qwen3.6... ✨ 27B dense model ✨ 262K native context, scalable to 1M+ tokens ✨ Vision + language in one model ✨ 77.2 on SWE-bench Verified 05310
Adina Yakup @adinayakup.bsky.social · 20/04/2026Kimi 2.6 is now available on @hf.co 🔥🎉 huggingface.co/moonshotai/K... ✨ 1T MoE / 32B active / 256K context ✨ Agent Swarm: 300 sub-agents × 4,000 steps ✨ Modified MIT 2316
Adina Yakup @adinayakup.bsky.social · 20/04/2026MOSS-VL 🔥 Vision model from Open MOSS Model: huggingface.co/collections/... Demo: huggingface.co/spaces/OpenM... ✨ 11B - Apache 2.0 ✨ Cross-attention + XRoPE (3D: time, height, width) ✨ Beats Qwen3-VL-8B by 8.3 pts on VSI-benchhuggingface.coMOSS-VL - a OpenMOSS-Team CollectionWe’re on a journey to advance and democratize artificial intelligence through open source and open science. 031
Adina Yakup @adinayakup.bsky.social · 14/04/2026Baidu just released ERNIE-Image on Hugging Face🔥 Model: huggingface.co/collections/... Demo: huggingface.co/spaces/baidu... ✨ 8B DiT - Image/Image Turbo ✨ Apache2.0 ✨ Strong text rendering for posters & UI-style images ✨ Structured outputs (comics, multi-panel scenes)huggingface.coERNIE-Image - a baidu CollectionThe serieas of image generation models, including text2img、img2img. 010
Adina Yakup @adinayakup.bsky.social · 31/03/2026A new large-scale RGB-D dataset from Ant Group: LingBot-Depth 🤖 huggingface.co/datasets/rob... ✨3M+ samples / 2.7TB ✨Real-world + simulation + VLA robotics data ✨Raw sensor depth + ground truthhuggingface.corobbyant/mdm_depth · Datasets at Hugging FaceWe’re on a journey to advance and democratize artificial intelligence through open source and open science. 010
Adina Yakup @adinayakup.bsky.social · 30/03/2026LongCat-AudioDiT 🔊New TTS from Meituan LongCat team huggingface.co/meituan-long... ✨ 1B & 3.5B - MIT license ✨ Diffusion + non-AR generation ✨ Operates directly in waveform latent space ✨ Simpler pipeline (no mel-spectrograms)huggingface.comeituan-longcat/LongCat-AudioDiT-1B · Hugging FaceWe’re on a journey to advance and democratize artificial intelligence through open source and open science. 021
Adina Yakup @adinayakup.bsky.social · 27/03/2026Matrix-Game 3.0🔥real-time interactive world models from Skywork huggingface.co/Skywork/Matr... ✨ MIT license ✨ 720p @ 40FPS with a 5B model ✨ Minute-long memory consistency ✨ Unreal + AAA + real-world data ✨ Scales up to 28B MoE 2132
Adina Yakup @adinayakup.bsky.social · 27/03/2026daVinci-LLM 🔥 The SII-GAIR team just shared the full training pipeline on @hf.co huggingface.co/SII-GAIR-NLP... ✨ 3B, competitive with 7B models ✨ 8T-token transparent training ✨ 200+ ablation studies ✨ Data Darwinism (L0–L9) framework 182