Sign in

hochreitersepp.bsky.social

@hochreitersepp.bsky.social
419 followers 178 following 50 posts
PostsRepliesMedia
hochreitersepp.bsky.social @hochreitersepp.bsky.social · 02/08/2026
xLSTM for three-dimensional bin packing problem: arxiv.org/abs/2607.28257 xLSTM mixes item and candidate tokens with a recurrent action mixer (LRAM). xLSTM based OPAL has mean space utilization of 0.49 (15.1% improvement), while being fast. Promising for high-fidelity robotics.
000
hochreitersepp.bsky.social @hochreitersepp.bsky.social · 24/07/2026
xLSTM for Brain MRI Analysis: arxiv.org/abs/2607.17782 Three-dimensional Bi-Directional xLSTM-UNet achieved second place overall and first in the meningioma segmentation task on FOMO 2025 leaderboard. xLSTM models long-range dependencies throughout the volumetric brain MRI.
020
hochreitersepp.bsky.social @hochreitersepp.bsky.social · 02/07/2026
1/5 🚀 Introducing TiRex-2 — our next-generation time series foundation model. Time series forecasting in the real world is streaming: • new observations continuously arrive, • variables interact, • some covariates are known into the future, • models must update predictions efficiently.
113
hochreitersepp.bsky.social @hochreitersepp.bsky.social · 11/06/2026
Comparison of sub-quadratic architectures xLSTM, Mamba-2, and Gated DeltaNet: arxiv.org/abs/2606.12364 Comparison of xLSTM, Mamba-2, and Gated DeltaNet on code pre-training, distillation, and time-series. xLSTM outperforms the others due to its gating scheme and state tracking.
001
Reposted by @hochreitersepp.bsky.social
Lukas Aichberger @aichberger.bsky.social · 29/05/2026
5/ Takeaway LLMs do not always need to externalize their thoughts. They can learn to reason in working memory instead, decoupling intermediate computation from autoregressive generation 💡 Full paper: arxiv.org/abs/2605.30343 Huge thanks to @hochreitersepp.bsky.social for the guidance!
arxiv.org
Unlocking the Working Memory of Large Language Models for Latent Reasoning
To improve the reasoning capabilities of large language models, test-time compute is typically scaled by generating intermediate tokens before the final answer. However, this couples reasoning to auto...
001
hochreitersepp.bsky.social @hochreitersepp.bsky.social · 21/04/2026
RNNs like xLSTM with vertically chunked inference strategy for efficient memory usage: arxiv.org/abs/2604.18199 Chunking enables a linear-time and constant-memory like for TFLA for xLSTM arxiv.org/abs/2503.14376 Chunking blocks via recurrent updates speeds up computation considerably.
arxiv.org
Linear-Time and Constant-Memory Text Embeddings Based on Recurrent Language Models
Transformer-based embedding models suffer from quadratic computational and linear memory complexity, limiting their utility for long sequences. We propose recurrent architectures as an efficient alter...
020
hochreitersepp.bsky.social @hochreitersepp.bsky.social · 31/03/2026
A high-performance trading architecture built around xLSTM: allenarch.dev/blog/combini... The model is used as a market-state encoder, while reinforcement learning handles the trading decisions. This performance of xLSTM should be confirmed by further investigations.
000
hochreitersepp.bsky.social @hochreitersepp.bsky.social · 22/03/2026
xLSTM is more expressive than Transformer, Mamba: arxiv.org/abs/2603.03612 *nonlinear RNNs: sLSTM, LSTM *DLPR linear RNNs: mLSTM, RWKV-7, DeltaNet *Non PNC1-complete: Mamba, Transformer “fundamental expressivity gaps between linear and nonlinear RNNs.” World models require nonlinear RNNs.
020
hochreitersepp.bsky.social @hochreitersepp.bsky.social · 17/03/2026
xLSTM Distillation: arxiv.org/abs/2603.15590 Near-lossless distillation of quadratic Transformer LLMs into linear-time xLSTM architectures enables cost- and energy-efficient alternatives without sacrificing performance. Efficient xLSTM variants of instruction-tuned Llama, Qwen, and Olmo models.
057
Reposted by @hochreitersepp.bsky.social
Günter Klambauer @gklambauer.bsky.social · 03/03/2026
Symbol-equivariant Recurrent Reasoning Models (SE-RRM) SE-RRM advances HRM and TRM -- guaranteed identical solutions for problems with permuted colors (ARC AGI) or digits (Sudoku). Coolest part: extrapolation to larger problem sizes!!! P: arxiv.org/abs/2603.02193 C: github.com/ml-jku/SE-RRM
051
hochreitersepp.bsky.social @hochreitersepp.bsky.social · 03/03/2026
xLSTM for Financial Time Series: arxiv.org/abs/2603.01820 "VLSTM achieved the highest overall Sharpe ratio" "VxLSTM and LPatchTST exhibited superior downside-adjusted characteristics" “xLSTM achieves the highest portfolio-level cost buffer" xLSTM excels in financial time series as TiRex does.
020
hochreitersepp.bsky.social @hochreitersepp.bsky.social · 15/01/2026
Drug design is significantly accelerated by ConGLUDe. Protein–small-molecule interactions can be screened orders of magnitude faster using the ConGLUDe approach.
021
hochreitersepp.bsky.social @hochreitersepp.bsky.social · 29/12/2025
xLSTM for Lensed Gravitational Waves: arxiv.org/abs/2512.21370 sLSTM models fine-grained temporal structures, while mLSTM finds large-scale global patterns. xLSTM achieves AUC beyond 0.99, a TPR above 98% a FPR below 1% and is robust against noise, lens type, lens mass. Cool xLSTM application.
031
hochreitersepp.bsky.social @hochreitersepp.bsky.social · 11/12/2025
xLSTM for Real-Time DNS Tunnel Detection: arxiv.org/abs/2512.09565 DNS-HyXNet = xLSTM for DNS tunnels. DNS-HyXNet has 99.99% accuracy, with F1-scores exceeding 99.96%, and per-sample detection latency of just 0.041 ms, confirming its scalability and real-time readiness. wow!
021
hochreitersepp.bsky.social @hochreitersepp.bsky.social · 20/11/2025
xLSTM for PINNs that learn PDEs: arxiv.org/abs/2511.12512 “Across four PDEs under matched size and budget, xLSTM-PINN consistently reduces MSE, RMSE, MAE, and MaxAE with markedly narrower error bands.” “cleaner boundary transitions with attenuated high-frequency ripples”
032
Reposted by @hochreitersepp.bsky.social
Günter Klambauer @gklambauer.bsky.social · 19/11/2025
Measuring AI Progress in Drug Discovery - A NEW LEADERBOARD IN TOWN 2015-2025: turns out that there's hardly any improvement. AI bubble? GPT is at 70% for this task, whereas the best methods get close to 85%. Leaderboard: huggingface.co/spaces/ml-jk... P: arxiv.org/abs/2511.14744
3127
hochreitersepp.bsky.social @hochreitersepp.bsky.social · 04/11/2025
xLSTM for Vehicle Trajectory Prediction: arxiv.org/abs/2511.00266 X-TRACK based on xLSTM achieves SOTA. “Compared to state-of-the-art baselines, X-TRACK achieves performance improvement by 79% at the 1-second prediction and 20% at the 5-second prediction in the case of highD” Again xLSTM excels.
031
hochreitersepp.bsky.social @hochreitersepp.bsky.social · 26/10/2025
xLSTM for robotic manipulation systems via diffusion-based imitation learning: arxiv.org/abs/2510.20406 PMP leverages xLSTM to denoise actions for robotics. “PMP not only achieves state-of-the-art performance but also offers significantly faster training and inference.” xLSTM excels in robotics.
020
hochreitersepp.bsky.social @hochreitersepp.bsky.social · 22/10/2025
Tenure Track in quantum informatics! Super cool position. Super cool team. World-class research. Scientifically outstanding work.
042
hochreitersepp.bsky.social @hochreitersepp.bsky.social · 21/10/2025
xLSTM for Toxic Comment Classification: arxiv.org/abs/2510.17018 “On the Jigsaw Toxic Comment benchmark, xLSTM attains 96.0% accuracy and 0.88 macro-F1, outperforming BERT by 33% on threat and 28% on identity_hate categories, with 15× fewer parameters and <50 ms inference latency.” xLSTM is fast!
000
hochreitersepp.bsky.social @hochreitersepp.bsky.social · 10/10/2025
gLSTM extends xLSTM to a graph neural network architecture: arxiv.org/abs/2510.08450 "gLSTM mitigates sensitivity over-squashing and capacity over-squashing." "gLSTM achieves comfortably state of the art results on the Diameter and Eccentricity Graph Property Prediction tasks"
020
hochreitersepp.bsky.social @hochreitersepp.bsky.social · 10/10/2025
xLSTM for Intrusion Detection: arxiv.org/abs/2510.08333 "The xLSTM-based IDS achieves an F1-score of 98.9%, surpassing the transformer-based model at 94.3%." xLSTM is faster than transformer when using fast kernels as provided in github.com/nx-ai/mlstm_... and github.com/NX-AI/flashrnn
010
hochreitersepp.bsky.social @hochreitersepp.bsky.social · 30/09/2025
xLSTM for long-term context using short sliding windows: arxiv.org/abs/2509.24552 "SWAX, a hybrid consisting of sliding-window attention and xLSTM." "SWAX trained with stochastic window sizes significantly outperforms regular window attention both on short and long-context problems."
000
hochreitersepp.bsky.social @hochreitersepp.bsky.social · 16/09/2025
xLSTM shines as an Electrocardiogram (ECG) foundation model: arxiv.org/abs/2509.10151 "xECG achieves superior performance over earlier approaches, defining a new baseline for future ECG foundation models." xLSTM is perfectly suited for time series prediction as shown by TiRex.
130
hochreitersepp.bsky.social @hochreitersepp.bsky.social · 03/09/2025
xLSTM excels in time series forecasting: arxiv.org/abs/2509.01187 . Introduces "stochastic xLSTM" (StoxLSTM). "StoxLSTM consistently outperforms state-of-the-art baselines with better robustness and stronger generalization ability." We know that xLSTM is king at time series from our TiRex.
020
hochreitersepp.bsky.social @hochreitersepp.bsky.social · 29/07/2025
xLSTM for Cellular Traffic Forecasting: arxiv.org/abs/2507.19513 "Empirical results showed a 23% MAE reduction over the original STN and a 30% improvement on unseen data, highlighting strong generalization." xLSTM shines again in time series forecasting.
020
hochreitersepp.bsky.social @hochreitersepp.bsky.social · 08/07/2025
xLSTM for Monaural Speech Enhancement: arxiv.org/abs/2507.04368 xLSTM has superior performance vs. Mamba and Transformers but is slower than Mamba. New Triton kernels: xLSTM is faster than MAMBA at training and inference: arxiv.org/abs/2503.13427 and arxiv.org/abs/2503.14376
010
hochreitersepp.bsky.social @hochreitersepp.bsky.social · 05/07/2025
xLSTM for Aspect-based Sentiment Analysis: arxiv.org/abs/2507.01213 Another success story of xLSTM. MEGA: xLSTM with Multihead Exponential Gated Fusion. Experiments on 3 benchmarks show that MEGA outperforms state-of-the-art baselines with superior accuracy and efficiency”
000
hochreitersepp.bsky.social @hochreitersepp.bsky.social · 01/07/2025
xLSTM for multivariate time series anomaly detection: arxiv.org/abs/2506.22837 “In our results, xLSTM showcases state-of-the-art accuracy, outperforming 23 popular anomaly detection baselines.” Again, xLSTM excels in time series analysis.
042
hochreitersepp.bsky.social @hochreitersepp.bsky.social · 12/06/2025
xLSTM for Human Action Segmentation: arxiv.org/abs/2506.09650 "HopaDIFF, leveraging a novel cross-input gate attentional xLSTM to enhance holistic-partial long-range reasoning" "HopaDIFF achieves state-of-theart results on RHAS133 in diverse evaluation settings."
030
hochreitersepp.bsky.social @hochreitersepp.bsky.social · 04/06/2025
Mein Buch “Was kann Künstliche Intelligenz?“ ist erschienen. Eine leicht zugängliche Einführung in das Thema Künstliche Intelligenz. LeserInnen – auch ohne technischen Hintergrund – wird erklärt, was KI eigentlich ist, welche Potenziale sie birgt und welche Auswirkungen sie hat.
033
hochreitersepp.bsky.social @hochreitersepp.bsky.social · 04/06/2025
We are soooo proud. Our European-developed TiRex is leading the field—significantly ahead of U.S. competitors like Amazon, Datadog, Salesforce, and Google, as well as Chinese models from companies such as Alibaba.
054
hochreitersepp.bsky.social @hochreitersepp.bsky.social · 02/06/2025
Attention!! Our TiRex time series model, built on xLSTM, is topping all major international leaderboards. A European-developed model is leading the field—significantly ahead of U.S. competitors like Amazon, Datadog, Salesforce, and Google, as well as Chinese models from Alibaba.
073
hochreitersepp.bsky.social @hochreitersepp.bsky.social · 02/06/2025
TiRex 🦖 time series xLSTM model ranked #1 on all leaderboards. ➡️ Outperforms models by Amazon, Google, Datadog, Salesforce, Alibaba ➡️ industrial applications ➡️ limited data ➡️ embedded AI and edge devices ➡️ Europe is leading Code: lnkd.in/eHXb-XwZ Paper: lnkd.in/e8e7xnri shorturl.at/jcQeq
linkedin.com
Introducing TiRex - xLSTM based time series model | NXAI
TiRex model at the top 🦖 We are proud of TiRex - our first time series model based on #xLSTM technology. Key take aways: 🥇 Ranked #1 on official international leaderboards ➡️ Outperforms models ...
055
Reposted by @hochreitersepp.bsky.social
Günter Klambauer @gklambauer.bsky.social · 30/05/2025
Recommended read for the weekend: Sepp Hochreiter's book on AI! Lots of fun anecdotes and easily accessible basics on AI! www.beneventopublishing.com/ecowing/prod...
163
hochreitersepp.bsky.social @hochreitersepp.bsky.social · 26/05/2025
xLSTM for the classification of assembly tasks: arxiv.org/abs/2505.18012 "xLSTM model demonstrated better generalization capabilities to new operators. The results clearly show that for this type of classification, the xLSTM model offers a slight edge over Transformers."
031
Reposted by @hochreitersepp.bsky.social
fses91.bsky.social @fses91.bsky.social · 22/05/2025
Happy to introduce 🔥LaM-SLidE🔥! We show how trajectories of spatial dynamical systems can be modeled in latent space by --> leveraging IDENTIFIERS. 📚Paper: arxiv.org/abs/2502.12128 💻Code: github.com/ml-jku/LaM-S... 📝Blog: ml-jku.github.io/LaM-SLidE/ 1/n
178
Reposted by @hochreitersepp.bsky.social
Sebastian Sanokowski @sanokows.bsky.social · 24/04/2025
1/11 Excited to present our latest work "Scalable Discrete Diffusion Samplers: Combinatorial Optimization and Statistical Physics" at #ICLR2025 on Fri 25 Apr at 10 am! #CombinatorialOptimization #StatisticalPhysics #DiffusionModels
1157
hochreitersepp.bsky.social @hochreitersepp.bsky.social · 24/04/2025
xLSTM for Multi-label ECG Classification: arxiv.org/abs/2504.16101 "This approach significantly improves ECG classification accuracy, thereby advancing clinical diagnostics and patient care." Cool.
020
hochreitersepp.bsky.social @hochreitersepp.bsky.social · 10/04/2025
Precall: 4 Tenure-Tracks in AI for females. Check it out.
044
hochreitersepp.bsky.social @hochreitersepp.bsky.social · 19/03/2025
Huge achievement: arxiv.org/abs/2503.14376 xLSTM kernels are now the fastest kernels both for training and inference. Faster than Flashattention or Mamba kernels. High arithmetic intensity. Optimized use of GPUs that set a new state of the art. Congratulations to the team. It is highly impressive.
arxiv.org
083
Reposted by @hochreitersepp.bsky.social
Jonathan Tirone @virtualnomad.bsky.social · 16/03/2025
Visited with legendary #AI pioneer @hochreitersepp.bsky.social at #JKULinz. We discussed the efficiency of #xLSTM against #GPT and why AI's future may be located on the edge of networks, rather than in centralised data centres via @bloomberg.com www.bloomberg.com/news/article...
bloomberg.com
AI Pioneer Wants Europe to Forge Its Own Nimbler Way Forward
One belief underlying the power-hungry approach to machine learning advanced by OpenAI and Mistral AI is that an artificial intelligence model must review its entire dataset before spitting out new in...
053
hochreitersepp.bsky.social @hochreitersepp.bsky.social · 14/03/2025
xLSTM for Automated Stock Trading: arxiv.org/abs/2503.09655 xLSTM outperforms LSTM. "These findings mark the potential of xLSTM for enhancing DRL-based stock trading systems." This reinforcement learning approach uses xLSTM in both actor and critic components, which increases the performance.
arxiv.org
A Deep Reinforcement Learning Approach to Automated Stock Trading, using xLSTM Networks
Traditional Long Short-Term Memory (LSTM) networks are effective for handling sequential data but have limitations such as gradient vanishing and difficulty in capturing long-term dependencies, which ...
093
hochreitersepp.bsky.social @hochreitersepp.bsky.social · 14/03/2025
Join Our Research Team in Linz! We are looking for 5 PostDocs and 10 PhDs in Machine Learning working on xLSTM, NLP, robustness, density ratio. Deadline: 04/20/25. More details: www.jku.at/en/lit-artif... #MachineLearning #DeepLearning #ResearchOpportunities #PhDPositions
jku.at
Deep Learning
0154
hochreitersepp.bsky.social @hochreitersepp.bsky.social · 19/02/2025
Exploration imitation learning architectures: Transformer, Mamba, xLSTM: arxiv.org/abs/2502.12330 *LIBERO: “xLSTM shows great potential” *RoboCasa: “xLSTM models, we achieved success rate of 53.6%, compared to 40.0% of BC-Transformer” *Point Clouds: “xLSTM model achieves a 60.9% success rate”
arxiv.org
X-IL: Exploring the Design Space of Imitation Learning Policies
Designing modern imitation learning (IL) policies requires making numerous decisions, including the selection of feature encoding, architecture, policy representation, and more. As the field rapidly a...
063
hochreitersepp.bsky.social @hochreitersepp.bsky.social · 17/02/2025
xLSTM for time series with Granger causality: arxiv.org/abs/2502.09981 xLSTM again shows superb performance at time series analysis. "Our experimental evaluations on three datasets demonstrate the overall efficacy of our proposed GC-xLSTM model."
arxiv.org
Exploring Neural Granger Causality with xLSTMs: Unveiling Temporal Dependencies in Complex Data
Causality in time series can be difficult to determine, especially in the presence of non-linear dependencies. The concept of Granger causality helps analyze potential relationships between variables,...
0153
hochreitersepp.bsky.social @hochreitersepp.bsky.social · 04/02/2025
xLSTM shines at tumor segmentation: arxiv.org/abs/2502.00314 “evaluated state-of-the-art segmentation methods, including U-Net and its enhanced variants with Transformers and Mamba. Our proposed ViLU-Net [vision xLSTM-Net] model achieved superior performance with reduced complexity.” Cool.
arxiv.org
A Study on the Performance of U-Net Modifications in Retroperitoneal Tumor Segmentation
The retroperitoneum hosts a variety of tumors, including rare benign and malignant types, which pose diagnostic and treatment challenges due to their infrequency and proximity to vital structures. Est...
031
hochreitersepp.bsky.social @hochreitersepp.bsky.social · 31/01/2025
Don't miss this workshop. Super interesting .
021
hochreitersepp.bsky.social @hochreitersepp.bsky.social · 31/01/2025
xLSTM for molecular property prediction: arxiv.org/abs/2501.18439 "AUROC improvement of 3.18% for classification tasks and an RMSE reduction of 3.83% across regression datasets compared to the baseline methods." Again xLSTM excels in a life science. Compare Bio-xLSTM: arxiv.org/abs/2411.04165.
lnkd.in
LinkedIn
This link will take you to a page that’s not on LinkedIn
040
hochreitersepp.bsky.social @hochreitersepp.bsky.social · 27/01/2025
xLSTM for knowledge tracing: arxiv.org/abs/2501.14256 xLSTM excels at another task. “DKT2 [the xLSTM] consistently outperforms 17 baseline models in various prediction tasks” Baseline models include Transformers, Mamba and graph-based methods. DKT2 exploits exponential gating and matrix memory.
arxiv.org
Revisiting Applicable and Comprehensive Knowledge Tracing in Large-Scale Data
Knowledge Tracing (KT) is a fundamental component of Intelligent Tutoring Systems (ITS), enabling the modeling of students' knowledge states to predict future performance. The introduction of Deep Kno...
011