Sign in

Underfox

@underfox3.bsky.social
814 followers 18 following 2.3K posts

Physicist, Telecom Engineering lover, HPC Enthusiast. Prog Rock/Metal fan. --- Independent tech analyst focused on semiconductors, patent analysis and emerging technologies.

PostsRepliesMedia
Underfox @underfox3.bsky.social · 03/09/2026
In this paper is presented CHIPSMORE, a multi-mode and multi-request CIM accelerator for LLM inference across diverse operating scenarios, including base-mode execution, LoRA adaptation, varying context lengths, and concurrent multi-request serving. arxiv.org/pdf/2608.30509
120
Underfox @underfox3.bsky.social · 03/09/2026
In this paper is introduced the first auto-tuning framework for Julia GPU kernels as an accessible Julia package, enabling systematic auto-tuning of hardware-agnostic kernels to target a wide variety of platforms, such as NVIDIA, AMD, Intel, and Apple. arxiv.org/pdf/2608.21227
120
Underfox @underfox3.bsky.social · 03/09/2026
In this paper is proposed NeuroPrefetcher, a storage-backed sparse inference runtime built around predictive delta prefetching for the regime where an LLM exceeds resident memory throughout execution. arxiv.org/pdf/2608.22643
110
Underfox @underfox3.bsky.social · 02/09/2026
Researchers have presented the first experimental ML-enabled Process-Technology Co-Optimization (PTCO) framework for optimizing 2D TMD FET fabrication directly from statistically meaningful experimental data rather than pure simulation data. arxiv.org/pdf/2609.00722
120
Underfox @underfox3.bsky.social · 02/09/2026
In this paper, AMD researchers have presented the first micro-architecture based performance characterization of SPEC CPU 2026 on AMD EPYC "Zen 5" processor. arxiv.org/pdf/2609.01527
172
Underfox @underfox3.bsky.social · 02/09/2026
In this paper is proposed GapFill, a specialized tool designed to assist professional anime colorists in addressing small, unpainted areas, integrating seamlessly into existing professional workflows. arxiv.org/pdf/2609.00800
110
Underfox @underfox3.bsky.social · 02/09/2026
[Basic Concepts] In this paper is presented a comprehensive and critical overview of on-chip photonic neural networks, spanning fundamental device technologies, system architectures, and emerging applications. advanced.onlinelibrary.wiley.com/doi/epdf/10....
020
Underfox @underfox3.bsky.social · 01/09/2026
In this paper is presented a benchmark of the latest RVV 1.0-capable hardware (SiFive X280, SpacemiT X60 (K1) and X100/A100 (K3), and T-Head C920v2 (Sophon SG2044)), using standard HPC benchmarks and synthetic workloads, comparing them to NVIDIA Grace. #HPC arxiv.org/pdf/2608.28097
270
Underfox @underfox3.bsky.social · 29/08/2026
In this paper, researchers have analyzed the current status of using do concurrent for GPU-accelerated Fortran applications across three major GPU vendors (NVIDIA, AMD, and Intel). arxiv.org/pdf/2608.20586
131
Underfox @underfox3.bsky.social · 29/08/2026
In this paper is presented HSMA-TRSM, a hierarchical shared memory-aware optimization framework for left-side lower-triangular triangular solve with multiple right-hand sides on NVIDIA A100, NVIDIA H800, and Hygon DCU Z100 accelerators. #HPC arxiv.org/pdf/2608.25469
110
Underfox @underfox3.bsky.social · 27/08/2026
If I have to choose between migrating to Windows 11 or buying a GPU from another company, I will certainly choose an Intel GPU. videocardz.com/newz/nvidia-...
videocardz.com
NVIDIA ends Windows 10 Game Ready support in October - VideoCardz.com
GeForce Windows 10 support moves to quarterly security updates after October. NVIDIA plans to provide critical fixes through October 2029.
031
Underfox @underfox3.bsky.social · 21/08/2026
In this paper is presented FIBER, a GPU architecture that decouples register ownership from parallel execution instances, comprising ISA extensions, lightweight microarchitectural enhancements, and a CUDA-compatible programming model. arxiv.org/pdf/2608.19628
110
Underfox @underfox3.bsky.social · 21/08/2026
In this paper is introduced a bilingual DSL for modeling compute kernels within a generic solver for hyperbolic PDEs, enabling programmers to express their numerics in an abstract manner in Python, with the physical PDE terms still in C, C++, or SymPy. arxiv.org/pdf/2608.19273
110
Underfox @underfox3.bsky.social · 21/08/2026
Ampere researchers presented the industrial-scale performance verification methodology applied across four generations of the AmpereOne custom CPU core, centered on the cycle-accurate correlation of the RTL design against a trace-driven performance model. arxiv.org/pdf/2608.19300
120
Underfox @underfox3.bsky.social · 20/08/2026
For the first time, researchers have demonstrated non-volatile, deterministic, and correlated electrical switching of the second- and third-order nonlinear anomalous Hall effects in few-layer WTe2 through ferroelectric polarization switching. arxiv.org/pdf/2608.18467
110
Underfox @underfox3.bsky.social · 20/08/2026
In this paper is introduced FlashAttention-V, a blocked FlashAttention algorithm for scalable vector architectures that exploits parallelism across attention heads to utilize longer vector lengths and maximize register utilization and reuse. arxiv.org/pdf/2608.18656
110
Underfox @underfox3.bsky.social · 20/08/2026
[Reading recommendation] In this paper is presented the technical proposal for the Atom Interferometer CERN Experiment (AICE), a O(100) m vertical atom interferometer to be installed against the wall of the PX46 access shaft to the LHC. arxiv.org/pdf/2608.18743
010
Underfox @underfox3.bsky.social · 19/08/2026
Amazon researchers developed a barrier-free allocation algorithm that replaces the barriers with precise, per-dependency wait conditions, operating on structured control-flow graphs with arbitrarily nested, dynamically bounded loops and conditionals. arxiv.org/pdf/2608.13757
110
Underfox @underfox3.bsky.social · 19/08/2026
In this paper is presented ReXpert, a ReRAM near-memory architecture for the MoE FFN pool that combines bounded core-local multicast pooling, routing-aware placement, and load-aware fetch, and provisions each communication level from the induced traffic. arxiv.org/pdf/2608.13962
110
Underfox @underfox3.bsky.social · 19/08/2026
In this paper, researchers have proposed DASH, a GPU memory architecture that connects HBM and HBF as main memory tiers for LLM inference, elevating HBF from a passive capacity extension to a main memory tier. arxiv.org/pdf/2608.14333
110
Underfox @underfox3.bsky.social · 18/08/2026
In this paper is presented SpSYRK and CommSpSYRK, the first distributed-memory algorithms for sparse SYRK that exploit the symmetry of the output matrix. #HPC arxiv.org/pdf/2608.09713
120
Underfox @underfox3.bsky.social · 17/08/2026
In this paper, Nvidia researchers have introduced RGBX-Next, a unified framework for generative forward and inverse rendering, capable of estimating G-buffers from images, videos, and streams, and rendering realistic RGB outputs from G-buffer inputs. arxiv.org/pdf/2608.13929
110
Underfox @underfox3.bsky.social · 17/08/2026
In this paper, researchers have proposed PFM (PIM-as-Flexible-Memory), a unified and flexible memory management solution for NPU-PIM systems that addresses the significant data-sharing and dynamic demands of LLM inference. arxiv.org/pdf/2608.06989
120
Underfox @underfox3.bsky.social · 17/08/2026
In this paper, researchers have proposed SLAC, the first fine-grained Prime+Probe CPU-to-GPU cache side-channel attack on Apple M-series SoCs targeting sensitive GPU workloads. arxiv.org/pdf/2608.09075
120
Underfox @underfox3.bsky.social · 16/08/2026
In this paper, researchers have demonstrated the first vertical β-Ga2O3 device architecture without the use of planarization etch back processes or mid-gap acceptor regions. arxiv.org/pdf/2608.12797
120
Underfox @underfox3.bsky.social · 16/08/2026
In this paper, researchers have demonstrated the first integration of squeezed-light generation and balanced homodyne detection on a heterogeneously integrated photonic chip. arxiv.org/pdf/2608.13218
120
Underfox @underfox3.bsky.social · 14/08/2026
For the first time, researchers have reported the growth of 4-inch wafer-scale magneto-optical Ce:YIG films on silicon substrates using confocal magnetron sputtering. arxiv.org/pdf/2608.09003
110
Underfox @underfox3.bsky.social · 13/08/2026
For the first time, researchers have developed a flexible, domain-engineerable polymer ferroelectric materials via a polar mesogenic approach. arxiv.org/pdf/2608.07942
110
Underfox @underfox3.bsky.social · 13/08/2026
In this paper, researchers have proposed NITRO, a high-performance NAND flash-based in-storage computing (ISC) architecture with enhanced activation buffering in DRAM to reduce read/write latency when using a slow flash memory array. arxiv.org/pdf/2608.11920
120
Underfox @underfox3.bsky.social · 13/08/2026
In this paper is presented the first demonstration of a multidimensional silicon photonic engine that achieves a communication capacity exceeding one terabit per second per wavelength. arxiv.org/pdf/2608.11639
110
Underfox @underfox3.bsky.social · 13/08/2026
For the first time, researchers have demonstrated the realization of symmetric discrete-variable quantum telecloning using a scalable silicon photonic platform. arxiv.org/pdf/2608.11718
120
Underfox @underfox3.bsky.social · 13/08/2026
In this paper, researchers have conducted a controlled evaluation of commercial AI detectors on published English abstracts, showing that they fail to detect humanized synthetic text. arxiv.org/pdf/2608.11256
110
Underfox @underfox3.bsky.social · 13/08/2026
In this paper is presented Load Hijack, a checkpoint-encoded scheduling attack that modifies only the router parameters of a clean MoE checkpoint and makes a private trigger concentrate assignments on one expert parallelism (EP) rank. arxiv.org/pdf/2608.10614
110
Underfox @underfox3.bsky.social · 12/08/2026
In this paper is presented FaCTz, the first GPU-based error-bounded lossy compressor that guarantees critical-point preservation, providing a block-wise mode optimized for throughput and a speculative per-point mode optimized for compression ratio. #HPC arxiv.org/pdf/2608.10586
120
Underfox @underfox3.bsky.social · 12/08/2026
In this paper is proposed AdaptCore, an adaptive framework for universally high-performance MatMul on Ascend NPUs that systematically decouples operator optimization into spatial tiling and instruction orchestration. arxiv.org/pdf/2608.10803
110
Underfox @underfox3.bsky.social · 12/08/2026
In this paper is proposed ReVolt, a dynamic operation unit (OU)-based framework for mitigating voltage droop violations in a PIM-based 2.5D multi-chiplet system, enabling both proactive and reactive OU adjustments that jointly minimize EDP. arxiv.org/pdf/2608.08496
110
Underfox @underfox3.bsky.social · 12/08/2026
In this paper is presented a controlled, Nsight-Compute-instrumented comparison of hand-written PTX Tensor-Core GEMM kernels against WMMA baselines across FP16, INT8, and INT4 on an NVIDIA L4. arxiv.org/pdf/2608.10103
110
Underfox @underfox3.bsky.social · 12/08/2026
In this paper, researchers have demonstrated wafer-scale monolithic 3D integration of three tiers of atomic-layer-deposited (ALD) indium oxide (InOx)-based devices, including ferroelectric, enhancement-mode, and depletion-mode FETs, on 200 mm silicon wafers. arxiv.org/pdf/2608.09508
110
Underfox @underfox3.bsky.social · 11/08/2026
Circular financing goes brrr... 🔥🔥🔥 NVIDIA partners with financial giants to mobilize $500 billion in AI infrastructure. www.hpcwire.com/off-the-wire...
hpcwire.com
NVIDIA Partners with Financial Giants to Mobilize $500B for AI Infrastructure - HPCwire
SANTA CLARA, Calif. and NEW YORK, Aug. 11, 2026 — NVIDIA has announced strategic partnerships to establish independent compute financing platforms with Apollo, BlackRock, Blackstone, Brookfield, Goldm...
051
Underfox @underfox3.bsky.social · 11/08/2026
In this paper, researchers have proposed a novel memory controller architecture and a RISC-V instruction set extension to optimize MLC NVM write operations by balancing speed and retention time. arxiv.org/pdf/2608.06725
130
Underfox @underfox3.bsky.social · 11/08/2026
In this paper is proposed G-Power, an architecture-level GPU power modeling framework that utilizes additional known chips to provide additional knowledge. arxiv.org/pdf/2608.06870
110
Underfox @underfox3.bsky.social · 10/08/2026
In this paper is presented HLSmith, an expert-guided agentic framework for translating C/C++ programs into optimized HLS accelerators. arxiv.org/pdf/2608.06791
110
Underfox @underfox3.bsky.social · 10/08/2026
In this paper is proposed Oz-FP4, a method for emulating FP64 DGEMM by constructing, on FP4 Tensor Cores, Ozaki schemes I and II, which realize high-precision matrix multiplication on low-precision arithmetic units. arxiv.org/pdf/2608.06812
110
Underfox @underfox3.bsky.social · 10/08/2026
In this paper, researchers have developed a self-consistent contact resistance model for metal-2D semiconductor-metal devices to capture the essential interface physics for channel lengths ranging from 5 to hundreds of nanometers. arxiv.org/pdf/2608.06793
130
Underfox @underfox3.bsky.social · 03/08/2026
Researchers have conducted a systematic characterization of LLM kernel memory access patterns through the lens of multi-partition NUMA effects, introducing a memory trace analysis methodology to derive workgroup-level data access and sharing behavior. arxiv.org/pdf/2607.28824
110
Underfox @underfox3.bsky.social · 31/07/2026
In this paper, Huawei researchers have proposed the first end-to-end FP4 reinforcement learning post-training framework for LLMs, in which both the rollout and training policies, including their forward and backward passes, operate at 4-bit precision. arxiv.org/pdf/2607.26515
110
Underfox @underfox3.bsky.social · 30/07/2026
In this paper are proposed two low-overhead instruction-reuse mechanisms for the NEORV32 RISC-V processor: a dynamic short-backward-branch-based loop cache and a software-managed static hot-code buffer with program-counter range matching. arxiv.org/pdf/2607.22792
110
Underfox @underfox3.bsky.social · 30/07/2026
In this paper, researchers have demonstrated the first monolithic non-volatile photonics platform on lithium tantalate-on-insulator (LTOI) for intrinsic, ferroelectric-domain-based non-volatile phase control without an added state-retentive material. arxiv.org/pdf/2607.23247
120
Underfox @underfox3.bsky.social · 30/07/2026
In this paper is presented the Marvell Photonic Fabric™ (PF™), a photonic-CXL hybrid architecture that replaces electrical switches with a passive fiber shuffle to deliver 32 TB of shared memory across 16 hosts via a switch-free full-crossbar topology. arxiv.org/pdf/2607.27187
120
Underfox @underfox3.bsky.social · 30/07/2026
In this paper is presented a cross-gen cost model for warp divergence in Ampere, Hopper and Blackwell GPUs, showing that divergence serializes linearly with a small constant per path and no super-linear reconvergence penalty up to a full 32-way split. arxiv.org/pdf/2607.23402
010