Underfox @underfox3.bsky.social · 03/09/2026The simulation results show that CHIPSMORE achieves up to 2.38× higher throughput and 27× higher energy efficiency on Mistral-7B inference compared with Nvidia H100 while eliminating weight replication for multi-request serving. 020
Underfox @underfox3.bsky.social · 03/09/2026Finally, a non-replicated multi-request execution pipeline exploits request-level parallelism without duplicating pretrained weights, while a state-aware resource reconfiguration mechanism selectively retains runtime states and power-gates inactive resources. 110
Underfox @underfox3.bsky.social · 03/09/2026A composable hierarchical KV memory scheme dynamically allocates router scratchpad, SRAM-DCIM, and eDRAM resources according to workload requirements, enabling scalable support for long-context and batched inference. 100
Underfox @underfox3.bsky.social · 03/09/2026CHIPSMORE employs heterogeneous processing elements consisting of resistive RAM analog CIM and static RAM digital CIM interconnected through a programmable Inter-PE computational network. 100
Underfox @underfox3.bsky.social · 03/09/2026In this paper is presented CHIPSMORE, a multi-mode and multi-request CIM accelerator for LLM inference across diverse operating scenarios, including base-mode execution, LoRA adaptation, varying context lengths, and concurrent multi-request serving. arxiv.org/pdf/2608.30509 120
Underfox @underfox3.bsky.social · 03/09/2026The evaluation results on a GPU kernel in the NextLA linear algebra library show that this approach achieve performance improvements of 3× to 7× over median parameter configurations. 010
Underfox @underfox3.bsky.social · 03/09/2026In this paper is introduced the first auto-tuning framework for Julia GPU kernels as an accessible Julia package, enabling systematic auto-tuning of hardware-agnostic kernels to target a wide variety of platforms, such as NVIDIA, AMD, Intel, and Apple. arxiv.org/pdf/2608.21227 120
Underfox @underfox3.bsky.social · 03/09/2026The evaluation results show that, in the memory-constrained regime, NeuroPrefetcher achieves 7.9–12.0x speedup over llama.cpp by replacing page-fault-driven whole-weight movement with application-scheduled sparse-row reads. 010
Underfox @underfox3.bsky.social · 03/09/2026NeuroPrefetcher makes weight movement explicit, predicting which sparse weights the next token will need, compares that predicted set against the weights already resident, and fetches only the difference in advance. 110
Underfox @underfox3.bsky.social · 03/09/2026In this paper is proposed NeuroPrefetcher, a storage-backed sparse inference runtime built around predictive delta prefetching for the regime where an LLM exceeds resident memory throughout execution. arxiv.org/pdf/2608.22643 110
Underfox @underfox3.bsky.social · 02/09/2026The experimental results shows close agreement with ML predictions, thus demonstrating the framework’s ability to efficiently guide gatestack optimization through iterative experimental feedback. 010
Underfox @underfox3.bsky.social · 02/09/2026These metrics are integrated into a multi-objective ranking framework and coupled with GPR to identify process–performance correlations and predict unexplored fabrication recipes within the experimental process space. 110
Underfox @underfox3.bsky.social · 02/09/2026The introduced transition voltage metric, V_Trans, complements conventional transistor metrics and enables more comprehensive process performance optimization. 100
Underfox @underfox3.bsky.social · 02/09/2026Researchers have presented the first experimental ML-enabled Process-Technology Co-Optimization (PTCO) framework for optimizing 2D TMD FET fabrication directly from statistically meaningful experimental data rather than pure simulation data. arxiv.org/pdf/2609.00722 120
Underfox @underfox3.bsky.social · 02/09/2026Finally, the results show that the scale-dependent effects, including SMT dispatch contention causing throughput reduction and L3 capacity interference, emerge only at full system utilization. 010
Underfox @underfox3.bsky.social · 02/09/20262 - High-efficiency compute workloads achieve exceptional dispatch utilization but suffer corresponding SMT contention at scale; 3 - Memory bandwidth-bound floating-point workloads exhaust DRAM bandwidth at scale while exhibiting low L3 hit rates even at single-copy. 110
Underfox @underfox3.bsky.social · 02/09/2026The multi-lens analysis reveals distinct behavioral clusters: 1 - Front-end control-flow-dominated integer workloads stress branch predictor throughput rather than accuracy; 110
Underfox @underfox3.bsky.social · 02/09/2026The results show that SPEC CPU 2026 exhibits substantial behavioral diversity, with IPC spanning a 3.3x range across benchmarks. 110
Underfox @underfox3.bsky.social · 02/09/2026The proposed characterization employs a multi-perspective methodology covering pipeline efficiency, control-flow behavior, cache hierarchy pressure, and instruction composition, analyzing both SPECrate and SPECspeed suites. 110
Underfox @underfox3.bsky.social · 02/09/2026In this paper, AMD researchers have presented the first micro-architecture based performance characterization of SPEC CPU 2026 on AMD EPYC "Zen 5" processor. arxiv.org/pdf/2609.01527 172
Underfox @underfox3.bsky.social · 02/09/2026GitHub GapFill marc2825.github.io/GapFill/marc2825.github.io[CHI '26] GapFill; 塗り残し解消ツール (Project Page)[CHI '26] GapFill はプロのアニメ彩色担当者が直面する, 小さな「塗り残し」への対処を支援するために設計された専用ツールです; [No Pixel Left Behind: Filling Gaps in Anime Colorization]; [GapFill: アニメ調彩色における塗り残しの解消を支援するツール];Animation production workflows ... 010
Underfox @underfox3.bsky.social · 02/09/2026The results also showing that usability may be driven by the combination of automatic detection and clear visualization of gaps, as well as user control over AI suggestions rather than prediction accuracy alone. 110
Underfox @underfox3.bsky.social · 02/09/2026The evaluation results with with 13 professional colorists demonstrated significant efficiency gains in gap-filling tasks compared to existing tools. 110
Underfox @underfox3.bsky.social · 02/09/2026GapFill enables automatic gap detection with circular highlights, along with temporary filling based on color suggestions using a domain-specific deep learning method. By hovering over a highlight, the corresponding region can be magnified for quick inspection without zooming in. 110
Underfox @underfox3.bsky.social · 02/09/2026In this paper is proposed GapFill, a specialized tool designed to assist professional anime colorists in addressing small, unpainted areas, integrating seamlessly into existing professional workflows. arxiv.org/pdf/2609.00800 110
Underfox @underfox3.bsky.social · 02/09/2026[Basic Concepts] In this paper is presented a comprehensive and critical overview of on-chip photonic neural networks, spanning fundamental device technologies, system architectures, and emerging applications. advanced.onlinelibrary.wiley.com/doi/epdf/10.... 020
Underfox @underfox3.bsky.social · 01/09/2026Overall, these results demonstrated a clear and rapidly accelerating path to the viability of RVV hardware in HPC in the near future. 020
Underfox @underfox3.bsky.social · 01/09/2026On the other hand, the FFTW evaluation highlights the distinct advantages of out-of-order execution, favoring the Sophon SG2044 and SpacemiT X100 cores for complex memory access patterns. 110
Underfox @underfox3.bsky.social · 01/09/2026The SpacemiT K3, however, emerges as a major exception and represents a profound generational leap over the preceding K1. By significantly improving its vector front-end and memory bandwidth, the K3 achieves compute efficiencies of up to 80% on both the X100 and A100 cores. 100
Underfox @underfox3.bsky.social · 01/09/2026These architectural limitations become apparent in the BLAS evaluation, where most platforms struggle to overcome the memory wall, yielding compute efficiencies between 30% and 50%. 100
Underfox @underfox3.bsky.social · 01/09/2026The results show that while RVV 1.0 delivers significant performance improvements over scalar execution, hardware-specific implementation challenges remain. 100
Underfox @underfox3.bsky.social · 01/09/2026In this paper is presented a benchmark of the latest RVV 1.0-capable hardware (SiFive X280, SpacemiT X60 (K1) and X100/A100 (K3), and T-Head C920v2 (Sophon SG2044)), using standard HPC benchmarks and synthetic workloads, comparing them to NVIDIA Grace. #HPC arxiv.org/pdf/2608.28097 270
Underfox @underfox3.bsky.social · 29/08/2026Overall, this work demonstrated that there is rapid advancement toward making GPU-accelerated scientific HPC code performance portable using the Fortran standard language. 030
Underfox @underfox3.bsky.social · 29/08/2026The results also show that the three GPU vendors can now GPU-accelerate pure Fortran (zero directives), but that manual data movement directives can help with performance and compatibility. 111
Underfox @underfox3.bsky.social · 29/08/2026On Intel GPUs, the pure Fortran code ran over 50% slower, however, the Intel unified memory feature is extremely new and is expected to improve with subsequent compiler updates. 100
Underfox @underfox3.bsky.social · 29/08/2026The results show that on NVIDIA and AMD GPUs, the two versions had very similar performance, even on multiple GPUs. 100
Underfox @underfox3.bsky.social · 29/08/2026The proposed code uses the Fortran language’s do concurrent loop construct, which allows the compilers to parallelize the loop on both multi-threaded CPUs and GPUs. Through the use of GPU-aware MPI, it was also able to run the code across multiple GPUs. 110
Underfox @underfox3.bsky.social · 29/08/2026In this paper, researchers have analyzed the current status of using do concurrent for GPU-accelerated Fortran applications across three major GPU vendors (NVIDIA, AMD, and Intel). arxiv.org/pdf/2608.20586 131
Underfox @underfox3.bsky.social · 29/08/2026The evaluation results on NVIDIA A100, H800, and Hygon DCU Z100 show that HSMA-TRSM achieves peak speedups of 2.05× over cuBLAS and 2.06× over rocBLAS. 010
Underfox @underfox3.bsky.social · 29/08/2026For large-scale problems, a diagonal-block decoupling optimization was introduced with an O(I_B) shared-memory footprint for diagonal block inversion, enabling adaptive block size selection based on matrix scale and hardware characteristics. 110
Underfox @underfox3.bsky.social · 29/08/2026For the small-scale regime, a pipelined compute-memory overlap mechanism was designed through loop unrolling and instruction reordering, and a dual thread-group seven-stage pipeline strategy was proposed to address shared memory constraints for double complex types. 100
Underfox @underfox3.bsky.social · 29/08/2026In this paper is presented HSMA-TRSM, a hierarchical shared memory-aware optimization framework for left-side lower-triangular triangular solve with multiple right-hand sides on NVIDIA A100, NVIDIA H800, and Hygon DCU Z100 accelerators. #HPC arxiv.org/pdf/2608.25469 110
Underfox @underfox3.bsky.social · 27/08/2026If I have to choose between migrating to Windows 11 or buying a GPU from another company, I will certainly choose an Intel GPU. videocardz.com/newz/nvidia-...videocardz.comNVIDIA ends Windows 10 Game Ready support in October - VideoCardz.comGeForce Windows 10 support moves to quarterly security updates after October. NVIDIA plans to provide critical fixes through October 2029. 031
Underfox @underfox3.bsky.social · 21/08/2026The evaluation results show that FIBER delivers 2.25×, 1.8×, and 2.09× end-to-end speedup under typical mixed-precision LLM serving scenarios against Ampere, Hopper, and Blackwell baselines, with kernel-level speedups of up to 2.49×. 010
Underfox @underfox3.bsky.social · 21/08/2026FIBER introduces a new parallel execution instance, the fiber, analogous to a SIMT thread but without private register ownership, carrying only minimal state. Unlike SIMT threads, fibers access an SM’s register file through a shared view. 110
Underfox @underfox3.bsky.social · 21/08/2026In this paper is presented FIBER, a GPU architecture that decouples register ownership from parallel execution instances, comprising ISA extensions, lightweight microarchitectural enhancements, and a CUDA-compatible programming model. arxiv.org/pdf/2608.19628 110
Underfox @underfox3.bsky.social · 21/08/2026The major gains that this approach introduces so far are a strict separation of concerns, support for rapid prototyping of novel numerics, cleaner code, and seamless support for both CPU and GPU kernels. 010
Underfox @underfox3.bsky.social · 21/08/2026It's important to note that this approach yields sufficiently efficient code, but does not yet unlock the full potential of a compiler-/DSL-based kernel encoding. 110
Underfox @underfox3.bsky.social · 21/08/2026The proposed approach keeps the numerical representation and the physics implementation separate for as long as possible, while delegating optimization to the compiler through existing MLIR optimization passes. 100
Underfox @underfox3.bsky.social · 21/08/2026In this paper is introduced a bilingual DSL for modeling compute kernels within a generic solver for hyperbolic PDEs, enabling programmers to express their numerics in an abstract manner in Python, with the physical PDE terms still in C, C++, or SymPy. arxiv.org/pdf/2608.19273 110