AoiZora is a compiler-mediated topology planner that improves low-latency video diffusion inference on TPU sub-slices. By aligning logical sharding with physical placement through the compilation flow, it reduces one-step denoising latency by up to 1.42x on TPU v5e sub-slices compared to existing methods.
AoiZora: Topology-Aware Auto-Parallel Optimization for Video Diffusion Inference
Nvidia presents FusionRelight for real-time portrait relighting via hybrid domain knowledge fusion
Researchers from Nvidia have introduced FusionRelight, a method for portrait relighting that combines physically plausible illumination transfer with identity preservation and compact real-time inference. The approach utilizes Hybrid Domain Knowledge Fusion (HDKF), a training framework that distills physics, reflectance, and realism priors from synthetic, One-Light-at-a-Time (OLAT), and in-the-wild data into a student model.
Lucebox and AMD beat Nvidia DGX Spark by 3.63x on DeepSeek V4 Flash
Lucebox has partnered with AMD to demonstrate a heterogeneous consumer hardware setup that outperforms the Nvidia DGX Spark in inference speed. The system combines an AMD Radeon AI PRO R9700 GPU with a Strix Halo processor to run the full 284B DeepSeek V4 Flash model.
Clustered Codebook Quantization for 2D Gaussian-based Image Compression
Researchers present Cluster-Guided Vector Quantization (CGVQ), a method designed to improve the rate-distortion performance of Gaussian primitive-based image compression. The approach partitions Gaussian parameters into homogeneous groups prior to quantization, addressing the inefficiency caused by storing large numbers of floating-point parameters per primitive.
NVIDIA tutorial demonstrates tile-based GPU programming with cuTile and Triton
A new tutorial explores NVIDIA's TileGym framework by building a practical Colab workflow that adapts to different hardware conditions. The guide probes the CUDA environment to determine if NVIDIA cuTile can run directly, falling back to Triton when standard Colab GPUs lack the required stack.
Shunt-modded RTX 6000 PRO MaxQ outperforms RTX 5090 in Anima and LLM benchmarks
A user compares the full compute performance of an RTX 5090 against shunt-modded and water-cooled RTX 6000 PRO MaxQ cards, as well as a rented RTX 6000 PRO WS edition. The test evaluates both Anima benchmark scores and LLM prompt processing speeds across various power limits.