Apple is reportedly planning an M7 Ultra chip that will support up to 1.5 TB of unified memory.
The planned specification highlights a significant increase in memory capacity for future Apple silicon.
Apple is reportedly planning an M7 Ultra chip that will support up to 1.5 TB of unified memory.
The planned specification highlights a significant increase in memory capacity for future Apple silicon.
The article calculates which open large language models fit into the new Apple M5 series Macs (Mac mini, Mac Studio) by analyzing memory constraints rather than running benchmarks. It determines capacity based on Q4_K_M GGUF file sizes, KV cache requirements, and runtime overhead against the usable GPU memory.
Researchers present BaseRT, a native Metal inference runtime for large language models on Apple Silicon that achieves the highest reported inference throughput to date. By utilizing chip-specific kernel fusion and unified memory-aware optimization, it overcomes the overhead found in existing frameworks like llama.cpp and MLX.
A recent report indicates that Apple plans to skip the release of M6 Pro and M6 Max chips in its upcoming lineup. Instead, the company intends to fast-track the development of the M7 chip series to better support local artificial intelligence workloads. This strategic shift suggests a prioritization of on-device AI capabilities over traditional performance increments for the Pro tier. The decision reflects Apple's growing emphasis on integrating advanced machine learning features directly into its hardware architecture. By accelerating the M7 timeline, Apple aims to provide more robust neural engine performance for running large language models locally. This move signals a significant pivot in Apple Silicon's development roadmap toward AI-centric design principles.
The llama.cpp project released version b11160, introducing an int8 cooperative matrix (coopmat1) implementation for Vulkan on AMD RDNA3 and RDNA4 architectures. This update includes a dedicated shader that probes coopmat values directly to improve performance.
At Meta Connect 2026, Meta presented its personal agent Muse as the center of a hardware-plus-agent strategy, introducing new features and devices while teasing a future frontier model. The company showcased Muse's expanded capabilities, including voice and real-time video interactions, alongside new hardware like VR glasses and the Charm keychain.
We use cookies to measure traffic and improve the site. You can accept or decline analytics cookies. Privacy policy