Recent
1346 →
vLLM Launcher provides a Windows desktop workbench for local LLM inference
llama.cpp b10331 fixes server get_info to report correct isolate working directory
Anthropic makes auto mode the default in Claude Code for Pro, Max, and Team plans
+1343 more
Inference efficiency
337 →
vLLM Launcher provides a Windows desktop workbench for local LLM inference
llama.cpp b10331 fixes server get_info to report correct isolate working directory
llama.cpp b10330 fuses rms_norm, mul, and rope on CUDA
+334 more
AI agents
320 →
Pokee AI releases Pokee-Isaac 28B, a 10M-token context model for in-boundary deployment
OpenAI models hacked infrastructure during training, then attacked HuggingFace
Claude Code v2.1.225 adds gateway spend-limit warnings and workspace trust prompts
+317 more
Research paper
275 →
U.S. Department of Energy launches Genesis Open Models Initiative and unveils Genesis-Science-1
U-OPSD enables unsupervised on-policy self-distillation for LLMs
M$^3$R-Bench introduces evidence-grounded benchmark for multimodal metaphor understanding
+272 more
Open weights
189 →
Mistral AI releases Shieldstral 1.0 3B, a policy-adaptive multimodal safety classifier
U.S. Department of Energy launches Genesis Open Models Initiative and unveils Genesis-Science-1
NVIDIA releases NOOA, an object-oriented Python framework for building AI agents
+186 more
Code generation
122 →
Anthropic makes auto mode the default in Claude Code for Pro, Max, and Team plans
Claude Code v2.1.225 adds gateway spend-limit warnings and workspace trust prompts
LoopTroop: local GUI for long AI coding tasks using LLM councils and Ralph loops
+119 more
Evaluation & benchmarks
118 →
TutorMoments evaluates whether AI tutors know when to help or hold back
M$^3$R-Bench introduces evidence-grounded benchmark for multimodal metaphor understanding
SciCode-Verified corrects benchmark defects to reveal models' true scientific-coding ability
+115 more
API & product launches
109 →
Pokee AI releases Pokee-Isaac 28B, a 10M-token context model for in-boundary deployment
xAI releases Imagine Image 2.0 with precise editing and layout capabilities
LiteLLM 1.15.13 fixes provider preservation and Anthropic cache token underreporting
+106 more
Safety & alignment
107 →
Anthropic makes auto mode the default in Claude Code for Pro, Max, and Team plans
OpenAI models hacked infrastructure during training, then attacked HuggingFace
Mistral AI releases Shieldstral 1.0 3B, a policy-adaptive multimodal safety classifier
+104 more
Training methods
105 →
U-OPSD enables unsupervised on-policy self-distillation for LLMs
Apple's AFM3 20B uses instruction-following pruning to activate ~20% of layers
NVIDIA NeMo releases Molt, a compact PyT-native agentic RL framework
+102 more
Hardware & chips
73 →
llama.cpp b10306 adds SYCL GLU flat path and consolidates kernels
llama.cpp b10255 extends oneDNN SDPA to non-FP16 KV caches
Lucebox and AMD beat Nvidia DGX Spark by 3.63x on DeepSeek V4 Flash
+70 more
Retrieval & RAG
72 →
Aletheion releases ASM-CM, a compact memory architecture for persistent AI agents
IntelShed combines hybrid RAG, GNN entity resolution, federated learning, and LLM multi-agent orchestration
TIS 2.0 eliminates position bias in RAG by reordering passages with token importance scores
+69 more
Benchmark results
72 →
DeepSeek V4 Flash 0731 ARC-AGI Results
DeepSeek releases V4-Flash-0731, outperforming V4-Pro at lower cost
DeepSeek-V4 Flash 0731 vs GPT-5.6 Luna on DeepSWE: Cost and Coding
+69 more
Multimodal
66 →
NVIDIA releases Alpamayo 2 Super, a 34B open VLA model for autonomous driving
Video-DeepResearch introduces multimodal agent for continuous video streams
Shieldstral introduces policy-adaptive 3B multimodal safety classifier
+63 more
Policy & regulation
45 →
OpenAI slashes Luna prices, Alibaba releases Qwen 3.8-Max, and DeepMind leadership changes
OpenAI partners with APA on youth mental health and AI safeguards
Jeff Dean, Sanjay Ghemawat, Oriol Vinyals, and Quoc Le depart DeepMind to cofound Discovery Loop
+42 more
Local inference
40 →
vLLM Launcher provides a Windows desktop workbench for local LLM inference
llama.cpp b10330 fuses rms_norm, mul, and rope on CUDA
llama.cpp b10328 adds initial server tool isolation via Docker
+37 more
Reasoning models
26 →
Meta's Muse Spark 1.2 and OpenAI's GPT-5.6 Sol updates
OpenAI's Astra solves 10 major open mathematics problems
AI9Stars releases G9v3-39A5B, an open weights model with 39B parameters and 5 active experts
+23 more
Training data
25 →
OpenGauntlet publishes catalog of 168 TTS and 106 STT systems
Monthly compliance matrix reveals 65% of 4,893 Indic datasets lack license tags
Calibra v0.7.1 adds dataset integrity checks for robot learning
+22 more
Voice & audio
23 →
llama.cpp b10270 adds Qwen3-TTS support with breaking llama-tts changes
NVIDIA releases NemotronLabs-VoiceChat-11B model on Hugging Face
OpenAI engineers GPT-Live, a full-duplex voice system that removes turn detectors
+20 more
Robotics
20 →
ROSA improves factory productivity by up to 12.06x via shared GPU-pool serving
Calibra v0.7.1 adds dataset integrity checks for robot learning
Google DeepMind releases Gemini Robotics 2 for whole body control and dexterity
+17 more