Recent
803 →
WebMRE benchmark reveals guide-action mutual reinforcement in web agents
SkillGym internalizes human skills into LLMs for real-world problem solving
Anthropic releases Claude Opus 5.5 system card detailing safety and capabilities
+800 more
Inference efficiency
287 →
llama.cpp v0.5.0 adds HRM-Text support, CUDA conv2d acceleration, and multi-address binding
Microsoft Research shows offloading physical AI inference improves robot performance and battery life
llama.cpp b11147 adds A8 Q6_K non-MoE dp4a OpenCL kernel
+284 more
AI agents
225 →
WebMRE benchmark reveals guide-action mutual reinforcement in web agents
SkillGym internalizes human skills into LLMs for real-world problem solving
Airbnb expands access to OpenAI GPT-6 Astra and other frontier models
+222 more
API & product launches
122 →
Google releases Gemini 3.8 Flash TTS models with custom voice cloning
OpenAI releases GPT-6 Sol and Luna; Anthropic launches Claude Opus 5.5
Google introduces Gemini 3.8 Flash and Flash-Lite text-to-speech models
+119 more
Open weights
103 →
NVIDIA releases Nemotron 3 Diarization, a 100M-parameter model tracking 8 speakers
Apple releases LensVLM-9B, a VLM that selectively expands compressed text images
NVIDIA releases open-weight Nemotron 3 Diarization model
+100 more
Local inference
103 →
llama.cpp v0.5.0 adds HRM-Text support, CUDA conv2d acceleration, and multi-address binding
llama.cpp b11136 adds OpenAI video_url support to llama-server
Sfeir releases Esus to offload deterministic tasks from small LLMs
+100 more
Safety & alignment
93 →
Anthropic releases Claude Opus 5.5 system card detailing safety and capabilities
Sam Altman urges UN Security Council on AI safety and international cooperation
OpenAI releases MentalHealthBench to evaluate AI mental health responses
+90 more
Code generation
83 →
Airbnb expands access to OpenAI GPT-6 Astra and other frontier models
Goose v1.52.0 adds live voice conversations and new model support
Claude Code v2.1.281 adds gateway features and fixes session stability
+80 more
Research paper
59 →
Anthropic's Claude discovers novel ART enzyme system in bacteriophages
AIDE^2 enables recursive self-improvement of AI research agents
Growing Harness turns recurring agent control into reusable code via failure-guided training
+56 more
Benchmark results
40 →
OpenAI releases GPT-6 Sol and Luna; Anthropic launches Claude Opus 5.5
OpenAI releases GPT-6 Sol and Luna with 50% lower API prices
Anthropic launches Claude Opus 5.5 with improved writing and lower costs
+37 more
Hardware & chips
37 →
Alibaba plans AI model with 5 trillion to 10 trillion parameters and unveils new chip
Huawei shelves global AI chip rollout as China's own demand outstrips supply
Proposal for progressive model growth and an open hardware profile for Transformers
+34 more
Training methods
37 →
SkillGym internalizes human skills into LLMs for real-world problem solving
Together AI enables fine-tuning a Jev-like classifier on Qwen3.5 4B for $17
NVIDIA Warp and MjWarp enable GPU-accelerated parallel robotics simulation
+34 more
Evaluation & benchmarks
35 →
WebMRE benchmark reveals guide-action mutual reinforcement in web agents
Anthropic releases Claude Opus 5.5 system card detailing safety and capabilities
OpenAI releases MentalHealthBench to evaluate AI mental health responses
+32 more
Multimodal
31 →
Apple releases LensVLM-9B, a VLM that selectively expands compressed text images
Qwen releases Qwen3.8-Omni-Flash, a native multimodal agentic model with 1M token context
Alibaba Qwen Team releases Qwen3.8-LiveTranslate with Interleave architecture
+28 more
Voice & audio
28 →
NVIDIA releases Nemotron 3 Diarization, a 100M-parameter model tracking 8 speakers
Google releases Gemini 3.8 Flash TTS models with custom voice cloning
Google introduces Gemini 3.8 Flash and Flash-Lite text-to-speech models
+25 more
Retrieval & RAG
25 →
Together AI enables fine-tuning a Jev-like classifier on Qwen3.5 4B for $17
YUCLAW 8.0.0 adds explicit snapshot mode for reproducible evidence demos
jbsalles experiments with SelMem, a selective memory system for LLMs
+22 more
Policy & regulation
24 →
Sam Altman urges UN Security Council on AI safety and international cooperation
DeepSeek and Moonshot AI face Beijing probe over potential data leaks to Anthropic
Lawsuit alleges Anthropic, OpenAI, SpaceX AI and Google colluded on AI slowdown
+21 more
Video generation
16 →
invideo improves color grading 3x with GPT-6 Astra
VideoX-Qwen introduces data-centric framework for instruction-based video editing
Higgsfield AI uses GPT-6 Astra to ship new video features in a day
+13 more
Reasoning models
13 →
TelecomGPT-R1: Unified Post-Training for Reasoning Across Heterogeneous Telecom Tasks
CosmicUndercurrent releases Principia-Structurae-Realitatis for testing LLM whole-system coordination
Proposal for a proposition decomposition and contradiction-driven reasoning system
+10 more
Image generation
12 →
Meta opens Muse connectors, releases SAM 3.1, and Gemini hacks systems
Alibaba Qwen releases unified 7B Qwen-Image-2.1 model for image generation and editing
Qwen releases open-weight Qwen-Image-2.1 image generation model
+9 more