Google Images introduces new gallery and Nano Banana image generation
To mark its 25th anniversary, Google is launching a redesigned, browseable home for Google Images alongside new image generation capabilities within AI Overviews.
To mark its 25th anniversary, Google is launching a redesigned, browseable home for Google Images alongside new image generation capabilities within AI Overviews.
Researchers at NVIDIA have introduced Spatial-IQ, a diagnostic framework designed to deconstruct the spatial intelligence of multimodal large language models (MLLMs). Unlike existing benchmarks that treat models as black boxes, this approach decomposes object counting in stacked 3D structures into nine perceptual and cognitive sub-tasks aligned with human developmental stages.
Researchers analyze gameplay videos from the competitive shooter game Delta Force to extract and understand emergent non-verbal communication, such as gestures and movement patterns. This work addresses a gap in prior studies that have largely focused on MOBA games or verbal coordination in other shooters.
At Galaxy Unpacked 2026, Google announced three updates for the new Samsung Galaxy Z Fold8 Ultra, Fold8, and Flip8, centering on the rollout of Gemini Intelligence. The company is expanding its task automation capabilities to support over 40 popular apps and introducing the Gemini Notebook app directly to these devices.
Researchers developed an immersive virtual reality task to assess pedestrian collision detection and avoidance in individuals with cerebral visual impairment (CVI) compared to control participants. The study tracked eye gaze, locomotor, and behavioral responses as subjects walked through a simulated shopping mall populated with crowds of varying densities.
Shieldstral is a 3B open-weights multimodal safety classifier that frames content moderation as a policy-adaptive question-answering task, allowing it to accept plain-language policies at inference time without retraining. It unifies text and image safety evaluation and delivers calibrated safety scores across diverse benchmarks while running efficiently on a single 16GB NVIDIA GPU.
On July 19, Alibaba's Qwen team previewed Qwen3.8-Max-Preview, a new flagship multimodal model with 2.4 trillion parameters. The announcement was made during the World AI Conference in Shanghai, arriving two days after Moonshot AI released its Kimi K3 model.
Thinking Machines has released Inkling, a large open-source multimodal language model available on Hugging Face that natively processes text, image, and audio inputs. The model features approximately 1 trillion total parameters with 41 billion active parameters, trained on 45 trillion tokens to support reasoning across modalities.
Thinking Machines Lab has released Inkling, a new multimodal mixture-of-experts model designed for token-efficient reasoning and native understanding of text, image, and audio inputs. Together AI is making the model available on its inference platform with day-zero access.
Tencent Robotics X, Futian Laboratory, and Tencent Hunyuan have released the technical report, inference code, and weights for Hy-Embodied-RxBrain-1.0, a ~6.2B-parameter unified multimodal foundation model for embodied cognition.
Tencent has released Hy-Embodied-VLM-1.0, an efficient Mixture-of-Experts vision–language foundation model designed for embodied agents operating in the physical world. The model activates only approximately 3 billion parameters per token out of a total of 30 billion, aiming to balance high inference efficiency with strong physical-world understanding.
NVIDIA has released Alpamayo 2 Super, a 34-billion-parameter vision-language-action (VLA) model designed for robotaxis and autonomous driving under the OpenMDW-1.1 license. The model targets long-tail events by combining a 32B Cosmos 3 Super Reasoner backbone with a 2.3B diffusion-based action decoder to generate trajectories, causal explanations, and meta-actions from multi-camera video.
Researchers introduce Video-DeepResearch (Video-DR), a framework extending multimodal agents from static images to continuous video streams, addressing modality bias and parametric knowledge leakage. The system employs a decoupled perception-exploration pipeline with stage-wise tool unlocking to enforce cross-frame visual grounding before web retrieval.
Douyin has released the technical report for its Multimodal Embedding (DME) model, which combines large-scale contrastive pre-training with latent reasoning to achieve strong discrimination and efficiency.
The article highlights Qwen3.8-Max's performance in a one-shot evaluation context, noting its ability to handle complex physical simulations.
MiniMax has made its MiniMax-H3 model available on Hugging Face. This general-purpose, omni-modal generative system supports the unified understanding of multimodal contexts composed of text, images, video, and audio.
Thinking Machines Lab has released Inkling-Small, an open-weights Mixture-of-Experts model with 276 billion total parameters and 12 billion active parameters. The model is trained to reason natively over text, images, and audio, featuring a 1M token context window and adjustable thinking effort.
The latest roundup of open model artifacts highlights significant releases from Thinking Machines and Tencent, alongside updates from Poolside and DeepSeek.
Hanyu (Olivia) Yu, a high school sophomore, is seeking an arXiv endorser in cs.LG or cs.CV to upload her independent preprint titled "AttnRoute-MoE: Attention-Prior Routing for Mixture-of-Experts Vision Transformers."
The llama.cpp project released build b10218, which includes the addition of minicpmv46 downsample support via pull request #25993. This update introduces a 4x ignore VIT merger and places the downsample mode inside the GGUF format.