Please confirm you are human
This browser or connection looks automated. Press and continuously hold the control for 3 seconds to enable Google-hosted web results and, when separately allowed, AI-assisted answers.
A successful check enables 100 search requests. Interactive access does not authorize scraping, systematic collection, or reuse of search output.
News
Serve Qwen3.8-2.4T-A95B, a 2.4T-Parameter Model, with Configurable Reasoning on NVIDIA GB300 NVL72
1+ hour, 24+ min ago (404+ words) Deploying a 2.4T parameter open-weight model requires data-center-scale accelerated compute. Inference at this scale depends on extreme co-design across chips, system architecture, and software. NVIDIA is working with the open-source ecosystem to bring the model to multinode deployments through optimized kernels,…...
NVIDIA JetPack 7.2.1 Adds Agentic Video Skills and T3000 Emulation
1+ day, 2+ hour ago (504+ words) Video is a core data path across NVIDIA Jetson applications, from robotics and intelligent video analytics to industrial automation, healthcare, media processing, and remote operations. A system may capture several cameras, decode network streams, run AI inference or conventional vision…...
NVIDIA Nemotron 3.5 Lightning Delivers Fast, Accurate Specialized Task Execution for Long-Running Agents
1+ day, 6+ hour ago (832+ words) Long-running AI agents spend most of their time on high-volume execution: tool calls, result validation, and subagent delegation. Using a frontier reasoning model for every execution step adds cost and latency. NVIDIA Nemotron 3.5 Lightning is an open 30B mixture-of-experts (MoE) model…...
Route AI Agents Across Models with NVIDIA NeMo Switchyard
1+ day, 6+ hour ago (1024+ words) Learn how NVIDIA NeMo Switchyard routes agent workloads across specialized and frontier models to balance performance, cost, and efficiency. Model routing addresses this challenge by orchestrating specialized and frontier models so that each task uses the model best suited to…...
Run Local Agentic AI Workflows with Meta???s Muse Glimmer on NVIDIA
2+ day, 6+ hour ago (444+ words) Meta returns to the open source ecosystem with the release of Muse Glimmer, a 30B open-weight dense model with a 120K+ context window built for local AI agentic work. Optimized to run across a range of NVIDIA edge, desktop, and workstation AI…...
Beyond VLAs: How World Action Models Reshape Robot Manipulation
1+ week, 5+ day ago (242+ words) This post explores how post-training can turn WAMs into specialized robot policy, how WAMs compare to VLAs, and why the open NVIDIA Cosmos 3 world model provides a strong foundation for building WAMs. In the VLA paradigm, a pretrained VLM provides…...
How to Run Isolated Tenant Kubernetes Clusters on Shared GPU Infrastructure
1+ week, 4+ day ago (836+ words) This post provides a pattern that preserves team autonomy without splitting the hardware. This solution involves a single control plane cluster with a GPU pool, GPU sharing with per-team quotas, and isolated Kubernetes control plane per team including an API…...
Co-Designing AI Model Attention for Fast, Interactive Long-Context Inference
1+ week, 4+ day ago (905+ words) As agentic and long-context workloads become common, the context lengths increase and attention consumes a larger share of inference time (Figure 1)....
NVIDIA Video Codec SDK 13.1: Zero-Copy Transcode, AV1 B-Frames, and Frame-Accurate Seek
1+ week, 5+ day ago (907+ words) Behind these experiences is a growing need for video pipelines that are faster, more efficient, and capable of handling increasingly complex formats and workloads. NVIDIA Video Codec SDK helps developers meet that challenge by providing access to GPU-accelerated video encoding…...
Four Ways to Deploy More Secure AI Agents
1+ week, 5+ day ago (731+ words) The NVIDIA AI Red Team shares how access controls, sandboxing, network egress restrictions, and secret management can reduce the risks of deploying AI agents at enterprise scale. Over the past six months, the NVIDIA AI Red Team has assessed multiple…...