Nvidia & Microsoft push local AI agents on Windows

Nvidia and Microsoft have announced a major push to run AI agents locally on Windows devices, unveiling the RTX Spark 1-petaflop superchip, streamlined software installers for popular open-source frameworks, and a local networking tool designed to share idle computer capacity across home networks.

The personal computer is undergoing a structural shift, moving away from traditional app launching toward background-running, autonomous software agents. Unveiled during announcements highlighting hardware and software partnerships, the strategy targets the friction that has historically kept advanced artificial intelligence locked inside cloud datacenters or restricted to specialized developer rigs.

RTX Spark Hardware and Microsoft Security Integration

At the center of the hardware rollout is the newly introduced RTX Spark superchip.

“For forty years, you launched apps. Click. Type. With RTX Spark and Microsoft Windows, you ask — and the PC does the work. RTX Spark brings everything NVIDIA has built — CUDA, RTX, our AI platform — into a single superchip. Local agents. Frontier models. Creative workflows. RTX games. All on a laptop. This is the new PC. The personal AI computer.”

NVIDIA RTX Spark
Photo: Nvidia

Jensen Huang, founder and CEO of NVIDIA

Systems built around the RTX Spark architecture are scheduled to ship in October through hardware partners including ASUS, Dell, HP, Lenovo, Microsoft Surface, and MSI, with models from Acer and GIGABYTE to follow. These slim Windows laptops and compact desktops pack up to 1 petaflop of AI compute and 128GB of unified memory. Adobe has already begun rearchitecting Photoshop and Premiere from the ground up to take advantage of the unified memory, Blackwell architecture, and TensorRT.

Because running powerful autonomous agents locally raises significant privacy and security questions, Microsoft and NVIDIA have built a dedicated runtime stack. The collaboration integrates new Windows security primitives with the NVIDIA OpenShell runtime to enforce policies and manage containerized execution.

Streamlining Local AI Agents and Boosting Inference Performance

Beyond flagship hardware, the initiative targets the complex configuration steps that previously discouraged everyday users from running open-source AI models locally. Three prominent agent applications—Hermes Agent, OpenClaw, and Perplexity Portable Computer—are introducing simplified setup routines on Windows systems equipped with qualifying NVIDIA graphics hardware.

Nvidia & Microsoft push local AI agents on Windows
Photo: NVIDIA Developer

Hermes Agent, developed by Nous Research, is introducing a one-click local setup across RTX and DGX systems on Windows. The software automatically detects the installed graphics processor, selects an appropriate model configuration, and executes it through an integrated version of llama.cpp pre-optimized with NVIDIA tuning. Similarly, OpenClaw and Perplexity Portable Computer are rolling out streamlined Windows application packages for GPUs featuring at least 24GB of VRAM.

To support heavier agentic workflows, NVIDIA has collaborated with the open-source llama.cpp and vLLM developer communities to accelerate inference.

Pooling Idle Home Computing Resources With NVIDIA PAIR

When an advanced AI prosumer deploys a local agent to handle research or coding tasks, that main agent often breaks the objective down into dozens of parallel subagent requests. Routing every call through a single local engine creates a severe execution bottleneck.

NVIDIA RTX Spark and Local AI Agents: On-Device AI Comes to Windows PCs

To solve this, NVIDIA introduced the NVIDIA Personal AI Router (PAIR) beta, an open-source virtual inference router for Windows, macOS, and Linux. PAIR automatically discovers participating computers on a private local network using mDNS and routes independent inference requests to whichever system has available capacity.

By tapping into underused home hardware—such as a secondary desktop, a workstation, or an Apple M4-equipped machine—PAIR distributes the computational load so that the primary computer remains free for gaming or creative work. The software proxies familiar interfaces like Ollama and LM Studio without requiring developers to rewrite their agent harnesses, letting users dynamically scale their available compute as client machines join or leave the network.

Photo of author

Sophie Lin - Technology Editor

Sophie is a tech innovator and acclaimed tech writer recognized by the Online News Association. She translates the fast-paced world of technology, AI, and digital trends into compelling stories for readers of all backgrounds.

চীনের জে-১০সি যুদ্ধবিমান কিনলে কী কী সুবিধা পাচ্ছে বাংলাদেশ

Leave a Comment

This site uses Akismet to reduce spam. Learn how your comment data is processed.