Home
News

Stop Relying On The Cloud: NVIDIA RTX Spark Changes How Local AI Agents Run

NVIDIA is pushing local AI agents closer to mainstream PC users, with new tools and hardware support announced at IFA 2026 alongside Microsoft and several software partners. The pitch is clear: users should be able to run capable AI agents on their own machines, with less setup work, faster inference and tighter control over private data.

The announcements centre on NVIDIA RTX and DGX systems, including new compact RTX Spark Windows PCs expected in October 2026 from partners such as Lenovo and Acer. NVIDIA is also bringing simplified local model support to widely used agent apps, improving llama.cpp and vLLM performance, and introducing a tool that can share AI workloads across multiple PCs on the same local network.

NVIDIA Takes Aim at Cloud AI With RTX Spark PCs

Local AI Agents Get Easier to Set up on NVIDIA GPUs

Running an AI agent locally has usually required a fair amount of technical knowledge. Users often had to choose a model, find a compatible inference server, adjust quantisation settings and troubleshoot performance. NVIDIA says that process is being simplified across Windows systems using RTX and DGX hardware.

Hermes Agent, OpenClaw and Perplexity Portable Computer are among the first agent platforms getting smoother local setup experiences. These integrations are built around llama.cpp and include NVIDIA’s latest inference optimisations, reducing the need for manual downloads and configuration. The goal is to make local agents practical for developers, creators and advanced PC users who do not want every workflow sent to the cloud.

Perplexity’s Portable Computer agent already allows users to run Perplexity locally on Linux systems such as NVIDIA DGX Spark. It packages models, orchestration and tools inside one app experience. NVIDIA says support is coming to Windows, while RTX GPU users with at least 24GB VRAM can run it on Linux now.

The privacy angle is important. Portable Computer can complete workflows locally without consuming cloud credits. When a task needs stronger reasoning or research, it can escalate parts of the work to more than 15 frontier cloud models. Perplexity says the agent asks for permission before sending content to the cloud, which should matter for users handling code, financial documents or company data.

Hermes Agent, developed by Nous Research, is also getting one-click local model setup on Windows. The agent will detect the NVIDIA GPU, select a suitable model and configuration, and run it through integrated llama.cpp. Linux support is planned. OpenClaw’s Windows app is being tuned for RTX GPUs with at least 24GB VRAM, with NVIDIA and Microsoft working with the project to reduce onboarding friction.

Faster Inference Could Make Local Agents Feel Less Sluggish

Performance remains one of the biggest barriers for local AI. Agents often split a task into multiple steps, call tools, process documents and maintain context. If inference is slow, the whole experience can feel unreliable compared with cloud services running on large data centre clusters.

NVIDIA says new llama.cpp optimisations deliver up to 1.9x higher throughput on a GeForce RTX 5090. The improvements include kernel optimisations, faster prefill and enhanced speculative decoding. For vLLM, NVIDIA claims 1.2x performance on RTX PRO 6000 Blackwell Workstation Edition and up to 1.4x on two DGX Spark clusters.

NVIDIA Takes Aim at Cloud AI With RTX Spark PCs

These upgrades are available through llama.cpp and vLLM backends. Users can also access them through popular local AI apps such as LM Studio and Ollama. That matters because many enthusiasts and developers already use these tools to download, test and run open models without building their own inference stack.

NVIDIA has also introduced PAIR, short for Personal AI Router. The free, open-source software discovers compatible systems on a local network and routes independent inference requests to whichever device has capacity. It supports Ollama and LM Studio, and works across Windows, macOS and Linux through graphical and terminal interfaces.

PAIR is designed for homes and small workspaces where more than one computer may be available. Instead of forcing every local agent request through the same GPU, PAIR can distribute parallel jobs across idle machines. Supported hardware includes GeForce RTX 20 Series and newer GPUs, NVIDIA RTX PRO workstation GPUs based on Turing or newer, DGX Spark and Apple M4 or newer silicon.

RTX Spark Windows PCs Arrive in October

The hardware side of the announcement is NVIDIA RTX Spark, a new class of Windows PCs expected to ship in October 2026. At IFA, Acer showed a compact desktop RTX Spark concept, while Lenovo announced the Yoga Pro 9n and Yoga 9n 2-in-1. NVIDIA says six OEMs are already lined up for October shipments.

RTX Spark systems are built around a 1 petaflop RTX Blackwell GPU, up to 128GB of unified memory and a 20-core Grace CPU. NVIDIA is positioning these machines as AI PCs for creators, gamers and local agents, with support for always-on background workflows through the Windows Agent framework.

For creative users, CyberLink’s PhotoDirector AI PC Mode will be among the early software examples. Coming to PhotoDirector 365 and optimised for RTX Spark, it will offer generative editing, image enhancement, object removal, background replacement, portrait refinement and local or cloud processing choices. On NVIDIA GPUs, the app uses TensorRT-RTX and FP8 to accelerate local AI tasks.

Gaming support is also part of NVIDIA’s RTX Spark push. Electronic Arts, Embark and Ubisoft are among the latest publishers and developers bringing major titles to RTX Spark Windows PCs. They join KRAFTON, NetEase, Riot Games and Xbox, which were named during earlier announcements at COMPUTEX.

NVIDIA’s broader local AI update also includes several new open or open-weight models. Nemotron 3.5 Lightning is a 30-billion-parameter model for RTX PCs, RTX PRO workstations, DGX Spark and Jetson. Meta’s Muse Glimmer is another 30-billion-parameter model for coding and agentic workloads, while DeepSeek v4 Flash is a 284-billion-parameter mixture-of-experts model with 13 billion active parameters.

The larger shift is that frontier-style AI is no longer being framed only as a cloud service. NVIDIA, Microsoft and their partners are trying to make local agents easier to install, faster to run and safer for sensitive work. The October arrival of RTX Spark PCs will show how much of that promise can move from demos to everyday use.

Best Mobiles in India

Notifications
Settings
Clear Notifications
Notifications
Use the toggle to switch on notifications
  • Block for 8 hours
  • Block for 12 hours
  • Block for 24 hours
  • Don't block
Gender
Select your Gender
  • Male
  • Female
  • Others
Age
Select your Age Range
  • Under 18
  • 18 to 25
  • 26 to 35
  • 36 to 45
  • 45 to 55
  • 55+
X