Artificial IntelligenceNews

NVIDIA Takes AI Agents Local With Nemotron and New Developer Tools

NVIDIA is stepping up its push to bring advanced artificial intelligence workloads onto local hardware, highlighting a growing ecosystem of open models, AI agents and developer tools that allow users to run AI closer to where data is created. The company says local AI is gaining momentum as developers and businesses look for greater control over data, lower inference costs and the ability to use AI without depending entirely on cloud infrastructure. NVIDIA’s latest showcase includes its Nemotron family of open models, new agentic applications and systems such as the DGX Spark designed to run sophisticated AI workloads locally.

The move also comes as AI agents become more capable of performing multi-step tasks, increasing the amount of sensitive information they may need to process. Running those workloads locally can allow organisations to keep data and computation within their own environments. At the centre of NVIDIA’s local AI strategy is the Nemotron family of open models.

NVIDIA says its open models are designed to give developers access to model weights, training datasets and techniques, allowing them to customise and deploy models rather than relying exclusively on proprietary cloud AI services. The company provides Nemotron models through platforms including Hugging Face and supports production deployment through NVIDIA’s AI software stack.

The models are also designed for agentic applications, where AI systems can reason, use tools and perform tasks across multiple steps. NVIDIA’s broader strategy is to provide the models alongside the hardware and software required to run them, creating an ecosystem that spans local development through to enterprise deployment.

One of the most significant examples highlighted in NVIDIA’s local AI push is Perplexity’s Portable Computer, which brings the company’s agentic AI platform to local hardware. The system runs the agent harness, orchestrator and models locally, initially on NVIDIA DGX Spark and compatible Linux systems equipped with NVIDIA RTX GPUs. Instead of sending every request to the cloud, tasks begin locally, with cloud models used when additional computing power or capabilities are required.

Perplexity’s approach is designed to give users greater control over where their data is processed. The system can work with files, execute code in a sandbox and interact with services such as Gmail and Slack while keeping the core agent workload on the local machine. When a task requires a more powerful cloud model, the system can request permission before sending that particular step to an external service.

Support for NVIDIA’s Nemotron 3.5 Lightning is also planned, expanding the range of models that can be used with the local agent platform. NVIDIA’s DGX Spark is central to the company’s effort to make high-performance AI computing available in a compact desktop form factor. The system is aimed at developers and AI professionals who need to experiment with and run larger models locally without building a traditional data-centre infrastructure.

The emergence of local agent platforms such as Perplexity’s Portable Computer demonstrates a potential shift in how these systems can be deployed. Rather than treating AI agents purely as cloud services, developers can increasingly run the orchestration, models and execution environment directly on hardware they control.

NVIDIA’s latest push reflects a broader change in the AI market. Local AI is increasingly moving beyond simple chatbots and image-generation applications towards agents capable of executing workflows. That creates different requirements for hardware. AI agents may need to maintain context, interact with applications, process files, run code and perform several inference steps before completing a task.

Keeping these operations local can offer benefits around privacy, latency and operating costs, while also allowing organisations to retain greater control over sensitive information. Cloud AI remains important for workloads that require larger models, fresh information or capabilities beyond the local system. The emerging model is therefore increasingly hybrid: run as much as possible locally and use cloud resources selectively when they provide a meaningful advantage.

The local AI push also gives NVIDIA an opportunity to extend its influence beyond the GPU itself. By combining open models such as Nemotron, developer frameworks, inference software and local AI systems, the company is positioning its hardware as part of an end-to-end platform for developing and deploying AI agents.

NVIDIA’s strategy comes as competition intensifies around open-weight models, with developers increasingly looking for alternatives to closed AI platforms. The company’s Nemotron programme is part of that wider shift, with NVIDIA positioning its models as customisable building blocks for developers and enterprises.

For businesses, the appeal is increasingly practical: sensitive workloads can remain on-premises or on controlled devices, while cloud AI can be brought in selectively when required. The result could be a more distributed AI landscape in which the question is no longer simply which cloud model to use, but which AI workloads should run locally, which should run in the cloud, and how the two environments should work together.

Show More

Chris Fernando

Chris N. Fernando is an experienced media professional with over two decades of journalistic experience. He is the Editor of Arabian Reseller magazine, the authoritative guide to the regional IT industry. Follow him on Twitter (@chris508) and Instagram (@chris2508).

Related Articles

Back to top button