NVIDIA Launches Nemotron 3.5 Lightning and Advances Nemotron 4

英伟达10亿美元投资诺基亚,强劲涨势能否延续?
Published on: Aug 12, 2026
Author: Amy Liu

NVIDIA (NVDA) is accelerating its expansion from an AI chip supplier to an open-source model ecosystem builder with the release of Nemotron 3.5 Lightning and the advancement of the trillion-parameter Nemotron 4 series. The core logic of this strategy is to drive compute demand through model proliferation. Although facing competition from customers and portfolio companies, and with current model performance still lagging behind the global cutting edge, the success or failure of Nemotron 4 will determine whether NVIDIA can build a complete AI technology stack covering hardware, software, and models, thereby sustaining its growth momentum in the GPU market over the long term.

Nemotron 3.5 Lightning is a mixture-of-experts model with 30 billion parameters, primarily designed to help developers build more intelligent and efficient AI agent applications. NVIDIA calls it the most efficient model in its class for long-running AI agent workloads.

Unlike traditional chatbots that primarily revolve around single-turn question-and-answer interactions, AI agents typically need to perform multi-step reasoning, call different tools, and continuously complete complex tasks. As a result, model operating efficiency, inference cost, and the ability to orchestrate between different models become more critical. With Nemotron 3.5 Lightning, NVIDIA aims to further penetrate this rapidly growing application scenario.

At the same time, the company has also released the open-source library NeMo Switchyard, which provides intelligent model routing capabilities for mainstream AI agent tools. Enterprises can build their own dedicated routing systems based on their specific needs. After deployment, NeMo Switchyard can automatically route requests to AI models whose capabilities and features better match the given task, without requiring developers to rewrite existing applications. NVIDIA states that the combination of Nemotron 3.5 Lightning and NeMo Switchyard gives enterprises more flexible control over how AI is deployed, where it runs, and at what efficiency, covering different computing environments including PCs, workstations, data centers, and the cloud.

Nemotron 4 in Development, with the Largest Model Potentially Reaching at Least 1 Trillion Parameters

Beyond Nemotron 3.5 Lightning, NVIDIA is also advancing the more aggressive next-generation Nemotron 4 series in terms of scale and performance. According to The Information, citing multiple people involved in the Nemotron project, NVIDIA hopes that the largest model in this series will be able to compete in performance with the world’s best open-source AI models. Multiple employees involved in the project expect that the largest Nemotron 4 model will have at least 1 trillion parameters, roughly twice the parameter scale of Nemotron 3 Ultra, which NVIDIA launched in June this year. The final specifications and release date for Nemotron 4 have not yet been determined. The report states that final training has not yet begun, and this process itself could take several months, with some employees estimating that the model could be ready as early as late this fall.

In March of this year, NVIDIA established the Nemotron Coalition, collaborating with multiple leading global AI labs to develop open-source frontier models. The company stated at the time that the first model co-created by the coalition would serve as the foundation for the future Nemotron 4 open-source model series. Kari Briski, NVIDIA’s Vice President of Generative AI, said that the company is investing in Nemotron because it believes that every enterprise and every country needs access to cutting-edge open-source models to enhance security, accelerate innovation, and provide a foundation that can be used consistently across different technology generations.

AI Financial Service Semiconductors Technology