Nemotron 3.5 Lightning model from NVIDIA available in Ollama for local agentic tasks with a context of 1M tokens
NVIDIA has made the open-source model Nemotron 3.5 Lightning (30B parameters, 3B active per token, MoE) available in Ollama for local AI agents. It supports a context of up to 1M tokens, runs on NVIDIA GPUs as well as Apple silicon and can be further trained for custom tasks.
The Nemotron 3.5 Lightning model from NVIDIA has been available in Ollama since 11 August 2026 and runs entirely on a local device. It is an open-source model with 30 billion parameters, only 3 billion of which are active per token, built on a hybrid Mixture-of-Experts architecture. According to NVIDIA, it is intended primarily for agents that run for long periods and go through repetitive steps – reading a file, calling a tool, sorting a result, retrying after an error – for which a large model is not needed.
The model supports a context of up to 1 million tokens, which, according to the manufacturer, leaves room for long tool-use histories within multi-step tasks. For faster inference, it uses speculative decoding with multi-token prediction (MTP), using DFlash or DSpark techniques, which, according to NVIDIA, is expected to deliver up to 4 times higher throughput compared with comparable open models. It runs on NVIDIA RTX PC, RTX PRO workstation, DGX Spark and DGX Station, as well as in data centers and in the cloud; for Apple silicon, Ollama offers the nemotron-3.5-lightning:30b-mlx variant. The model was developed in collaboration with Nemotron Coalition and trained on coding, tool calling, instruction following and multi-turn conversation tasks. It is an open model trained on open data that can be further trained (post-trained) for a specific task and run anywhere from edge devices to a data center.
To run it, simply download Ollama and use the commands for a specific harness, for example „ollama launch claude --model nemotron-3.5-lightning" for Claude Code, or similarly for OpenClaw, Hermes Agent or OpenCode. The same pattern also works for models running in the Ollama cloud, so an agent can send an individual step to a larger model without changing the rest of the configuration.
According to NVIDIA, Nemotron 3.5 Lightning offers 4 times higher throughput and 30 % faster task completion compared with other leading open models of similar size, along with leading accuracy in agentic, coding and reasoning tasks. According to the source, detailed benchmark results and test configurations are provided in the launch blog from NVIDIA, which was not part of the text being processed.
Why it matters
The model allows repetitive steps performed by AI agents (reading files, calling tools, checking results) to run directly on local hardware without sending data to the cloud, addressing privacy and latency concerns. The low number of active parameters (3B out of 30B) and support for a long context of 1M tokens also make it possible to deploy such a model on more commonly available GPUs than those needed for dense models of comparable size, while more demanding steps can be redirected through the same API and CLI to a larger hosted model.
Release card
Nemotron 3.5 Lightning
NVIDIA
- Context
- up to 1M tokens
- Inputs
- text (agentic tasks, tool calling, coding)
- long-running personal assistants working with email, calendars and bookings
- coding sub-agents for running tests, searching code and refactoring within existing harnesses
- security operations such as alert enrichment, incident classification and indicator correlation
- according to the source, steps that require a larger model should be passed to a separately hosted larger model through the same CLI/API
According to the source, it offers 4 times higher throughput and 30 % faster task completion compared with other leading open models of similar size, along with leading accuracy in agentic, coding and reasoning tasks.
The card summarizes information from the article and any dated corrections, with a link to the original source. It is not our assessment of the model. It does not yet have a dedicated editorial profile. Model selection and other announcements →
Two audiences, two different impacts
What this means
For individuals
Developers can download the model through Ollama and run it locally on their own GPU (RTX PC, RTX PRO workstation, DGX Spark or Station, as well as Apple silicon through the MLX variant), including integration with tools such as Claude Code, OpenCode or Hermes Agent, without data leaving the device.
For a business
According to the description from NVIDIA, companies can move high-volume steps in agentic tasks (reading files, calling tools, sorting and retrying failed steps) to a locally running model and send only the more demanding steps through the same API and CLI to a larger hosted model, reducing reliance on cloud token costs.
DevelopmentCheck the original
Event sources
clearly official source · 1 publisher, 0 independent. We count feeds from the same owner only once.