Skip to content
important New models

Nemotron 3.5 Lightning model from NVIDIA available in Ollama for local agentic tasks with a context of 1M tokens

clearly official source

NVIDIA has made the open-source model Nemotron 3.5 Lightning (30B parameters, 3B active per token, MoE) available in Ollama for local AI agents. It supports a context of up to 1M tokens, runs on NVIDIA GPUs as well as Apple silicon and can be further trained for custom tasks.

The Nemotron 3.5 Lightning model from NVIDIA has been available in Ollama since 11 August 2026 and runs entirely on a local device. It is an open-source model with 30 billion parameters, only 3 billion of which are active per token, built on a hybrid Mixture-of-Experts architecture. According to NVIDIA, it is intended primarily for agents that run for long periods and go through repetitive steps – reading a file, calling a tool, sorting a result, retrying after an error – for which a large model is not needed.

The model supports a context of up to 1 million tokens, which, according to the manufacturer, leaves room for long tool-use histories within multi-step tasks. For faster inference, it uses speculative decoding with multi-token prediction (MTP), using DFlash or DSpark techniques, which, according to NVIDIA, is expected to deliver up to 4 times higher throughput compared with comparable open models. It runs on NVIDIA RTX PC, RTX PRO workstation, DGX Spark and DGX Station, as well as in data centers and in the cloud; for Apple silicon, Ollama offers the nemotron-3.5-lightning:30b-mlx variant. The model was developed in collaboration with Nemotron Coalition and trained on coding, tool calling, instruction following and multi-turn conversation tasks. It is an open model trained on open data that can be further trained (post-trained) for a specific task and run anywhere from edge devices to a data center.

To run it, simply download Ollama and use the commands for a specific harness, for example „ollama launch claude --model nemotron-3.5-lightning" for Claude Code, or similarly for OpenClaw, Hermes Agent or OpenCode. The same pattern also works for models running in the Ollama cloud, so an agent can send an individual step to a larger model without changing the rest of the configuration.

According to NVIDIA, Nemotron 3.5 Lightning offers 4 times higher throughput and 30 % faster task completion compared with other leading open models of similar size, along with leading accuracy in agentic, coding and reasoning tasks. According to the source, detailed benchmark results and test configurations are provided in the launch blog from NVIDIA, which was not part of the text being processed.

What changed

Why it matters

The model allows repetitive steps performed by AI agents (reading files, calling tools, checking results) to run directly on local hardware without sending data to the cloud, addressing privacy and latency concerns. The low number of active parameters (3B out of 30B) and support for a long context of 1M tokens also make it possible to deploy such a model on more commonly available GPUs than those needed for dense models of comparable size, while more demanding steps can be redirected through the same API and CLI to a larger hosted model.

Release card

Nemotron 3.5 Lightning

NVIDIA

open weights
Context
up to 1M tokens
Inputs
text (agentic tasks, tool calling, coding)
According to the sources, it is suitable for
  • long-running personal assistants working with email, calendars and bookings
  • coding sub-agents for running tests, searching code and refactoring within existing harnesses
  • security operations such as alert enrichment, incident classification and indicator correlation
Documented limits
  • according to the source, steps that require a larger model should be passed to a separately hosted larger model through the same CLI/API

According to the source, it offers 4 times higher throughput and 30 % faster task completion compared with other leading open models of similar size, along with leading accuracy in agentic, coding and reasoning tasks.

The card summarizes information from the article and any dated corrections, with a link to the original source. It is not our assessment of the model. It does not yet have a dedicated editorial profile. Model selection and other announcements →

Two audiences, two different impacts

What this means

01

For individuals

Developers can download the model through Ollama and run it locally on their own GPU (RTX PC, RTX PRO workstation, DGX Spark or Station, as well as Apple silicon through the MLX variant), including integration with tools such as Claude Code, OpenCode or Hermes Agent, without data leaving the device.

What to do Try the model with the command „ollama run nemotron-3.5-lightning" or connect it to an existing agentic tool (Claude Code, OpenCode, Hermes Agent, OpenClaw) through Ollama.
More practical updates →
02

For a business

According to the description from NVIDIA, companies can move high-volume steps in agentic tasks (reading files, calling tools, sorting and retrying failed steps) to a locally running model and send only the more demanding steps through the same API and CLI to a larger hosted model, reducing reliance on cloud token costs.

Development
What to decide Consider a pilot deployment of the model for high-volume, low-risk steps in agentic tasks (e.g. security operations, coding sub-agents) and compare costs and throughput with the existing cloud solution.
More business impacts →
AI agents local LLM Nemotron 3.5 Lightning NVIDIA Ollama open-source

Check the original

Event sources

clearly official source · 1 publisher, 0 independent. We count feeds from the same owner only once.

1
Ollama blog (local models) primary source · first detected NVIDIA Nemotron 3.5 Lightning