Skip to content
worth noting New models

NVIDIA has made Nemotron 3.5 Lightning available for agent workflows in Amazon SageMaker JumpStart

clearly official source

NVIDIA has made the open model Nemotron 3.5 Lightning (30B parameters, 3B active, MoE architecture) available in Amazon SageMaker JumpStart. According to the vendor, it offers up to 4× higher throughput and 30 % faster task completion for high-volume agent workflows.

AWS has made the Nemotron 3.5 Lightning model from NVIDIA available in Amazon SageMaker JumpStart. Deployment requires no manual configuration of serving infrastructure – the model can also be launched directly from the Hugging Face page by selecting deployment to Amazon SageMaker AI. It is an open model intended for high-volume, continuously running agents, derived through distillation from Nemotron 3 Ultra and developed in collaboration with Nemotron Coalition; because it is open and trained on open data, NVIDIA says it can be further fine-tuned on proprietary data, with ownership of the resulting weights.

The model has a hybrid Mixture-of-Experts architecture with a total of 30 billion parameters, of which only 3 billion are active per pass, allowing it to run on a single supported GPU (e.g. instances ml.g6e.12xlarge, ml.g6e.24xlarge, ml.p4d.24xlarge or ml.p5.48xlarge). It supports a context window of up to 1M tokens, handles text input and output, and uses DFlash speculative decoding to reduce latency. According to NVIDIA, the model achieves up to 4 times higher throughput and up to 30 % faster task completion on high-volume agent tasks compared with open models in the same class. NVFP4 and BF16 variants are available in SageMaker JumpStart.

According to results published by NVIDIA, the accuracy of the NVFP4 variant remains close to that of the BF16 variant across benchmarks such as MMLU Pro (81.94 vs. 81.62), GPQA Diamond (75.44 vs. 75.57), SWE-bench Verified (51.56 vs. 52.80), PinchBench (85.37 vs. 83.43), IFBench (71.88 vs. 72.88) or AA-LCR (52.00 vs. 49.19). NVIDIA states that the evaluation used a unified testing framework and that the results may differ from figures reported by other vendors. Deploying the model creates an endpoint in SageMaker AI, with running costs charged according to Amazon SageMaker AI pricing; the source does not provide specific rates.

Details of the deployment procedure and additional parameters can be found in the source article.

What changed

Why it matters

In agent systems, a large proportion of model calls involve simple, repetitive steps (classification, data extraction, rule checks), for which deploying a frontier model is unnecessarily costly and slow. A smaller specialized model available directly through SageMaker JumpStart allows companies to separate these high-volume steps from planning and orchestration, where a stronger model is still needed, and thus, according to NVIDIA, reduce both latency and costs without having to build their own serving infrastructure.

Release card

Nemotron 3.5 Lightning

NVIDIA

open weights
Specifications
30B total / 3B active (MoE)
Context
Up to 1M tokens
Inputs
Text input, text output
Documented measurements
  • MMLU Pro BF16 81,94 / NVFP4 81,62 Comparison of model accuracy in the BF16 and NVFP4 variants on the MMLU Pro benchmark.
  • GPQA Diamond BF16 75,44 / NVFP4 75,57 Comparison of model accuracy in the BF16 and NVFP4 variants on the GPQA Diamond benchmark.
  • SWE-bench Verified BF16 51,56 / NVFP4 52,80 Comparison of model accuracy in the BF16 and NVFP4 variants on the SWE-bench Verified benchmark.
  • PinchBench BF16 85,37 / NVFP4 83,43 Comparison of model accuracy in the BF16 and NVFP4 variants on the PinchBench benchmark.
  • IFBench BF16 71,88 / NVFP4 72,88 Comparison of model accuracy in the BF16 and NVFP4 variants on the IFBench benchmark.
According to the sources, it is suitable for
  • High-volume, repetitive steps in agent workflows (e.g. classifying alerts, extracting fields from a form, checking a record against a policy)
  • Long, multi-turn agent sessions thanks to a context window of up to 1M tokens and lower latency thanks to DFlash speculative decoding
  • Deployment on a single supported GPU without configuring serving infrastructure through Amazon SageMaker JumpStart
Documented limits
  • Planning multi-step workflows or orchestrating sub-agents, which the source says require frontier-level reasoning

According to NVIDIA, this is the fastest open model in its class for continuously running agents; the article contrasts a “system of models” approach (assigning specialized, high-volume steps to the Lightning model) with routing every step of an agent workflow through a single large frontier model.

The card summarizes information from the article and any dated corrections, with a link to the original source. It is not our assessment of the model. It does not yet have a dedicated editorial profile. Model selection and other announcements →

Two audiences, two different impacts

What this means

01

For individuals

Developers building agent applications have another model option for routine, high-volume workflow steps that can be deployed without manually configuring serving infrastructure.

What to do Review the model specifications (MoE architecture, 1M-token context window, NVFP4/BF16 variants) and consider trying deployment through SageMaker JumpStart or Hugging Face for a specific agent task with a high volume of calls.
More practical updates →
02

For a business

Companies running agent systems on AWS gain the option to separate inexpensive, high-volume steps (classification, extraction, rule checks) from costly frontier model calls, which NVIDIA says reduces latency and the cost of running agents; the model is open and can be fine-tuned on proprietary data through NVIDIA NeMo.

Development
What to decide Evaluate whether to deploy Nemotron 3.5 Lightning through SageMaker JumpStart instead of a more expensive frontier model for high-volume, repetitive steps in agent workflows (classification, data extraction, checks against rules), and…
More business impacts →
AI agents AWS SageMaker JumpStart mixture-of-experts model distillation NVIDIA Nemotron open-source

Check the original

Event sources

clearly official source · 1 publisher, 0 independent. We count feeds from the same owner only once.

1
AWS Machine Learning Blog primary source · first detected NVIDIA Nemotron 3.5 Lightning now available in Amazon SageMaker JumpStart