Nvidia releases Nemotron 3.5 Lightning, 30B MoE model with 4x faster output
Nvidia launches efficient open-source model and NeMo Switchyard router for multi-agent AI systems.
Conversation activity · last 24 hours peak 31/hr
Summary, timeline and people extracted by Claude from 44 items across 7 sources · 5h ago. Quotes are verbatim.
Nvidia unveiled Nemotron 3.5 Lightning, a 30-billion-parameter mixture-of-experts open model claiming up to 4x faster output speeds and 30% faster agentic task completion versus competing models in its class, alongside NeMo Switchyard, an open-source routing library for intelligent model selection in agent systems. The releases target enterprises deploying multi-model AI systems that require task-specific routing and easy customization on proprietary data.
- Nemotron 3.5 Lightning is an open, customizable 30B-parameter MoE model targeting high-volume agentic AI tasks, claiming 4x faster output and 30% faster task completion versus competing models.
- NeMo Switchyard is a new open-source routing library that intelligently directs requests to the most suitable model in a multi-model system without requiring application rewrites.
- The model emphasizes ease of customization and low-cost domain-specific fine-tuning, with early partners achieving trained agents for under $100 in hours.
- Nvidia released training datasets, recipes, and post-training frameworks alongside the model to enable enterprise customization on proprietary data while maintaining deployment flexibility across local, edge, and cloud infrastructure.
How it unfolded
-
Event Nvidia announces Nemotron 3.5 Lightning and NeMo Switchyard
Nvidia released Nemotron 3.5 Lightning, a 30B-parameter MoE model designed for agentic AI workloads, alongside NeMo Switchyard, an open-source model routing library for intelligent request direction across multiple models.
“What we're hearing is that Lightning is remarkably easy to customize”
Kari Briski, VP of Generative AI at Nvidia · Techmeme ↗ -
Report Model performance claims and customization examples
Nvidia claims the model delivers up to 4x faster output speed and 30% faster agentic task completion compared with other models in its weight class. Early partner examples show rapid customization: CodeRabbit trained a router agent for $85 in around two hours using a single H100 card.
-
Report Open-source model architecture and deployment flexibility
Nemotron 3.5 Lightning is an open and customizable model that can run on local systems including RTX PCs, DGX workstations, and edge devices, or scale across data centers and cloud environments. Nvidia is releasing datasets, post-training datasets, and training recipes used to develop the model.
-
Report Early adoption by enterprise partners
Multiple organizations have begun customizing Nemotron 3.5 Lightning for domain-specific tasks, including CrowdStrike for cybersecurity, Harvey with Trajectory for legal services, CodeRabbit for code review, Lila Sciences for physical and life sciences reasoning, and Fastino Labs for software development, finance, and healthcare.
What people are saying verbatim
“What we're hearing is that Lightning is remarkably easy to customize”
Kari Briski, VP of Generative AI at Nvidia · SiliconANGLE ↗ · Aug 10
“CodeRabbit Inc., according to Briski, used Nvidia's standard auto model recipe, trained for one epoch, and produced a router agent for $85 in around two hours.”
SiliconANGLE (reporting Kari Briski), News coverage · SiliconANGLE ↗ · Aug 10
“This is turning post-training and fine-tuning experience into loading the dishwasher and hitting the button.”
SiliconANGLE (editorial summary of Nvidia's positioning), News coverage · SiliconANGLE ↗ · Aug 10
“Nemotron 3.5 Lightning, a 30-billion-parameter mixture-of-experts model, helps create smarter and more efficient agentic applications.”
Nvidia Corp., Company blog announcement · Nvidia blog ↗ · Aug 10
“Nemotron 3.5 Lightning is the successor to @nvidia Nemotron 3 Nano 30B”
@artificialanlys, Social media analyst · X (Twitter) ↗ · Aug 10
“Enterprises can use it to build a router based on their specific needs. When deployed, NeMo Switchyard can intelligently direct each request to the most capable and suitable model for the job without requiring developers to rewrite their applications.”
Nvidia Corp., Company blog announcement · Nvidia blog ↗ · Aug 10
Voices from the web unedited
-
Nvidia unveils first open-source AI model since CEO Jensen Huang entered the chat
-
NVIDIA has just released the first Nemotron 3.5 model: Nemotron 3.5 Lightning, a highly efficient small open weights model with performance similar to gpt-oss-120b at around a quarter of the total parameters Nemotron 3.5 Lightning is the successor to @nvidia Nemotron 3 Nano 30B [image]
-
Ubuntu users can now run NVIDIA’s Nemotron 3.5 Lightning model through Canonical’s newly released inference Snap. https:// linuxiac.com/nvidia-nemotron-3 -5-lightning-lands-on-ubuntu-via-snap/ # linux # ubuntu # nvidia # ai # opensource
-
One day after Meta released Muse Glimmer, NVIDIA launched Nemotron 3.5 Lightning and NeMo Switchyard, showing a different approach to local agentic AI. Really cool. Lightning is a 30B MoE with only 3B active parameters, distilled from Nemotron 3 Ultra and designed for the [image]
-
NVIDIA Nemotron 3.5 Lightning is now live on OpenRouter. A 30B hybrid MoE with 3B active params, distilled from Nemotron 3 Ultra. Built for high-volume, specialized AI agent workloads, delivering up to 4× higher throughput and up to 30% faster task completion compared to similar
-
In collaboration with @nvidia we're releasing two new open weight models: Fastino-Nemotron-3.5-Lightning- Finance and Fastino-Nemotron-3.5-Lightning- Healthcare. Working closely with the Nemotron team, we developed both models on Nemotron 3.5 Lightning using the Fastino [image]
-
Nvidia Nemotron 3.5 Lightning delivers leading accuracy on agentic coding tasks and up to 4x the speed of comparable open models. With open weights, datasets, and recipes, Nemotron is open and easy to customize. Welcome to Pi, Nemotron 3.5 Lightning! [image]
-
Introducing NVIDIA Nemotron 3.5 Lightning⚡ An open 30B MoE model with 3B active parameters, built for always-on agents to complete high-volume, specialized tasks faster. It delivers up to 4x the output speed of similar-sized models. [image]