< Back to situation

[REVISION HISTORY]

Nvidia Nemotron AI model and software development

Updated 3 times since CLSTR started tracking revisions of this situation.

What changed

2026-08-19 10:54 UTC → 2026-08-20 11:31 UTC · added removed

Nvidia Nemotron AI model and software development

Nvidia is expanding its presence in the AI software and open-weight model market through the development and release of its Nemotron model family. Reports indicate the company is working on Nemotron 4, a next-generation series that may include a version with at least one trillion parameters. This development follows significant financial commitments, with Nvidia reportedly tripling its cloud spending on in-house model training to $28 billion through 2031. In a recent release, To optimize AI agent efficiency, Nvidia launched released Nemotron 3.5 Lightning, an open-source a 30-billion-parameter mixture-of-experts (MoE) model. This 30-billion-parameter model is designed for efficiency in AI agent systems, activating only 3 billion parameters per token to optimize speed. To support this, Nvidia introduced the NeMo Switchyard routing library By utilizing proprietary NVFP4 4-bit floating-point precision and utilized Quantization-Aware Distillation (QAD) to compress (QAD), the model, reducing its size model reduces memory requirements from 66 GB to 22 GB while maintaining accuracy. The model and achieves up to four times the throughput of full-precision versions by utilizing proprietary NVFP4 4-bit floating-point precision. versions. Accompanying this release is NeMo Switchyard, an open-source routing library designed to direct requests to the most suitable models based on latency, cost, and quality. Nvidia claims this router can achieve a “74% cost reduction” compared to using frontier models exclusively, though it may result in a “6% reduction in accuracy.” Implementation challenges have been noted, including potential latency of approximately 700 milliseconds per request and the fact that the judging model used for routing can consume up to 21.2% of total costs.

Versions

  1. 2026-08-20 11:31 UTC Nvidia Nemotron AI model and software development
  2. 2026-08-19 10:54 UTC Nvidia Nemotron AI model development
  3. 2026-08-19 10:53 UTC Nvidia Nemotron AI model development
  4. 2026-08-18 00:44 UTC Nvidia Nemotron AI model development

Only revisions since CLSTR began indexing content versions appear here. Select a version to see what changed compared to the one before it.