TNG Technology Consulting GmbH

TNG repacks Nemotron 3.5 Lightning as GGUF

September 11th, 2026

Following our Qwen3.6-27B release, we also converted NVIDIA's Nemotron 3.5 Lightning 30B-A3B into GGUF (GPT-generated unified format) for local use. Via llama.cpp, it runs natively on Blackwell hardware, with the full 1M token context fitting on a single 24 GB GPU. The optional multi-token prediction raises throughput without affecting output quality.

We preserved NVIDIA's 4-bit quantization bit-exactly and carried over the remaining weights with minimal loss, so the result seems to stay very close to the NVIDIA original.

The model weights and more details are available on Hugging Face.