TNG Technology Consulting GmbH

TNG releases Qwen3.6-27B-NVFP4-GGUF

July 28th, 2026

We repackaged NVIDIA’s Qwen3.6-27B-NVFP4 into GGUF format. The primary advantage is inference speed on local setups.

It comes with a performance increase of up to 50% faster decode at 46.5 tokens/second on an RTX 5090 laptop, while coming close to the original precision. Running natively with 4-bit quantization on NVIDIA Blackwell, it fits easily into 24 GB of video RAM. With that, a large context of 160k tokens multi token prediction can be preserved.

The weights and more information are available on Hugging Face.