deployment
GGUF
A file format for storing quantized machine learning models, designed for efficient local inference using the llama.cpp library. GGUF replaced the older GGML format and supports metadata, multiple tensor types, and various quantization levels.
In practice
Downloading a 4-bit GGUF version of Llama 3 allows running it on a laptop with 16GB of RAM using llama.cpp.