this post was submitted on 21 Nov 2023
1 points (100.0% liked)
LocalLLaMA
3 readers
1 users here now
Community to discuss about Llama, the family of large language models created by Meta AI.
founded 1 year ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
I've tested pretty much all of the available quantization methods and I prefer exllamav2 for everything I run on GPU, it's fast and gives high quality results. If anyone wants to experiment with some different calibration parquets, I've taken a portion of the PIPPA data and converted it into various prompt formats, along with a portion of the synthia instruction/response pairs that I've also converted into different prompt formats. I've only tested them on OpenHermes, but they did make coherent models that all produce different generation output from the same prompt.
https://desync.xyz/calsets.html