Litterae.eu
Humanities & IT


LLMs for home computers

Until recently, interacting with Large Language Models (LLMs) was limited to online services such as ChatGPT, LeChat, Copilot, Gemini... However, thanks to the GGUF compression format, it is now possible to run relatively lightweight yet high-performance LLMs on your home computer. This opens up new possibilities for interaction and use.

Before we explore some of the most well-known models suitable for domestic use, a brief technical explanation is necessary.
The various models are typically classified by two labels: B and Q.
B indicates the billions of parameters used by the model to reason; in simpler terms, it represents the size of the LLM's "brain." An 8B model is more educated and logical than a 4B model, requiring more RAM to run on your PC.
Q, on the other hand, indicates quantization - the level of compression and consequently the operational quality of the model. A high value such as Q8 represents lower compression and thus a more accurate model in terms of consistency and reasoning; a low value like Q4 makes a model lighter and faster to handle for your PC but sacrifices some precision in reasoning and consistency.

Here are the results of our tests, conducted on each model using both terminal CLI via llama.ccp and graphical interface with Jan, LM Studio, and gpt4all.
For English, the Llama-3.1 family dominates the home scene, especially in uncensored versions.
The DarkIdol Llama 3.1 8B Q8 quantized by QuantFactory is currently the best choice for your home computer: accurate, completely uncensored (for a creative person, a fundamental aspect), able to handle very long conversations (in the order of 80 thousand words, equivalent to a novel of almost 600 pages), and ideal for customization via system message as a specific assistant. For smaller PCs, you can also opt for the version Q4 from QuantFactory, but be sure to pay attention to consistency of reasoning.
A balanced alternative to DarkIdol is the Cognitive Computations Dolphin Mistral 24B Q4 quantized by Bartowski; it proves to be very coherent and fluid, offering a broader knowledge base thanks to its 24 billion parameters, which generally manages to offset the less accuracy of greater compression.
A great lightweight option for those with limited resources is the Jan v3.5 4B Q4 quantized by the same Jan; a flexible LLM suitable for customization via system message, although consistency must be monitored carefully.

As for Italian, despite the fact that the training data of the LLMs remain mostly in English and Italian is a more complex language grammatically, there are also good domestic solutions available.
The first choice is the Llama-3 8B Ita Q8 still quantized by QuantFactory. Based on LLama-3, it can handle extremely long conversations, is optimal for customization via system messages, and has an excellent level of logic and consistency. There are rare grammar flaws, but they are really minimal.
A good alternative is the ANITA NEXT 24B Q4 quantized by Marco Polignano, which has excellent language care and good logical consistency despite the increased compression.
For future reference, keep an eye on the Minerva 7B instruct v1.0 Q8 quantized by the NLP department of Sapienza, the Mistral Ita 7B Q8 quantized by QuantFactory, and the Modello Italia 9B GGML Q8 quantized by Francesco Baldassarri. On paper, they all have good potential for domestic use, although at the moment they have actually turned out to be unusable due to various technical lacks.

So you can already interact with LLMs locally on your PC without relying on external services and thus gaining full control of your data; not only in English but also in Italian.

Puoi leggere questo articolo anche in italiano.


Site designed by litterae.eu. © 2004-2026. All rights reserved.
Info GDPR EU 2016/679: no cookies used, no personal data collected.
p.iva / vat number: 02757940206