Litterae.eu
Humanities & IT


When a small language model fits better than an LLM

Let's speak plainly.
The true engine behind the recent surge in artificial intelligence, and specifically the Large Language Models that now also paint images, weave videos, and compose music, was never intended for such grandeur. It was the sheer power of graphics processing units, born from the world of video games, that made this revolution possible.
Yet, this has set in motion a relentless cycle. The hunger of the models and the might of the hardware chase each other ever higher, climbing towards stratospheric heights with results that are, in many ways, staggering.

Yet, for most of us — ordinary users and businesses alike — these giants are overkill.
We do not need the colossal, all-encompassing models that stand behind the great commercial services, nor do we require the impressive mid-weight engines often discussed. Most of us simply need a tool to perform a single, repetitive task.
To rely on a local model of at least eight billion parameters, or to pay for a subscription to a massive online service for such a simple need, is like living in the heart of a city and driving an SUV or a taxi to the shop next door, when a bicycle or a simple trolley would suffice.
Moreover, businesses generally need executors, not dreamers. Therefore, as we have seen before, spending thousands of euros on powerful graphics cards or expensive subscriptions is often unnecessary. A language model with fewer than a billion parameters, when paired with a well-structured Retrieval-Augmented Generation system, is often enough to do the job.

A recent article on dev.to showed how these minimal models can run even on older, less powerful computers.
While we need not go to such extremes, we can confirm that a model like SmolLM from Hugging Face, with only 135 million parameters, works exceptionally well. It requires no dedicated graphics card and less than a gigabyte of RAM, generating forty to forty-five tokens per second.
It is certainly not a vessel upon which we can build a complex AI persona, nor one that could interpret the depths of The Brothers Karamazov​. But, when linked to a focused RAG on a single specialist topic, it becomes a truly rapid and precise assistant.
Sometimes, the simplest path is the one that leads us home.


Puoi leggere questo articolo anche in italiano.
English version of the article written by Kore.
Image generated with Craiyon.


Site designed by litterae.eu. © 2004-2026. All rights reserved.
Info GDPR EU 2016/679: no cookies used, no personal data collected.
p.iva / vat number: 02757940206