How to run an LLM locally: what fits on your hardware
The main thing between you and a working local setup is picking a model that fits your memory. Get that wrong and you will conclude local models are useless.
Read MoreIndependent AI news, with every claim linked to its primary source
Independent AI news, with every claim linked to its primary source
Evergreen explainers, guides and comparisons. The reference material that stays true after the news cycle moves on.
The main thing between you and a working local setup is picking a model that fits your memory. Get that wrong and you will conclude local models are useless.
Read MoreFine-tuning teaches behaviour. It does not teach facts reliably, and confusing those two is where the money goes.
Read MoreNothing is composed. The image emerges everywhere at once, which is exactly why counting and text placement go wrong.
Read MoreMost teams reach for a dedicated vector store where a table in the database they already run would do. Here is where the threshold sits.
Read MoreA citation has a recognisable shape, and producing that shape is easy. A three-step workflow that survives deadline pressure.
Read MoreThe loop takes twenty lines. Error compounding, absent memory, no stopping condition and excess permissions are where the work actually is.
Read MoreOutput tokens cost up to five times input. A support assistant at 10,000 conversations a day runs $2,700 or $9,000 a month depending only on model choice.
Read MoreClassifiers measure how ordinary writing looks, which is not the same as who wrote it. Watermarking is better and still proves processing, not authorship.
Read MoreA transformer does fixed computation per token, so extra tokens are the only way to spend more effort. That constraint explains the gains and the bill.
Read MorePublic human text runs out between 2026 and 2032. Generated data fixes that only where a checker exists, which is a much narrower set of tasks than the pitch suggests.
Read MoreOpenAI found a 1.3B tuned model was preferred over 175B GPT-3. The three stages, the alignment tax, and whose preferences get encoded.
Read MoreA trillion-parameter model can cost the same to run as a small one, because only a fraction of it fires per token. How sparsity works and what it costs.
Read More