Local AI is real and I run it daily — but when a frontier model offered to optimise my llama-server script, invented its own benchmark, and found 15%, I was reminded why the big models are not redundant. Then it reminded me of something else.
Sparse models, MoE, quantisation, and better local runtimes have changed what is possible on modest hardware. Local AI is no longer only for people running dual RTX 3090 rigs — but it is still very much for developers who can supervise the machine when it starts confidently sawing through the floorboards.