Sparse models, MoE, quantisation, and better local runtimes have changed what is possible on modest hardware. Local AI is no longer only for people running dual RTX 3090 rigs — but it is still very much for developers who can supervise the machine when it starts confidently sawing through the floorboards.
AWS swapped hierarchical fat-tree fabrics for quasi-random flat meshes with passive optical ShuffleBoxes, cutting routers 69% — and the logic will spread.