Local AI on 8GB of VRAM: this is how I do
My local AI setup AKA My journey of making a frankly unreasonable number of experiments in making 8GB of VRAM behave like more.
Welcoming our new AI overlords.
Practical writing on AI, software development, and the tech that survives contact with reality — while keeping a cynical eye on our new robot overlords.
Local AI is real and I run it daily — but when a frontier model offered to optimise my llama-server script, invented its own benchmark, and found 15%, I was reminded why the big models are not redundant. Then it reminded me of something else.
My local AI setup AKA My journey of making a frankly unreasonable number of experiments in making 8GB of VRAM behave like more.
Local AI is not one magic download. It is a stack of models, inference engines, clients, prompts, hardware limits, and occasionally MCP plumbing.
When companies turned AI token spend into a performance metric, engineers started burning tokens to hit quotas
Dropbox's Nova, LinkedIn's MCP tooling, and GitHub's Copilot app show coding agents becoming a fleet that needs orchestration, sandboxing, and context plumbing.
New York's data center moratorium, Utah's forced downsizing, and capital fleeing to India and orbit show permitting, not GPUs, now gates AI scale.
Sparse models, MoE, quantisation, and better local runtimes have changed what is possible on modest hardware. Local AI is no longer only for people running dual RTX 3090 rigs — but it is still very much for developers who can supervise the machine when it starts confidently sawing through the floorboards.
AWS swapped hierarchical fat-tree fabrics for quasi-random flat meshes with passive optical ShuffleBoxes, cutting routers 69% — and the logic will spread.
Google Cloud deleted a $124B fund's entire infrastructure, replicas included. The only thing that saved it was a backup Google didn't control.
Bot traffic has quietly surpassed human traffic on the web, yet most infrastructure — rate limiting heuristics, analytics pipelines, CDN cache-fill assumptions
Uber's $1,500-per-tool monthly cap on Claude Code isn't about frugality — it's the first public admission that no one can measure agentic coding's return.
A sabotaged jqwik release and a critical Starlette flaw expose one blind spot: coding agents run third-party code under a threat model nobody designed for.