Why Local AI Matters
Running models on your own hardware is not a hobby project. It is sovereignty. Your data stays. Your inference does not depend on an API.
Every conversation you send to a cloud API is a conversation you do not own. The prompt, the response, the context window — all of it lives on someone else’s infrastructure, governed by someone else’s terms, priced at someone else’s discretion. Today the API costs $0.003 per thousand tokens. Tomorrow it costs $0.03. Next quarter the terms of service change and your fine-tuned workflow violates a clause that did not exist when you built it. You have no recourse. You rented the capability. The landlord sets the rules.
I run models locally. Not as an experiment. As a position.
The Hardware Equation
Apple silicon changed what “local AI” means. The M4 Max with 128GB of unified memory — shared between CPU and GPU with no bus bottleneck — runs a 32B parameter model at inference speeds that are genuinely usable. Not fast enough for a consumer chat product serving millions. Fast enough for a single practitioner doing real work. The difference matters.
Unified memory is the key. A discrete GPU setup with 24GB of VRAM forces you into quantised models or model sharding — splitting the network across devices, adding latency at every boundary. The M4 Max gives you 128GB of memory that the GPU cores can address directly. A 32B model in 4-bit quantisation fits comfortably. A 14B model runs in full precision. The ceiling is high enough for serious inference and climbing with every chip generation.
MLX — Apple’s machine learning framework, open-sourced, designed specifically for Apple silicon — makes the software native. No CUDA dependency. No Docker container wrapping a Linux environment on a Mac. Native Metal acceleration, lazy evaluation, and a NumPy-compatible API that researchers actually want to use. The ecosystem is small but growing fast. Quantised models in MLX format load in seconds and run without the overhead of translation layers.
I run Qwen 2.5 32B for reasoning-heavy tasks. Llama for general-purpose inference. Both quantised to 4-bit, both running through MLX, both producing output I can use in production pipelines without sending a single token off my machine.
The Sovereignty Argument
This is not a technical preference. This is a structural position about who controls the intelligence layer of your work.
When I run a model locally, my data does not leave. The client briefs, the internal documents, the code repositories, the personal writing — none of it touches an external server. Privacy is not a policy I trust someone else to enforce. It is a physical fact of the architecture. The bits stay on my disk. The inference happens on my chip. The results exist in my memory. No telemetry. No audit trail on someone else’s dashboard. No “we may use your data to improve our models” buried in paragraph fourteen of the terms.
The sovereignty argument is the same argument I make about owned channels over rented reach. Your website instead of your Medium page. Your newsletter instead of your LinkedIn posts. Your negatives instead of your Instagram archive. The pattern is consistent: anything built on infrastructure you do not control can be taken from you by a decision you were not consulted on.
Cloud APIs are rented reach for intelligence. The model you call today might be deprecated tomorrow. The pricing you budgeted against might double. The moderation layer might start refusing prompts that were fine last month. Every dependency on an external API is a dependency on someone else’s roadmap, and roadmaps change without your permission.
The Tradeoffs
I am not pretending local inference matches frontier cloud models. It does not. Claude, GPT-4, Gemini — the largest cloud models operate at a scale that no laptop will replicate. The gap is real. A 32B local model does not reason at the depth of a 400B+ cloud model. Context windows are shorter. Tool use is less reliable. Multi-step planning is weaker.
But the question is not “which model is better.” The question is “which model do I control.” For tasks that require my data to stay local — document analysis, code review on proprietary repositories, personal knowledge management, draft generation on sensitive material — the local model is not a compromise. It is the only option that respects the constraint.
And the gap is closing. Every quarter, the open-weight models improve. Qwen 2.5 at 32B outperforms what GPT-4 could do two years ago on most benchmarks. The frontier moves, but the local floor rises faster than most people track.
Control the Stack
The principle is simple and it applies everywhere I work. If you depend on a platform, you are subject to the platform. If you control the stack, you control the work.
I host my own site. I run my own models. I keep my own negatives. I maintain my own instruments. The cost is maintenance — updates, configuration, the overhead of being your own infrastructure team. The benefit is that no one can change the terms of my practice with an email notification.
Local AI is not a hobby. It is the same conviction that keeps the guitar in the case at home instead of rented from a studio, the same conviction that keeps the domain registered in my name instead of published on someone else’s platform. Own the tools. Own the work.
More from this domain
12 Aug 2026
MachinesThe Nine-Stage Mind
Nine stages around a 27B local model. On my own tasks it beats a trillion-parameter model alone. Architecture beats parameter count.
9 Jul 2026
MachinesThe Harness Is the Product
Every team has Claude Sonnet 4.5. A 20-step pipeline at 95% per step finishes 36% of the time. The harness is where the product lives.
11 Jun 2026
MachinesM3: Frontier Coding, Not Frontier Freedom
428B params, 1M context, 80.5% on SWE-Bench Verified by its own card. M3 wins long-context coding — then drops MIT for a license with strings.