Local and On-Device Inference: Run Open Models Yourself
Local inference runs open-weight models on hardware you own at zero token cost. These guides cover GLM-5.2, Gemma 4, Apple Core AI, and the RAM you need. Below are all 5 guides on Local & On-Device Inference published on this site, newest first, each one a hands-on walkthrough with configuration you can copy.
All Local & On-Device Inference guides (5)
Moonshot’s 2.8T open-weight model runs for agentic coding two ways: routed into Claude Code via one env block, or the native Kimi Code CLI. Pricing, benchmarks, local-run reality, and an honest hybrid verdict.
GLM-5.2 is 2026’s top open-weight coding model. Run it with llama.cpp and Unsloth quants - the hardware you need, the right quant for coding, and when a cloud API still makes more sense.
Apple Core AI runs open-weight models like Qwen and Mistral on Apple Silicon with zero token cost. Convert PyTorch with coreai-torch, load via the Swift API, and quantize for mobile.
Tested all 4 Gemma 4 model sizes locally. Includes RAM requirements, Ollama setup, comparison with Llama 4 and Mistral, and a practical guide to picking the right variant.
Set up your own personal AI assistant that works on WhatsApp, Telegram, Discord, and more. Self-hosted with persistent memory and proactive notifications.
Explore other topics
Every guide on this site is grouped into one of these hubs.
Looking for everything at once?
The full archive lists every article in publication order, across all seven topics.