AMD just changed the local AI game.
The Ryzen AI Max+ 395 is being called the first x86 chip that can run a 200B+ parameter model on one piece of silicon.
No separate graphics card needed.
The CPU and GPU share 128GB of unified memory, giving Linux users around 110GB of usable VRAM.
That means machines like the GMKtec EVO-X2 can run models like Qwen3 235B fully, DeepSeek V3 comfortably, and Llama 3.3 70B with room to spare.
AMD also claimed the chip beat an NVIDIA RTX 5080 by more than 3x on DeepSeek R1 inference.
A lunchbox-sized PC outrunning a $1,000 discrete GPU on a real AI workload is wild.
And the economics make it even crazier.
A heavy AI user might pay $200/month for Claude Code Max, $200/month for ChatGPT Pro, $20/month for Cursor, and $20/month for Gemini.
That’s $5,280 a year.
At that point, a local AI box can pay for itself in around 9 to 10 months.
Install Ollama, pull the model, point Claude Code at localhost, and suddenly you have the same workflow with one m...
Suggested Credits
Tags, Events, and Projects