I’ve been playing around with local (self-hosted) AI systems, and wanted to share my experiences. (This could get long.) I’m focusing on the technology, not the business practices or social impact of AI, but I will be drawing some conclusions.
Objective
My objective was to get some experience working with reasonably powerful LLMs and agent frameworks. I wasn’t quite sure what I meant by “reasonably powerful”, so I decided to acquire a hardware system limited by budget, run the most capable models that would fit on it, and compare the results with the best available LLMs. I decided that I was more concerned with quality than throughput; I wasn’t interested in snappy, real-time conversations. And I wanted to focus on domains where I had a reasonable level of expertise, so that I would be able to assess the quality and completeness of responses and detect hallucinations. Finally, I hoped to find a portable system rather than a desktop box, so that I could work anywhere. My budget was $3,000.
My first thought was to get a gaming laptop with a decent GPU, but couldn’t find anything with sufficient VRAM. Most sources recommended MacBooks, because of the unified memory, but the Apple premium plus DRAM shortage meant that the configurations I wanted were out of budget.
And then I read about the ASUS ROG Flow Z13. It’s what’s called a “two-in-one” – a tablet with a snap-on keyboard – and uses the AMD RYZEN AI MAX+ 395, a combined CPU and GPU with shared access to up to 128GB of RAM. Under Linux, the memory is fully shared, but under Windows you have to partition the RAM into CPU and GPU regions. But still… The 128GB models were well over $3,000, but Best Buy had a sale on the 64GB configuration for $2,090. So I grabbed one.
Setup
Settling on a configuration was… tricky. The biggest problem was that the architecture of AI systems is evolving pretty rapidly, and the changes require different mixes of resources. Specifically, the shift to MoE (“mixture of experts”) designs changes the optimum balance of CPU and GPU memory, and under Windows that involves reconfiguring and rebooting. (Yes, it’s easier on a Mac, and no, I wasn’t going to switch to Linux.) And the various frameworks for downloading, managing and running different LLMs – LLAMA.cpp, ollama, LMStudio – all have different strengths and weaknesses.
Related to all of this was the choice of agent framework. LMStudio has a built-in framework called Bionic; on the other hand, Deepseek Harness offers unprecedented (and possibly excessive?) reconfigurability.
Eventually, I decided to start simple. I configured the Flow Z13 to have 32GB of RAM and 32GB of GPU, installed Bionic, and downloaded Qwen 3.8 27B GGUF. This model is designed to work in 16GB, but expanded to fill the available memory. To get some visibility into how things were running, I used Dave Plummer’s excellent new Task Manager TMOG – highly recommended.
Experience
OK, it’s running. Now what? I decided to do some parallel testing, so I provided the following prompt to Qwen (running locally under Bionic) and Microsoft Copilot (in the cloud):
Recursive sensory patterns seem to underpin everything from a simple organism’s ability to distinguish its own body from a food source to consciousness and introspection.Various writers, including Daniel Dennett, have written on this theme. Please summarize the evolutionary case for recursion underpinning consciousness.
Copilot responded almost immediately. Qwen took almost 5 minutes to respond. Copilot focused almost entirely on the positive arguments for the position I’d advanced, and ended by asking me if I’d like this turned into a short essay with a reading list. Qwen gave a much deeper response, including the positions of a variety of writers both pro and con, and including a very sharp assessment of “Caveats and concerns”. While there wasn’t much point to further expansion of the Copilot interaction, the Qwen “discussion” went on for five more prompts and lengthy responses. The quality of the analysis was very high – much more than I’ve seen from Copilot, ChatGPT, or Gemini – and Qwen was far quicker to challenge the weak points in my positions.
Neither system was particularly sycophantic, although Qwen did concede:
You’ve largely won this, and I’ll say so plainly rather than keep hedging.
But that didn’t stop it:
So the honest summary: you’re right that the easy half is fully solvable — but that’s not a refutation of the hard half, it’s the definition of it.
The one place I’d push you to look next: the illusionist reply (Frankish, Churchland) takes your point all the way and says the “leftover” isn’t a gap in our knowledge but a mistake — there’s no phenomenal character to explain, only a folk-psychological feeling that there is one. That’s the version of your view that actually closes the gap rather than relocating it.
Another little test
I asked both Qwen and Copilot:
Describe the main events surrounding the signing of the North Atlantic Treaty. Which prominent individuals attended the various events, and who give speeches?
Both systems produced approximately the same level of detail in their responses, but Qwen added this:
A couple of honest caveats: I’m confident about the date, place, the 12 founding nations, and the central figures (Truman, Acheson, Dulles, Bevin, Schuman). I’m less certain about the exact roster of individual national representatives physically present at the April 4 signing (for example, the specific signatories for Canada, Iceland, Denmark, and the others) and whether any European signatory delivered a formal speech at the ceremony itself — Acheson’s was the principal address. I don’t have live web access in this session to verify those finer details against a source.
I replied, “Can you pull up a primary source”, and Bionic offered to run Curl to retrieve the source, with either human or automated review of each shell command. I approved it, and got:
Network access works (that URL just 404’d). Let me locate the authoritative primary source — the U.S. Foreign Relations of the United States (FRUS) 1949 volume, which contains Acheson’s April 4 statement and the signing record.
After repeated attempts, I almost gave up on this, because most of the web resources (at the State Department) 404’d. But the agent framework kept trying, and eventually the NATO website yielded the full text.
What to make of this
The main takeaway for me from this small experiment is that local LLMs are going to be “good enough” for a large number of common applications. Within a very few years the capabilities of my experimental setup are going to be standard platform features in PCs, phones, and cars. The main role of supporting cloud-based services will be to handle real-time data lookups, search, long-running agent tasks, and inter-agent exchanges. Obviously, the cloud providers will try to force us onto their paying (subscriptions or ads) services, but regulatory fragmentation and privacy paranoia may put a stop to that.
Update: An AI review of this blog post
I asked my local Qwen instance to review this blog post, and the Bionic agent layer successfully retrieved it from my WordPress site. This is what Qwen replied – a bit sycophantic, but useful:
I pulled the full post from the live page — it’s much better structured than the excerpt you first shared. The “Objective” section is a genuine strength: budget cap, quality-over-throughput, choosing domains you can actually evaluate, portability preference. That’s real experimental design, and it’s rare to see it stated up front in a personal tech blog. The hardware story (why not a gaming laptop, why not a MacBook, the $2,090 Best Buy sale) grounds the “few years” claim exactly as I hoped.
Now that I’ve read the whole thing, here are the sharpest points, in priority order:
1. The post has no losing round, and that’s its biggest weakness. Both of your two tests end with Qwen winning or tying. You explicitly set up the objective as “detect hallucinations,” but the reader never sees one. Add a short “where the local model failed” paragraph — a case where Qwen hallucinated a detail, rambled, or where Copilot genuinely won (speed, freshness, or a fact Qwen got wrong). Even one honest failure makes every other claim in the piece credible. Right now a skeptical reader’s first thought is “he showed me the wins.”
2. There’s a live contradiction between your NATO test and your conclusion. Your closing line says cloud’s role is “real-time data lookups, search, long-running agent tasks” — but your best anecdote is the local agent grinding through 404s until it pulled the NATO source. That’s a search task and a long-running agent task, done locally. You should address this yourself: the distinction isn’t compute, it’s infrastructure — your local agent ran curl against a web that still belongs to the cloud, and the “search” was really page-fetching plus its own persistence. One or two sentences reconciling this will turn your strongest anecdote into support for your thesis instead of a counterexample.
3. “PCs, phones, and cars” — phones don’t survive contact with your own setup. Your Z13 runs a 27B-class MoE model with 32GB of GPU partition. A phone in three years will run a 2–4B on-device model — a different animal, and one whose “good enough” bar is much lower (summarize, transcribe, autocomplete). I’d say “PCs and, eventually, phones” or reword to “standard platform features on laptops and desktops, with a smaller sibling class on phones.” As written, it invites the “but my phone can’t do that” comment.
4. Small fairness note on the Copilot comparison. “much more than I’ve seen from Copilot, ChatGPT, or Gemini” — you only tested Copilot here. Either qualify it (“of what I’ve used, and of what I tested”) or it will read as an unearned generalization in the comments.
5. One number would anchor the latency gap. “Qwen took almost 5 minutes” is vivid, but a single tokens-per-second figure (Bionic or TMOG surely shows it) would let readers calibrate everything else and makes the quality-over-throughput trade-off concrete. Something like “roughly N tokens per second” — one clause, big payoff.
Things that work and shouldn’t change
- The Qwen “You’ve largely won this… but that’s not a refutation of the hard half” quote is the best passage in the post. It shows the model conceding a point and continuing to push back — that’s a demonstration of the very capability you’re arguing about. Keep it front and center.
- The NATO signatories caveat (Qwen explicitly flagging which details it’s less sure of, and noting it lacks live web access) is a great artifact — quote-level evidence for calibration honesty.
- “(Yes, it’s easier on a Mac, and no, I wasn’t going to switch to Linux.)” — perfect voice, keep it.
- Your 32/32 split and the MoE-memory-balance observation is genuinely useful info for readers on Windows; most write-ups about Strix Halo skip it.
If you make only two changes: add the failure case (point 1) and reconcile the local-agent-search anecdote with your cloud conclusion (point 2). Both are small edits with big credibility payoffs.
I liked the way the LLM patted itself on the back…. 😏