Local Model or Claude? How to Tell If Your Mac's AI Is Earning Its Keep
James Ballard · September 26, 2026
If you've followed this series, your assistant is now answering you from a model running on your own Mac. You started with the overview of bringing your own AI model. You checked whether your Mac was up to it, connected everything step by step and built the daily habits that keep it running.
After a few weeks, the next question is whether the setup is actually working for you. It's a fair question. A local model gives you real control, but it also comes with trade-offs you only notice once you live with them. This post helps you evaluate your setup honestly and diagnose problems calmly. It also covers how to decide, without guilt, when a job belongs with Claude instead.
First, a Quick Recap of What Runs Where
It helps to be clear about which part of your setup does what, because it shapes every decision below.
- Your conversations go to the model running in oMLX on your Mac. The model does its work on hardware you own, and there's no per-message charge.
- Dex, the AI sysadmin that monitors, patches and repairs your setup, still uses Claude through your own Anthropic key. That's why the key stays connected. Dex's maintenance work costs far less than conversations do, but it isn't free, and it's billed to your Anthropic account.
- Your messages still travel through the apps you chat in, such as Telegram, WhatsApp or email. The local model changes where your assistant does its thinking. It doesn't change how those apps carry your messages.
Keep that picture in mind. Most of what follows is about the first bullet: whether the model on your Mac is the right brain for the work you're giving it.
Three Questions to Ask After Your First Few Weeks
1. Is it actually getting the job done?
With a local model, quality depends almost entirely on which model you picked. In our own testing, a less-capable model once reported that it had finished a step it hadn't actually done. Weaker models especially struggle with multi-step work, such as reading a file, acting on what's in it, and then reporting back. They can also struggle with images.
A simple way to check is to give your assistant a few tasks where you already know the right answer:
- Summarize an email attachment you've already read.
- Pull three specific details out of a document.
- Complete a two-step request, then check that both steps happened.
If the answers are solid, great. If the model skips steps, invents details or confidently claims success, that tells you something useful. It doesn't mean you did anything wrong. It means this model is out of its depth for that kind of task.
2. Is the speed livable?
In our testing on an M-series Mac, replies typically took about 10 seconds, with a range of roughly 5 to 46 seconds. Your numbers will depend on your machine and your model. Bigger models usually give better answers but reply more slowly.
For quick questions and drafting, a 10-second wait is often fine. For rapid back-and-forth, it can start to feel sluggish. Only you can decide where your limit is, so notice whether the wait bothers you in real use.
3. Is the Mac staying awake, with the model running?
This is the practical question that matters most. Your assistant can only answer while your Mac is awake and oMLX is running your model. After a restart, the model doesn't start again on its own. You have to reopen oMLX yourself. This is the single most likely thing to go wrong, and that includes the restarts that come with system updates.
If you're reopening oMLX without thinking about it, you've built the habit. If you keep discovering hours later that your assistant has been silent, this setup is asking more of you than it's giving back.
When Your Assistant Goes Quiet: A Calm Checklist
Here's what actually happens when the model on your Mac becomes unreachable, so it doesn't catch you off guard.
- There's no automatic switch to Claude. If the model stops, your assistant stops answering. It won't quietly hand your conversation to the cloud instead.
- A message gives up after about 3 minutes if it can't get a reply from your model.
- We typically notice within about 12 minutes and send you one email. There's one important exception. If your assistant and your model are on the same Mac and that Mac goes to sleep, the email can't go out either. In that case, the unanswered message is your only sign that something's wrong.
When you notice the silence, work through these steps in order:
- Is the Mac awake? Wake it and check that it hasn't been set to sleep on its own schedule.
- Did it restart? If so, reopen oMLX. A restart is the most common cause of a silent assistant.
- Is your model loaded and running in oMLX? Opening the app isn't enough if the model itself isn't running.
- If your assistant runs on a separate computer, such as a Raspberry Pi or a spare machine, is that computer on and connected to the same home network as your Mac?
- Send a short test message and give it a little time to answer.
If all of that checks out and it's still quiet, ask Dex in chat to help look into it.
Matching the Work to the Right Brain
You don't have to be loyal to one model. A sensible way to think about it:
Tasks that tend to suit a capable local model:
- Quick questions and everyday drafting
- Rewording, tidying up or summarizing text you've pasted in
- Conversations where keeping the thinking on your own hardware matters most to you
Tasks where Claude may serve you better:
- Multi-step jobs involving files, attachments or several actions in a row
- Work involving images
- Anything where a confident wrong answer would cause real problems
Whichever model you use, double-check anything important before you rely on it. No AI model gets everything right.
If you have the memory, a larger model may narrow the gap. A large model needs about 15GB to itself. That's comfortable on a 32GB Mac, but a 16GB machine can only run a noticeably smaller model. Our hardware guide covers this in more detail.
How Switching Actually Works
Switching between your local model and Claude is a single setting in your dashboard, one click in each direction. There's no model name to type, because the setting reads the list of models directly from your Mac.
The two directions behave differently in one way:
- Switching back to Claude starts a fresh conversation. Your assistant won't carry the current thread's context over to Claude.
- Switching to your local model doesn't start a fresh conversation.
So if you plan to switch to Claude for a big task, do it at a natural stopping point rather than in the middle of a detailed exchange.
Also remember that when Claude handles your conversations, per-message charges resume on your Anthropic account. Dex's monthly budget alert and runaway-spend circuit breaker keep an eye on that spending either way. The tier you choose in the dashboard affects the cost too. For the full picture of what you pay and to whom, see our breakdown of what a private AI assistant really costs per month.
The Honest Framing: Control, Not Savings
It's tempting to judge this setup by how much you save, but that's the wrong measure. Your LaunchMy.ai subscription stays the same whether you use a local model or not. You'll still pay a little for Dex's Claude usage. The only running cost for your conversations is electricity, but in exchange you're committing to a Mac that stays awake and a model you reopen after every restart.
The real benefit is control. You choose the model, and your assistant does its thinking on hardware you own. If that matters to you, the trade-offs may well be worth it. If keeping your Mac awake and reopening oMLX has become more hassle than the control is worth to you, switching back to Claude is a completely reasonable choice. You can come back to your local model later.
A Reminder of What's Supported
To set clear expectations:
- oMLX is the only local model tool we support and have tested.
- Two setups are supported: your assistant on the same Mac as the model, which we recommend, or on another computer on the same home network.
- Cloud assistants can't use a local model, because a cloud server can't reach your home network. An existing cloud assistant also can't be moved onto your Mac. You'd need to set up a new assistant on your own computer.
- We support the connection between your assistant and your model. We don't support your Mac itself or the particular model you chose.
Where to Go From Here
At this point you know what running a local model is actually like, beyond the theory. Keep the tasks that work well on your Mac, send the harder jobs to Claude when it makes sense, and lean on Dex for the maintenance work. For more on that side of things, read how Dex keeps your assistant running without a terminal.