Launch offer: 40% off for 6 months — just $5.99/mo (reg. $9.99)Claim the deal

← All posts

When Your Mac's Local AI Model Stops Answering: What Happens and How to Get It Back

James Ballard · September 26, 2026

At some point your local model will stop answering. You send your assistant a message on Telegram, and nothing comes back.

This is the sixth post in our series on running your assistant on a model hosted on your own Mac. We've covered what bring-your-own-model is, whether your Mac can handle it, how to set it up, daily habits that make it work, and how to decide whether it's earning its keep.

This post covers the part most guides skip: what happens when the model goes quiet, how you'll find out, and how to get back up and running.

Why this deserves its own post

When your assistant uses Claude, the model lives in a data center you never think about. When it uses a model on your Mac, your Mac is the data center. That brings real benefits: no per-message charge, and conversations that stay on your machine. It also means the model is only available while your Mac is awake, logged in, and running the model app.

We made a deliberate choice here. If your local model stops, your assistant does not quietly switch over to Claude. Some people are surprised by that. The reason is that you picked a local model for a reason, often privacy. Sending your messages to the cloud behind your back would undo that choice. So when the model stops, the assistant stops answering until you decide what to do.

What actually happens when the model stops

Here's the sequence, based on how the feature works today.

1. You'll probably notice first. You send a message and it goes unanswered. Your assistant waits for the model and gives up after about three minutes. In most cases, that silence is your first sign that something's wrong.

2. We detect it within about 12 minutes. If your assistant is running but can't reach the model, Launchpad detects the problem within about 12 minutes. You'll get one email explaining what's going on. It's a single email, not a stream of alerts.

3. The exception: one Mac that's gone to sleep. If you use the recommended one-Mac setup, your assistant and model share the same machine. If that Mac falls asleep, your assistant is asleep too, so nothing can notice the problem and no alert is sent at all. The unanswered message is your only signal.

If you use the two-machine setup, the situation is different. In that setup, your assistant runs on something like a Raspberry Pi or spare computer and reaches the Mac over your home wi-fi. If the Mac drops off, the assistant is still awake to notice, so the detection email applies.

The most likely culprit: a restart, not sleep

Many people assume sleep will cause most outages. In practice, a Mac restart is the more likely cause.

After your Mac reboots, the model doesn't start again by itself. You need to reopen oMLX (the free Mac app that runs your model) and make sure your model is loaded and running. Until you do, your assistant has nothing to talk to.

This is easy to miss. macOS can restart for a system update overnight, and everything looks normal the next morning. Our own test machine sat in exactly this state for six days. If it can happen to the people who built the feature, it can happen to you.

Your recovery checklist

When your assistant goes quiet, work through these checks in order. The cause usually turns up in the first three.

1. Is the Mac awake?

Wake it up. If you're on a laptop, check whether the lid was closed.

2. Did the Mac restart?

If you see a login screen, or apps you had open are gone, it probably restarted. Log in and go to step 3.

3. Is oMLX open, with your model running?

Open oMLX and confirm your model is loaded. Your assistant reads its list of models straight from oMLX, so you don't need to retype anything in the dashboard. You just need the model to be running.

4. Two-machine setup only: are both machines on the same home network?

The two-machine setup works only across the same home network. If your Mac joined a different wi-fi network, such as a guest network, or left the house, your assistant can't reach it. Reaching your model over the internet is not supported. It isn't a hidden setting; the product refuses that setup.

5. Send your message again

Once the model is back, re-send whatever went unanswered. The original message timed out, so you'll likely need to send it again rather than wait for a late reply.

What we can and can't help with

Launchpad can tell you whether your assistant is able to reach your Mac. That covers the most useful question: "Is the connection working?"

What we can't do is support your Mac itself or the model you chose. If oMLX misbehaves, a model runs slowly, or your Mac has a hardware issue, that's outside what we can fix. That's the trade-off of running a model on hardware you own. You get full control, and you also take on the upkeep. Dex, your built-in AI sysadmin, still looks after the maintenance and repair of your assistant's setup. But the model and the Mac hosting it are your responsibility.

The one-click escape hatch: switching back to Claude

Sometimes you don't have time to troubleshoot. Maybe oMLX crashed while you're busy, or you're on the two-machine setup and your Mac is out of reach, and you need your assistant now.

Switching your assistant back to Claude takes one click in the dashboard. Three things to know before you do it:

  • It only helps if the machine running your assistant is awake. In the one-Mac setup, your assistant lives on that Mac. If the Mac is asleep or switched off, your assistant is too, and switching to Claude won't bring it back until the Mac is awake again. The switch is most useful when your Mac is on but the model has stopped, or in the two-machine setup.
  • Switching back to Claude starts a fresh conversation. Your old conversations are kept, not deleted, but the assistant won't pick up mid-thread.
  • Switching to your local model later does not reset the conversation. So the fresh start happens only in one direction.

Your Anthropic key needs to stay connected either way. Dex keeps using Claude for maintenance and repairs even when your conversations run locally. That costs much less than conversations do, but it isn't zero. That's also why there's no such thing as a "completely free" setup here, even though your everyday replies carry no per-message charge.

Habits that prevent most outages

You can't prevent every outage, but a few habits cut down how often they happen. Menu names vary between macOS versions, so treat these as directions to look rather than exact clicks.

  • Stop your Mac from sleeping while it's hosting the model. On a desktop Mac, look in System Settings under Energy (or Battery on laptops) for options that prevent automatic sleep. Laptops generally sleep when the lid closes, so a Mac that stays open or docked works better as a model host.
  • Consider having oMLX open at login. macOS lets you add apps to Login Items in System Settings. Afterwards, check whether oMLX reloads your model automatically. If it doesn't, you'll still need to start the model yourself. Also note that if your Mac requires a password after restarting, login items won't run until someone actually logs in.
  • Check in after macOS updates. If you know an update installed, send your assistant a quick "you there?" message. It takes ten seconds and could save you days of silence.
  • Be skeptical of "done." This isn't an outage issue, but it matters for trust. Smaller models can be confidently wrong. In our testing, one reported a step as finished when it wasn't. Smaller models also struggle with multi-step jobs like reading files or handling pictures. For those tasks, Claude is the safer choice. Our post on when to switch goes deeper.

A reality check on speed

People sometimes mistake a slow reply for a broken one. In our testing on one M-series Mac, a typical reply took about 10 seconds, with a range of 5 to 46 seconds across a single 22-message session. Your numbers will depend on your machine and the model you run. If a reply is slow but arrives, the model is working. If nothing arrives after a few minutes, that's when to work through the checklist.

A few common questions

Can I use my local model with an assistant that runs in the cloud? No. Bring-your-own-model works only when your assistant runs on your own hardware: on the same Mac or another computer in the same home. An assistant already running on a cloud server can't be moved onto a Mac. You'd need to set up a new assistant on your own hardware, which our BYOD announcement explains.

Does using a local model change my subscription? No. It's a setting, not a separate plan or add-on, and your subscription stays the same.

Why not just fall back to Claude automatically? An automatic fallback would send your conversations to the cloud without asking. We'd rather you notice, decide, and click.

The bottom line

A local model gives you more control, and more control means a little more responsibility. The failures are predictable: a restart, a sleeping Mac, a closed app. The fixes are mostly simple: wake the Mac, reopen oMLX, resend your message. And when your Mac is awake but the model isn't cooperating, Claude is one click away.

If you'd rather not think about any of this, that's a perfectly reasonable choice. The standard Claude setup doesn't depend on a model running on your Mac, and Dex helps keep it maintained without you touching a terminal. But if you run a local model, knowing how it fails and how to recover is what makes it practical to live with.