Launch offer: 40% off for 6 months — just $5.99/mo (reg. $9.99)Claim the deal

← All posts

Living With a Local AI Model on Your Mac: Practical Habits That Make It Work

James Ballard · September 26, 2026

Getting a local AI model running on your Mac is the hard part, and if you've followed our earlier guides, you've already done it. This post covers what comes next: the day-to-day habits that separate a setup you enjoy from one that quietly stops answering on a Tuesday afternoon.

If you're just arriving, start with the earlier posts. Bring Your Own LLM: Use Your Mac's AI Capabilities With Launch My AI explains what the feature is and who it's for. Can Your Mac Run a Local AI Model? An Honest Hardware Guide covers memory and Mac models. How to Run Your AI Assistant on a Local Model on Your Mac walks through setup.

A quick reminder of what we're talking about. "Bring your own AI model" is an advanced, opt-in setting in your LaunchMy.ai dashboard. It points your assistant's everyday replies at a model running on your own Apple-silicon Mac, through the Mac app oMLX, instead of Claude in the cloud. It's not a separate plan and doesn't change your subscription. People choose it for control and privacy: your conversations are answered on hardware you own. It isn't meant to be a bargain.

Habit #1: Open oMLX every time your Mac restarts

If you remember only one tip from this post, make it this one.

After a restart, your assistant comes back on its own. The model does not. oMLX has to be opened by hand. Until you open it, your assistant has nothing to think with, and your messages go unanswered.

In our experience, this is the single most likely reason a local setup "breaks." A few ways to make the habit stick:

  • Pair it with something you already do. If you open your email first thing after a restart, open oMLX right before it.
  • Watch for the overnight software update. macOS updates often restart your Mac while you sleep. If your assistant is quiet in the morning, check oMLX before anything else.
  • Leave yourself a reminder. A sticky note on the monitor works fine: "Restarted? Open oMLX."

Habit #2: Keep your Mac awake

A sleeping Mac can't answer your messages. If your Mac sleeps, loses its network connection, or the model stops, your assistant stops answering.

In System Settings, look at your Mac's energy or battery options. Turn on the setting that prevents automatic sleep while the Mac is plugged in. The exact wording varies by Mac and macOS version.

On a MacBook, closing the lid usually puts it to sleep. If a laptop is your model machine, keep it plugged in and open, or treat your assistant as available only while you're using the laptop.

Your Mac's display can turn off. What matters is that the computer itself stays awake.

Habit #3: Know how you'll find out it's down (and when you won't)

Here's how an outage plays out, because it affects how you should react to silence:

  1. You'll probably notice first. A message to your assistant gives up after about 3 minutes if the model isn't answering.
  2. We detect it within about 12 minutes and send you one email.
  3. On a one-Mac setup, a sleeping Mac sends no alert. The part of your setup that would send the email is asleep too.

So if your assistant goes quiet, don't wait for an email. Check the Mac first: is it awake, is it online, is oMLX open? Those are the usual culprits. For a fuller walkthrough, see When Your Mac's Local AI Model Stops Answering: What Happens and How to Get It Back.

Our team can tell you whether your assistant can reach your Mac. We can't support the Mac itself or the particular model you've chosen. When you own every piece of the setup, the Mac and the model are yours to look after.

Habit #4: Send a warm-up message

In our testing, a typical reply took about 10 seconds. That's the median across a 22-turn session, with replies ranging from 5 to 46 seconds. Speed depends heavily on your Mac and the model you've chosen.

The first reply after the model loads can be much slower than that. After you open oMLX, send a throwaway "hi" and let it answer while you make coffee. Your first real question should then come back closer to typical speed.

It also helps to set your own expectations. A local model feels more like texting a friend who takes a moment to reply than using an instant search box. If you need snappy back-and-forth, see Habit #5.

Habit #5: Match the job to the model

This is where you get the most out of a local setup. Local models vary a lot, and results depend on the one you choose. As a general rule, less-capable models are better suited to lighter tasks and struggle with complicated ones. For help deciding which jobs belong where, see Local Model or Claude? How to Tell If Your Mac's AI Is Earning Its Keep.

Lighter tasks that are generally a better fit:

  • Talking through an idea or a decision
  • Drafting a quick message or reply
  • Summarizing something you paste in
  • Simple questions you'd rather keep on your own machine

Where less-capable models tend to struggle:

  • Multi-step jobs, like reading a file, acting on it, then reporting back
  • Working with pictures
  • Anything where one wrong step quietly spoils the result

For the harder work, switching back to Claude takes one click in the dashboard. Before you click, know how conversations carry over:

  • **Switching to your local model keeps your current conversation.**
  • **Switching back to Claude starts a fresh conversation.**

So plan ahead. If a session will start with a demanding, multi-step task, begin on Claude. Once the heavy lifting is done, switch to your local model for the lighter back-and-forth, and your conversation carries over. Going the other direction means starting over.

Your assistant also won't switch to Claude on its own if your local model stops. It stays on the model you chose until you change it in the dashboard.

Habit #6: Trust, but verify

Less-capable models can be confidently wrong. In one case, a model reported a step as finished when it hadn't happened. For anything that matters, check the result yourself. Did the file actually arrive? Is the draft actually saved?

You may also see a model include its private "thinking" in its replies: long passages where it reasons to itself before answering. Some models do this. If it bothers you, try a different model in oMLX. Your dashboard reads the list of available models straight from oMLX, so you pick from a list instead of typing names.

Habit #7: Leave memory headroom

As the hardware guide explains, a large model needs about 15GB of memory on its own. That's comfortable on a 32GB Mac. A 16GB Mac only works with a noticeably smaller model.

In daily use, the model shares memory with everything else you're running. A browser with forty tabs, a video call and a photo editor all compete with it. If replies suddenly get much slower during a busy workday, your Mac may be running short on memory. Close a few heavy apps, or consider a smaller model for everyday use.

Habit #8: If you use two machines, keep them on the same home network

There are two supported setups:

  1. One Mac, nothing else. Your assistant and the model both run on the same Mac. This is the one we recommend: fewer moving parts, fewer things to check.
  2. Two machines on one home network. Your assistant runs on a Raspberry Pi or spare computer and reaches the Mac over your home Wi-Fi.

With two machines, both have to stay on the same home network. Watch for:

  • The Mac leaving the house. If your model machine is a MacBook, your assistant loses its model the moment you take the laptop to a coffee shop.
  • Guest networks. Some routers keep devices on a guest network separated from the main one. Put both machines on the same network.
  • Router restarts. After a power blip, confirm both machines reconnected.

Reaching your model over the internet isn't available. The product won't set it up. For the same reason, an assistant running on a cloud server can't use a local model. If your existing assistant is in the cloud, you'd set up a new assistant on your own hardware first. Keeping both would count as a second assistant on your plan. The BYOD announcement covers running your assistant on your own computer.

Habit #9: Keep your Claude key connected and topped up

Even with a local model answering your conversations, your Anthropic (Claude) key is still required. Dex, the AI systems administrator that comes with every assistant, keeps using Claude for its own AI work. Dex's usage costs far less than conversations would, but it isn't zero, and it's billed to your own Anthropic account.

If your Anthropic credit runs out, Dex's AI-powered work can't run. Keep an eye on your balance in your Anthropic account. Dex's AI usage and cost monitoring, including its monthly budget alert, helps you keep track of what you're spending.

So "local" doesn't mean "free." Your running costs are electricity for your Mac, plus a small amount of Claude usage for Dex, on top of your normal LaunchMy.ai subscription. For the full picture, see How Much Does It Really Cost to Set Up Your Own Private AI Assistant Per Month?

A one-minute daily routine

Here's the whole post, condensed:

  • After any restart: open oMLX, then send a warm-up message.
  • If your assistant goes quiet: check that the Mac is awake, online and running oMLX, before anything else.
  • Before a demanding task: start on Claude. Switch to local afterward to keep the conversation.
  • For anything important: double-check what the model says it did.
  • Once a week: glance at your Anthropic balance and Dex's weekly trend report.

Is this the right setup for you?

Be honest with yourself. If you want an assistant that needs less looking after, without habits like these, the standard setup is likely the better fit. That means Claude, running on a cloud server in your own account, with Dex doing the maintenance. You pay your cloud provider for that server directly, separate from your subscription. How to Keep Your Private AI Assistant Running Without Touching a Terminal shows what that looks like.

If you enjoy tinkering, want your everyday conversations answered on a machine you physically own, and don't mind opening an app after a restart, a local model can be rewarding. You choose the model, and the replies come from your own Mac rather than a cloud AI service. Your messages still travel through whichever chat app you use to reach your assistant, and Dex still relies on Claude for its maintenance work. But the everyday thinking happens on hardware you control.