Run AI on your own computer: what you need and why
Download an open-weight model and run it on your own machine: private, offline and with no per-use cost. Here’s what you need, and what you give up.
In 30 seconds
- Local AI keeps your prompts on your machine, works offline and costs nothing per use.
- Memory is the limit: the model must fit in it. A graphics chip makes it much faster.
- Smaller and Jargon busterQuantisation: Storing a model’s numbers less precisely, so it needs less memory and runs faster, at the cost of a little quality. models trade some quality for size. Check licences, and download only from trusted sources.
Every question you type into a cloud assistant travels to someone else’s computers. For most things that’s fine. But for a confidential contract, or a long flight with no wifi, there’s another option: download an Jargon busterOpen-weight model: A model whose numbers are published, so anyone can download and run it themselves. and run it on your own machine.
Why run it yourself
- PrivacyYour prompts and documents are processed on your machine, not an AI company’s servers. Check the app’s own policy too.
- Works offlineOnce the model is downloaded, it runs with no internet connection: on a plane, a train or a remote site.
- No per-use costNo subscription or charge per message. You pay for the hardware and the electricity, however much you use it.
- ControlYou pick the exact model and version. It won’t change or disappear overnight, and you can adjust its settings.
What you need
You need a fairly recent computer, and memory matters most: to run at a usable speed, the whole model has to fit in memory, with room to spare for the conversation. A graphics chip, or Jargon busterGPU: Graphics processing unit: a chip built to do many small calculations at once. Designed for games and video, it also runs AI models far faster than a normal processor., speeds things up a lot. With a separate graphics card, what counts is the card’s own memory; without one, the model runs on the main processor, more slowly.
Macs with Apple’s own chips share memory between the processor and the graphics chip, which suits local AI. Model files are big, so you’ll need disk space too.
A model’s size is given in Jargon busterParameter: One of the numbers inside a model, set during training. A model’s size is given as its number of parameters, such as 8B for eight billion., such as 8B for eight billion: more usually means more capable, but bigger and slower. Many models also come in Jargon busterQuantisation: Storing a model’s numbers less precisely, so it needs less memory and runs faster, at the cost of a little quality. versions, labelled with codes such as Q4 or Q8. The lower the number, the smaller the file and the faster it runs, but the more quality you lose.
Getting started
Free apps such as LM Studio and Ollama download models and run them for you. Their catalogues include well-known families, such as Meta’s Llama, Google’s Gemma and models from Mistral. Both can also run a local server, so other apps on your computer can use the model.
Your first local model
- Check your machineNote how much memory you have, and your graphics card’s own memory if there is one.
- Install a trusted appGet it from the developer’s official website, not a link in a forum or a video.
- Start smallPick a small, quantised model from the app’s catalogue. Some apps warn you when one is too big for your machine.
- Test, then size upTry it on a task you know the answer to. If it’s quick and you want better answers, try the next size up.
Here’s a document I know well: [paste a few pages that contain nothing confidential]. Summarise it in five bullet points, list every date and figure it mentions, and tell me the one question it leaves unanswered.
Licences and security
Open-weight doesn’t mean free for any use. Licences differ: some let you do almost anything, while others restrict commercial use or set rules on how the model may be used. Read it before you use a model for work.
- Download only from trusted places: your app’s built-in catalogue or the model maker’s official page. Copies uploaded by unknown accounts may have been tampered with.
- Prefer formats built to hold only numbers, such as GGUF or safetensors. Some older formats can run hidden code when the file is opened.
- Keep the app updated, and don’t open its local server to your network or the internet.
- You’re responsible for the output. No company is screening what the model says, and it can still be confidently wrong.
When the cloud is still better
| On your computer | In the cloud |
|---|---|
| Prompts and files stay on your machine | Data goes to the provider, under its terms |
| Works offline | Needs a connection |
| No per-use cost, once you have the hardware | Free plans, subscriptions or pay per use |
| Smaller models, weaker on hard tasks | The most capable models |
| Knows nothing after its training | Often searches the web for current answers |
| You handle updates, safety and security | The provider handles them, and adds features |
So the cloud still wins for the hardest reasoning and coding, for up-to-date answers, for voice and image tools, and when your computer isn’t up to it. You don’t have to choose: use local for private drafts and offline work, and the cloud for the heavy lifting. At work, ask IT first: personal data on a laptop is still covered by Jargon busterUK GDPR: The UK law on personal data. It sets how organisations must collect, use and protect information about people..
Check yourself
3 quick questions nothing is savedSpotted a mistake? Tell us and an editor will check it.