Local artificial intelligence: how to protect your data and work without the internet
How to use AI offline when financial statements and client data cannot go into somebody else’s cloud
For owners and executives who want to automate routine work without letting corporate information out of the company. After this article you will understand how to start a digital assistant right on your working laptop, how to pick a model to match your hardware, and why this option often turns out cheaper than a subscription.
You are travelling, the connection comes and goes, and yesterday’s meeting has to be worked through today. A cloud service simply will not open at a moment like that.
With a local model the picture is different. You load your drafts into it and ask for a business letter, or for the notes to be broken into specific tasks for the team. You throw in photographs of receipts and screenshots and get a table. Not a single byte leaves your device while it happens.
In AKORDO’s practice the question of security comes up regularly and in roughly the same words: how to use the new tools without putting financial statements and the client base at risk. Local AI removes exactly that dilemma, because it runs on your laptop’s own capacity and depends neither on the internet nor on somebody else’s privacy policy.
How a local model differs from a cloud one
When you use a popular service through a browser, your request goes to a large company’s servers, where an LLM processes it (a large language model able to understand and generate text on the basis of large bodies of data). You depend on connection speed, on having a subscription, and on what the provider decided to write in its data processing policy.
A local model is a file you download onto your own disk. Three differences follow from that, and they matter to a business.
Privacy. You put financial documents and client data into the model, because the processing happens inside the device.
Autonomy. The tool works on a plane, on a train, and where Wi-Fi works every other time.
Cost. You pay neither per request nor by monthly subscription, because you are using your own hardware. Paid for earlier, but your own.
Google, Meta, and Microsoft publish models like these openly, building ecosystems and a research base around them. For you that means developers all over the world are adapting these models to specific languages and to narrow professional tasks, including medical and corporate systems where regulation forbids passing data outside.
How to pick a model for your hardware
To avoid overloading the laptop and being disappointed by the quality, you need to understand two characteristics.
The first is the number of parameters, marked with the letter B and counted in billions. Parameters are the neural connections the model uses when it picks the next word. Put simply, it is the size of its brain.
Models of 1B to 4B parameters are light and start on an ordinary working laptop. They are fast, but they fall noticeably behind on complex reasoning. Models from 7B to 14B and above need stronger hardware, a MacBook with 16 GB of RAM for example, and in return they reason far more deeply.
The second characteristic is called quantization. That is the compression of a model so it takes less space and runs faster. A model in Q4 format is compressed and therefore lighter for the computer. Compression is never free in quality terms: the more aggressively a model is compressed, the more it simplifies its answers, so with very small formats it is worth checking the result more carefully.
For a start on business tasks we recommend 4B parameters with Q4 quantization. That is a sensible balance of speed and quality for the first week. Also watch for the IT or Instruct marking in the name: it means the model has been trained to carry out instructions rather than simply continue a text.
What you can hand it
A local model does not replace powerful cloud systems. It covers specific stretches of the pipeline well, meaning the sequence of stages in your process.
Correspondence and notes. A draft turns into a business letter, a long text into a short summary, a recording of a meeting into a list of tasks with owners.
Documents and figures. The model pulls data out of receipts, screenshots, and tables. For calculation and work with numbers, specialized small models do well, including those from Microsoft.
Texts in different languages. Individual families of models, the Chinese ones among them, give high quality in structuring information.
Visual analysis. Google’s models recognize images, which covers digitizing paper reports and financial documents without going online.
This is the routine that takes hours every day and needs no deep thinking. It is what to hand over first.
What you need to start
You will not have to program. Three things are needed: the model file, an application to run it, and your computer.
The models sit on Hugging Face, the largest library of neural networks in the world. In logic it is an app store, only for AI. They are convenient to run through the free applications Ollama or LM Studio, whose interface looks like an ordinary chat.
The order of operations is this:
- Look at your device’s specification: the type of processor and the amount of RAM.
- Choose a model in the application’s catalog marked as compatible with your hardware, usually a green check mark.
- Download it in one click or through the terminal, meaning the window where you give the computer commands directly.
- Turn the internet off and start work.
If the model answers slowly or the computer starts to stall, it is too big for your machine. Take a version with fewer parameters or stronger compression.
Is this about you?
Recognize yourself in three points or more and it is worth trying:
- you work with financial statements, clients’ personal data, or information that cannot be passed to a cloud;
- your schedule has a lot of travel, flights, and places with an unstable connection;
- you have to process many operations of the same kind with text or images: receipts, formatting letters, transcripts;
- you want to build AI into internal corporate systems;
- subscriptions to cloud services for the whole team come out expensive.
Where to start
Install Ollama on your working computer and download a 4B-parameter model from the Gemma or Qwen family. Then give it real tasks for a week: an answer to a letter, working through a screenshot of a document, a summary of notes after a meeting. You will see both the speed on your own hardware and the ceiling on quality, at no cost and with no risk to the data.
One separate question is worth closing before the tool goes to the team. The word “open” in a model’s description does not always mean permission for commercial use: licences differ between families. Check the terms of the specific model you are taking before you put it into a business process.
Key takeaways
- A local model processes everything inside the computer, so confidential documents never leave the company perimeter.
- The choice of model comes down to hardware: for an ordinary working laptop the reference point is 4 billion parameters with Q4 quantization.
- Ollama and LM Studio let you run a neural network with no programming knowledge, and the interface is no harder than a chat.
- A local assistant covers document routine and offline work well, rather than replacing powerful cloud systems.
The AKORDO team has worked with fintech companies in three formats: an AI audit and roadmap, product consulting for fintech teams and four AI systems in production. The guide to the AI knowledge assistant explains how to prepare documentation before launch. If customer data will sit in a CRM, the EU CRM requirements checklist shows which GDPR duties belong in the specification.