What if you could have an AI conversation on your phone without sending a single message to the internet?
I tried it, and it was surprisingly reassuring. I downloaded a language model, put my phone into airplane mode, and started chatting. No cloud processing or internet connection was required. The AI was running right there on my phone.
It comes with some compromises, but having an AI that works completely offline and keeps the conversation on your device is more useful than I expected.
I put an AI model on my phone
Putting local AI to the test
Android offers several ways to run a local LLM. MLC Chat is a great option if you want to utilize supported phone hardware for faster inference.
Google AI Edge Gallery takes a different approach, offering Google’s on-device models and features such as image and audio capabilities.
You can also try advanced options, such as Maid, LM Playground, or running Ollama through Termux if you don’t mind a more technical setup.
I chose PocketPal AI because I wanted something straightforward that let me choose the model.
The app is open source, with its code publicly available on GitHub. For most, though, the easiest way to install it is through the Google Play Store.
PocketPal supports GGUF models from Hugging Face, so you can choose models such as Gemma, Qwen, and Phi and keep the one you want on your phone.
How I set up PocketPal on my phone
Getting my first local model running


A local LLM isn’t like installing an Android app. Models can take up significant storage and require enough memory to load and run. PocketPal recommends at least 6GB of RAM for smaller models and 8GB or more for larger ones.
Newer phones with more RAM and faster chips are better suited to running AI models locally.
After installing PocketPal, tap Download Model on the home screen. You’ll see the available models, along with details such as their size, parameter count, and capabilities.
Tap a model to see more information and tap Download to save it to your phone. After the download, tap Load on the card to load it into memory and start chatting.
I chose Gemma 3 1B, downloaded it, and tapped Load when it was ready.
To add more models, I opened the hamburger menu, tapped Models, and used the option to add a model. PocketPal gives you a few ways to do this: Add from Hugging Face, Add local model, or Add remote model.
My conversations stayed on my phone
I could chat even when I was offline


After I downloaded Gemma 3 1B, I could start chatting with it in PocketPal. The responses were noticeably slower than what I’m used to from cloud-based AI, but that was also the trade-off I expected from running the model on my phone.
PocketPal uses llama.cpp to run GGUF models locally, using the phone’s available CPU, GPU, or supported NPU hardware.
From the chat screen, I could tap the arrow icon in the text box and switch between models.
The same menu gives you access to Pals, which are preconfigured AI personalities. I could switch to Pip, a general-purpose assistant designed to run locally, or try Lookie, which is a more unusual Pal that can analyze video from the phone’s camera.
Lookie is designed for real-time, on-device video analysis. However, its performance depends on the model and the phone you’re using.
PocketPal also lets you add voices, so you aren’t limited to typing and reading responses. Its voice options include Kitten, Kokoro, Supertonic, and your phone’s system voices.
What makes offline AI worth it?
A slower AI can still be useful
For me, the biggest appeal is privacy. Some conversations are better kept on my phone, and with a local model, my prompts and responses don’t leave the device for processing.
The other advantage became apparent when I switched off my connection. After the model is on the phone, I can use it on a flight, somewhere with poor reception, or when I don’t want to connect to a cloud service.
There is also something appealing about having a model I control. I can download different GGUF models, keep them on my device, and switch between them. I don’t have to depend on the model a cloud AI service offers.
Still, I know I’m giving up some things. The responses on my phone were slower. A small local model can’t compete with the largest cloud models on every task.
The biggest downside of keeping everything offline is that the AI is cut off from the outside world. My local model couldn’t tell me what was happening on the news today, look up a breaking story, or answer questions that depended on current information.
My phone doesn’t need the cloud for every AI chat
Running an LLM on my phone won’t replace Gemini or ChatGPT. Its responses were slower, the smaller model wasn’t as capable, and being completely offline meant it couldn’t tell me anything about current events.
However, I could have a conversation without sending those prompts to a server, and I could keep using it when my phone had no internet connection.
Now that I know how easy it is to run an AI model locally, I’m more interested in experimenting with what else my phone can handle.


