Features How it works Design Models Compare Privacy Get the app
Now on the App Store

Run AI on your phone.
Fully private.

Open-source models with native tool calling and on-device agentic loops. No cloud, no account, nothing ever leaves your phone.

Zero tracking Works offline No account needed
100% on-device
Agentic loop
How it works

From download to answer in under a minute

No sign-up screen and no setup wizard — pick a model and start talking.

1

Pick a model

Choose from 16 curated open-source models sized for your phone, from a fast 500M to a sharp 8B.

2

It downloads once

A single one-time download, cached on your device. Every conversation after that makes zero network calls.

3

Chat, call tools, get it done

Ask it to search, do the maths, or manage your calendar — it plans, calls the right tool, and answers, entirely on your phone.

Why LocalInference

Everything a cloud assistant does. Nothing it shouldn't.

Built for people who want a capable assistant without handing their data to a server.

Advanced

Native tool calling

Models call structured functions directly — no brittle prompt hacks. Web search, maths, contacts and calendar, called the way they were meant to be.

Advanced

On-device agentic loops

Multi-step reasoning that plans, calls tools, reads the result, and decides what's next — the full loop runs on your phone, visible at every step.

Private by design

Your conversations never leave your phone. No servers, no analytics, no crash reporting. We couldn't read your chats if we wanted to.

Works fully offline

Download a model once, use it forever without a connection. Built for flights, dead zones, and anywhere privacy matters most.

Choose your model

16 curated open-source models from 500M to 8B parameters. Pick lightweight and fast, or bigger and sharper — your hardware, your call.

Always free

No subscriptions, no premium tier, no freemium trap. Open-source models, open pricing: zero, forever.

Built with care

It looks like it belongs on your phone

One set of design decisions, rendered in two native visual languages — not a single design ported twice.

iOS

Real system materials

The nav bar, composer and menus are live blur surfaces using native iOS materials — depth from material and shadow, never a fake border.

Android

Proper Material 3

Tonal elevation and M3 easing instead of blur — because blur isn't part of Material 3, and it costs a full-screen readback per frame on hardware that can't afford it.

Everyone

Accessible by default

Reduce Motion, Reduce Transparency and Increase Contrast are all honoured. Every colour pair is asserted against WCAG AA in the test suite — a change that drops contrast fails the build.

Model library

Powered by leading open-source models

Swap models per conversation. Every one runs entirely on-device.

500MSmolVLM2 0.6BQwen3 1BGemma 3 1.2BLFM2 1.5BDeepSeek R1 Distill 3BLlama 3.2 8BQwen3 + 9 more
How it stacks up

Privacy of on-device. Capability of the cloud.

LocalInference is the only on-device app with native tool calling and agentic loops.

Feature LocalInference ChatGPT Claude Locally AI
Fully private Yes No No Yes
Works offline Yes No No Yes
No subscription Yes No No Yes
Native tool calling Yes Yes Yes No
Agentic loops on-device Yes No No No

Swipe to see all columns →

16
Open-source models
0
Bytes sent by default
$0
Subscription cost
2
Features no on-device rival has
FAQ

Questions, answered

Still deciding? Here's what people usually ask before switching.

Does anything ever leave my phone?

No — by default, nothing does. Chats, images, and settings are stored locally. The only network calls are model downloads you request, tool calls you trigger (like a web search), or an optional cloud API key you add yourself.

Will it work on my phone?

LocalInference runs on modern iPhones and Android devices. Smaller models (500M–2B) run smoothly on most phones from the last few years; larger models benefit from more RAM.

What's the catch — is it really free?

Yes. The models are open-source and run on hardware you already own, so there's no server cost to pass on to you. No subscription, no in-app purchase, no ads.

What is "native tool calling," really?

Some on-device assistants fake tool use by asking the model to output text and hoping it looks like a function call. LocalInference uses each model's actual structured tool-calling format, so requests to search, check your calendar, or look something up are reliable — not a guess.

Can I use my own OpenAI/Anthropic key?

Yes. Bring your own API key for a supported cloud provider and switch between on-device and cloud models per conversation. Cloud calls go straight from your device to that provider — never through us.

Your phone. Your model. Your data.

Free on the App Store, today. Android is on the way.

Get the app