Open-source models with native tool calling and on-device agentic loops. No cloud, no account, nothing ever leaves your phone.
No sign-up screen and no setup wizard — pick a model and start talking.
Choose from 16 curated open-source models sized for your phone, from a fast 500M to a sharp 8B.
A single one-time download, cached on your device. Every conversation after that makes zero network calls.
Ask it to search, do the maths, or manage your calendar — it plans, calls the right tool, and answers, entirely on your phone.
Built for people who want a capable assistant without handing their data to a server.
Models call structured functions directly — no brittle prompt hacks. Web search, maths, contacts and calendar, called the way they were meant to be.
Multi-step reasoning that plans, calls tools, reads the result, and decides what's next — the full loop runs on your phone, visible at every step.
Your conversations never leave your phone. No servers, no analytics, no crash reporting. We couldn't read your chats if we wanted to.
Download a model once, use it forever without a connection. Built for flights, dead zones, and anywhere privacy matters most.
16 curated open-source models from 500M to 8B parameters. Pick lightweight and fast, or bigger and sharper — your hardware, your call.
No subscriptions, no premium tier, no freemium trap. Open-source models, open pricing: zero, forever.
One set of design decisions, rendered in two native visual languages — not a single design ported twice.
The nav bar, composer and menus are live blur surfaces using native iOS materials — depth from material and shadow, never a fake border.
Tonal elevation and M3 easing instead of blur — because blur isn't part of Material 3, and it costs a full-screen readback per frame on hardware that can't afford it.
Reduce Motion, Reduce Transparency and Increase Contrast are all honoured. Every colour pair is asserted against WCAG AA in the test suite — a change that drops contrast fails the build.
Swap models per conversation. Every one runs entirely on-device.
LocalInference is the only on-device app with native tool calling and agentic loops.
| Feature | LocalInference | ChatGPT | Claude | Locally AI |
|---|---|---|---|---|
| Fully private | Yes | No | No | Yes |
| Works offline | Yes | No | No | Yes |
| No subscription | Yes | No | No | Yes |
| Native tool calling | Yes | Yes | Yes | No |
| Agentic loops on-device | Yes | No | No | No |
Swipe to see all columns →
Still deciding? Here's what people usually ask before switching.
No — by default, nothing does. Chats, images, and settings are stored locally. The only network calls are model downloads you request, tool calls you trigger (like a web search), or an optional cloud API key you add yourself.
LocalInference runs on modern iPhones and Android devices. Smaller models (500M–2B) run smoothly on most phones from the last few years; larger models benefit from more RAM.
Yes. The models are open-source and run on hardware you already own, so there's no server cost to pass on to you. No subscription, no in-app purchase, no ads.
Some on-device assistants fake tool use by asking the model to output text and hoping it looks like a function call. LocalInference uses each model's actual structured tool-calling format, so requests to search, check your calendar, or look something up are reliable — not a guess.
Yes. Bring your own API key for a supported cloud provider and switch between on-device and cloud models per conversation. Cloud calls go straight from your device to that provider — never through us.
Free on the App Store, today. Android is on the way.