On October 16, 2026, Google will switch off Gemini 2.5 Flash and Gemini 2.5 Pro on its developer API. Any app still calling those model names after that date will get errors instead of answers. For a small team that shipped an AI feature last year and then moved on to other work, that is an easy date to miss.
It is also a fair picture of what building on Google Gemini AI is like right now. The models are fast, reasonably priced, and they get better every few months. They are also a moving target. The model you tested in March may be retired by October. A setting your code relied on may be removed in the next version. The price you budgeted for may change on the first of January.
None of that is a reason to stay away. It is a reason to go in with open eyes. This guide is written for founders, developers, content teams and business owners who are deciding whether to build with Gemini, and for people who already have and want to do it better. We will cover what Gemini actually is, the different ways to plug it into an app, what it costs, and the less exciting questions that decide whether an AI feature survives contact with real users. What happens when data is missing? When two inputs disagree? When a decision has to be made in under a second? When ten thousand people show up at once?
First, what does "Gemini" refer to?
The name covers several different things, and mixing them up causes a lot of confused planning meetings.
▪ The Gemini app. This is Google's chatbot, the one people open on their phone or in a browser. Google said in August 2026 that it had passed one billion monthly users.
▪ The Gemini models. These are the engines underneath. Each one is a generative AI model, which means software trained on huge amounts of text, images, audio and code so that it can produce new content in response to a request, instead of looking up a stored answer.
▪ The developer products. These are the doors developers use to reach the models from their own software: the developer API, Google AI Studio, Firebase AI Logic, the Agent Platform on Google Cloud, and Gemini Nano on Android phones.
When people talk about Google Gemini AI in the context of app development, they almost always mean the second and third items. Your users never need to open the Gemini app. They use your app, and your app quietly talks to a Gemini model in the background.
The current lineup in plain terms
Google's models page, updated in late September 2026, recommends two models for new projects: Gemini 3.8 Flash and Gemini 3.5 Flash-Lite. Flash is the capable all-rounder. Flash-Lite is the cheaper, quicker option for simple jobs you run in large numbers, such as tagging support tickets or cleaning up product descriptions.
Gemini 3.8 Flash can read up to one million tokens in a single request and write up to 64,000 tokens back. A token is a small chunk of text, roughly three quarters of an English word on average. One million tokens is somewhere around 750,000 words, which is several long novels, a large contract bundle, or a mid-sized codebase in one go.
The model also has adjustable "thinking levels" (low, medium and high). Thinking here means the model works through a problem in steps before it answers. More thinking gives better results on hard, multi-step tasks, but the answer takes longer and costs more because those extra steps use tokens too.
Beyond text, the family includes models for live voice conversations, image generation (branded Nano Banana), video generation (Veo), text-to-speech, transcription, and embeddings. An embedding turns a piece of text into a list of numbers that captures its meaning, so your app can find "similar" documents even when they share no exact words. Search and recommendation features are usually built on embeddings.
The two developer figures look like they disagree, and it is worth being clear about why. Google's 13 million counts everyone who has ever built something with its generative models, including one-off experiments. The 2.4 million figure is a third-party estimate of active developers in a given period. They measure different things, and the smaller number comes from an outside source rather than Google, so treat it as a rough indicator rather than a hard count.
Four ways to put Gemini inside your app
Google offers several paths to the same models. Picking the right one early saves a painful migration later. Here is how each of these Google AI tools fits a different kind of project.
1. The developer API with a key from Google AI Studio. AI Studio is a browser workspace where you type prompts, try different models, upload files and see results straight away. When something works, it hands you the matching code. You then call the Gemini API from your own server with an API key. This is the fastest way to start, and there is a free tier for testing. Requests go to a global pool, which means Google may process them in any of its regions.
2. The Agent Platform on Google Cloud (formerly Vertex AI). Same models, wrapped in enterprise controls. You get data residency options (choosing which country or region processes your data), Google Cloud's access management, audit logs and billing through your existing cloud account. Larger companies, and anyone in finance, health or government work, usually end up here.
3. Firebase AI Logic. Built for mobile and web apps that want to call Gemini directly from the app without building and running a backend server. Firebase manages the connection so your API key is not sitting inside the app. From November 2, 2026, Firebase will require App Check to use AI Logic. App Check is a service that confirms a request really came from your genuine app and not from someone who copied your configuration.
4. Gemini Nano on the device. On supported Android phones, a small Gemini model runs directly on the handset through a system service called AICore. You reach it through the ML Kit GenAI APIs, which offer ready-made tasks such as summarising, proofreading, rewriting, describing images and recognising speech, along with a lower-level Prompt API for custom instructions. Nothing leaves the phone and there is no charge per request. The trade-off is that the model is much smaller, the task list is narrower, and only newer phones support it.
Pro tip: Never ship a raw API key inside a mobile app or a web page. Anyone can pull it out of the app package or the browser's network tab and run up your bill. Call the model from your own server, use Firebase AI Logic with App Check turned on, or, for live voice features, issue short-lived "ephemeral" tokens that expire after a single session.
What actually changes for people building apps
Features that used to need a specialist team
A few years ago, reading data off a crumpled receipt meant training a dedicated image model on thousands of labelled examples. Summarising sales calls needed a speech-to-text system plus a separate text model. Tagging videos needed a third system again. Each of those was a project with its own data, its own specialists and its own maintenance.
With Gemini, one model accepts text, images, audio, video and PDFs in the same request. A receipt-scanning feature can now be a photo, a clear instruction, and a structured output request. Structured output means you give the model a template (called a schema) that says "return the merchant name as text, the total as a number, and the date in this format", and the model fills it in. Your app gets tidy data it can save to a database, instead of a paragraph it has to pick apart.
This is the real shift for small teams. The skill that matters most is no longer training models. It is describing the task clearly, testing it against messy real examples, and building sensible handling around the answers.
Your app can take actions, not only chat
Function calling lets you tell the model about actions your app can perform, such as check_order_status or book_appointment, with a short description of each. When a user asks "where is my parcel?", the model does not guess. It replies with a request to run check_order_status for the order number it found in the message. Your code runs the lookup and passes the result back, and the model writes the final answer using real data.
On top of that, the platform includes built-in tools. The model can search Google to ground its answer in current information, look up places through Google Maps, run code to do calculations, read a web page you point it at, and search through files you have uploaded. Google also now offers managed agents, where the Antigravity agent (running on Gemini 3.8 Flash by default) can carry out longer multi-step jobs in a remote environment.
The prototype became the easy part
With AI Studio's Build mode and the other Google AI tools, a working demo can come together in an afternoon. That changes where the effort goes. The demo is cheap. The distance between a demo that impresses in a meeting and a feature that behaves well for thousands of strangers is where the real work sits, and it is the subject of the rest of this guide.
Key takeaway
Gemini lowers the cost of trying an AI feature almost to zero. It does not lower the cost of making that feature reliable. Budget your time for testing with real, messy inputs, not for the first working version.
The hard parts that never show up in the demo
1. Data gaps: when the model does not have what it needs
Every generative AI model is trained on data up to a certain point in time and knows nothing about your company unless you tell it. When information is missing, the model does not stop and say so by default. It tends to produce the most likely-sounding answer, which is how you end up with a confident, wrong return policy in your support bot.
There are three practical ways to close the gap. First, give the model the facts it needs in the request itself, such as the customer's order history or the relevant help article. Second, use the File Search tool or your own retrieval system so the model pulls from your documents instead of its memory. Third, switch on grounding with Google Search for questions about current events, prices or anything that changes often. Search grounding is billed separately on Gemini 3 models, so check the pricing page before turning it on everywhere.
The subtler fix is in your schema. If a field might be missing, make it optional and add an explicit "not found" value. Tell the model that leaving a field empty is acceptable and guessing is not. An invoice extractor that is allowed to return "due date: not found" is far more useful than one that invents a plausible date.
2. Conflicting signals: when inputs disagree
Multimodal apps run into a problem text-only apps rarely had. A user uploads a photo of a damaged blue chair and writes "the red chair arrived broken." Your database says the order contained a green chair. Search results say the product was discontinued. Which one should the model believe?
Left alone, the model will pick one without telling you. The fix is to decide the order of trust yourself and write it into the system instruction, the standing set of rules you send with every request. For example: order records from our database come first, the customer's photo second, the customer's text third, web results last. Then ask the model to report any conflict in a dedicated output field rather than silently resolving it.
The final call on anything with money or safety attached should stay in your own code. Let the model flag "photo and order record disagree on colour", and let a simple rule decide whether that goes to a human agent. Models are good at spotting inconsistency. Your business rules should decide what to do about it.
3. Real-time decisions: when speed matters more than depth
A voice assistant that pauses for four seconds feels broken. A fraud check that holds up a payment for eight seconds loses the sale. Gemini gives you a few levers for these moments.
▪ Pick the lightest thinking level that works. Google suggests the low level for latency-critical work like real-time chat and incident alerts. Note that Gemini 3.8 Flash does not accept the "minimal" level and returns an error if you send it.
▪ Stream the answer. Streaming sends words to the screen as they are generated, so the user sees a response starting within moments instead of waiting for the whole thing.
▪ Use the Live API for voice. It keeps a continuous two-way audio connection open, so the model can listen and speak without the stop-start rhythm of separate requests.
▪ Use a smaller model for the first pass. Flash-Lite can make a quick yes-or-no call, and only the harder cases get passed to Flash.
▪ Move simple tasks onto the phone. Gemini Nano has no network round trip at all, which makes it a good fit for things like smart replies or proofreading as someone types.
Always set a timeout and decide in advance what the app does when it runs out. For a fraud check, that might mean falling back to your existing rules engine. For a chat reply, it might mean showing "still working on it" and continuing in the background.
4. Exceptions and edge cases
The problems that cause support tickets are rarely the obvious ones. These are the ones teams run into most:
▪ Retired models. Model names are shut down on fixed dates. Gemini 2.0 Flash went offline on June 1, 2026, and the 2.5 family goes on October 16 on the developer API (October 20 on the Agent Platform). Keep the model name in a config setting, not scattered through your code, and check the deprecations page every month.
▪ Removed settings. The migration guide for Gemini 3.8 Flash tells developers to strip out the temperature, top_p and top_k settings and to replace the old thinking budget with the new thinking level. Code written for older models may need changes before it works on newer ones.
▪ Safety filters on legitimate content. A pharmacy app discussing dosages or a legal app summarising a violent crime can trip content filters. Test these cases on purpose and adjust the safety settings where your use case allows it.
▪ Malformed tool calls. Occasionally the model returns a function call your code cannot parse. Validate every call against your expected format before running it, and ask the model to retry when the check fails.
▪ On-device surprises. Gemini Nano may not be downloaded yet on a phone that supports it, and several apps can share the one on-device model, so your request may wait in a queue. Always write a cloud fallback path for the moment the local model is not ready.
▪ Language and format drift. A prompt tested only in English may behave differently with Hindi, Tamil or mixed-language input. Test with the actual languages your users write in.
Watch out: Preview and experimental models are useful for testing new features, but they come with tighter rate limits and can change or disappear quickly. Put stable, generally available models behind anything customers depend on.
5. Behaviour under pressure and at scale
This is where many launches stumble. Google measures usage on three main counts: requests per minute, input tokens per minute, and requests per day. Crossing any one of them triggers an error, even if you are well under the other two. Two details surprise people. Limits apply per project, not per API key, so creating extra keys does not buy extra capacity. And the daily request count resets at midnight Pacific time, which works out to early afternoon in India.
There is also a spending limit measured over a rolling ten-minute window. According to Google's rate limits page, a project on Tier 1 can spend $10 in any ten minutes, Tier 2 can spend $50, and Tier 3 can spend $200. Projects move up tiers automatically as they spend more over time. When you hit any of these ceilings, the Gemini API returns a "429 RESOURCE_EXHAUSTED" error.
A well-built app treats that error as normal weather, not a disaster:
▪ Retry with growing gaps. Wait one second, then two, then four, with a small random offset so thousands of devices do not all retry at the same instant.
▪ Send non-urgent work to the Batch API. Batch jobs have their own separate limits, run at a lower price, and are ideal for overnight tasks such as summarising the day's tickets.
▪ Cache repeated context. If every request includes the same 50-page product manual, context caching stores it once so you do not pay full price to resend it each time.
▪ Pick the right service level. Google offers flex inference (cheaper, slower when the system is busy) and priority inference (faster and more dependable at peak times).
▪ Watch token growth. Google says openly that Gemini 3.8 Flash can use more tokens on long, complex tasks by design, because it checks its own work. A feature that looked cheap in testing can cost noticeably more on hard real-world inputs.
What a busy feature might cost
Here is a rough worked example. Say a support assistant handles 50,000 conversations a day, and each one uses about 2,000 input tokens and 400 output tokens. That adds up to 100 million input tokens and 20 million output tokens daily.
These figures use Google's published Gemini 3.8 Flash prices: $0.75 per million input tokens and $3.75 per million output tokens during the introductory period, rising to $1.50 and $7.50. Reasoning steps generally count toward output, so heavier thinking levels push the output line up. The bigger lesson is in the second row. If your financial model was built on this year's introductory rate, the cost doubles on New Year's Day unless you change models, trim prompts or cache more.
Cloud Gemini, Gemini Nano or an open model?
Gemini in the cloud is not the only choice, even within Google's own family. Google also releases Gemma, a set of open models you can download and run on your own servers. Gemini Nano, according to Google's Android team, shares its architecture with Gemma 4 but is tuned specifically for phones. Each option makes a different trade between power, privacy, cost and effort.
Plenty of good apps mix all three. A notes app might proofread on the phone with Nano, summarise long meetings in the cloud with Flash, and keep a self-hosted Gemma model for a hospital client whose data cannot leave the building. Choosing a generative AI model is less about finding the single best one and more about matching each task to where it should run.
A pre-launch checklist
Before you ship any Gemini-powered feature to real users, run through these points:
1. Keep the model name in configuration and put the shutdown date for that model in your team calendar.
2. Keep keys on the server, or use Firebase AI Logic with App Check switched on.
3. Use structured output with a schema that allows "not found" values.
4. Write down the order of trust for conflicting inputs in your system instruction.
5. Set a timeout and a fallback for every call to the Gemini API, including a plan for the 429 error.
6. Test with at least a hundred real, messy examples in every language your users write in.
7. Model your costs at standard pricing, not introductory pricing, and at three times your expected traffic.
8. Log inputs, outputs and token counts (with personal data removed) so you can find out why something went wrong.
Who should build on Gemini, and who should wait
Gemini is a strong fit if your app deals with mixed media, such as photos, voice notes, PDFs and video, because handling all of them in one model saves a great deal of integration work. It also suits teams already on Google Cloud or Firebase, Android-first products that can use Nano for offline features, and products that benefit from long inputs, such as contract review or codebase analysis.
It is a weaker fit if your contracts forbid sending data to any outside provider (look at Gemma or Nano instead), if your team cannot commit to regular upgrades as models are retired, or if your feature needs identical output every single time. Generative models vary a little from one answer to the next, and a regulated calculation is better handled by ordinary code with the model only explaining the result.
For many startups, the honest answer is to start with Google Gemini AI on the developer API, keep your code independent of any single model name, and move to the Agent Platform when customers start asking where their data is processed.
Conclusion
Gemini has changed what a small team can build. A two-person startup can now ship features such as document reading, voice conversation and image understanding that once needed a machine learning department. The models are capable, the free tier makes experiments cheap, and the on-device option opens up private, offline features on Android.
What it has not changed is the need for careful engineering. Models get retired on fixed dates. Prices change. Inputs go missing or contradict each other, users expect instant answers, and traffic spikes arrive without warning. The teams that do well treat Gemini like any other outside service: they plan for its failures, watch its costs, and keep the final say on important decisions in their own code. With the range of Google AI tools now available, from AI Studio for quick tests to Nano for on-device work, the tools are rarely the limiting factor. Planning usually is.


