Here is a workflow that plays out in a lot of small AI teams. A developer trains a model that sorts support tickets into categories. It works well inside a Jupyter notebook. Then the head of support asks to try it, and the developer realizes the only way to share it is to send a notebook file, explain how to install Python, and hope for the best. A week later the model is still sitting on one laptop.
The model was never the hard part. The hard part was the screen around it: the text box, the upload button, the results panel, and a link that someone else can open.
Gradio exists to close that gap. It is a free, open-source Python library that wraps any Python function in a web page. You write the logic, tell Gradio what goes in and what comes out, and it builds the page for you. This article works as a practical guide and a hands-on Gradio tutorial at the same time. It covers how the library works, where it fits, and where it starts to strain, for founders, developers, and office teams who want to put AI in front of real people without hiring a front-end team first.
What Gradio does, in one sentence
Gradio takes a Python function and gives it a web page. That is the whole idea.
A function, for readers who don't write code, is a small piece of a program that takes something in and gives something back. A spam filter function takes an email and returns "spam" or "not spam". An image function might take a photo and return a caption. Gradio looks at what your function expects and what it returns, then draws matching boxes, buttons, and panels in a browser.
Before tools like this existed, a basic machine learning UI meant writing a separate web server, building HTML forms, handling file uploads, and wiring it all together with JavaScript. A data scientist could easily spend more time on that wrapper than on the model. Gradio removes most of that work, so the person who built the model can also build the screen.
You work with three main building blocks, called Interface, Blocks, and ChatInterface, which are compared later in this article. Each piece on the screen, such as a text box or an image uploader, is called a component. Your app is simply a set of components connected to your code.
The numbers behind Gradio
Usage figures for Gradio come almost entirely from Hugging Face, which owns the project. Nobody independent audits them, so treat them as the company's own reporting.
▪ In April 2025, Gradio co-founder Abubakar Abid wrote on the Hugging Face blog that more than 1 million developers use Gradio every month.
▪ In October 2024, coverage of the Gradio 5 launch by ADTmag reported more than 2 million monthly users and over 470,000 applications built with it.
▪ Hugging Face job listings published in 2026 describe its platform as hosting 1.5 million Gradio apps.
▪ The main GitHub repository had about 43,200 stars and 3,600 forks as of July 29, 2026, according to star-history.com.
You'll notice the figures don't line up. The 2 million number came earlier than the 1 million number, which makes no sense if both measure the same thing. The likely explanation is that "monthly users" and "monthly developers" count different groups, but neither source defines its method, so we can't confirm that. If you quote one of these figures in a pitch deck, name the source and the date next to it.
The April 2025 post also lists some well-known open-source projects built on Gradio, including the AUTOMATIC1111 Stable Diffusion web interface, Fooocus, Oobabooga's text generation web UI, and LLaMA-Factory. Many people have used a Gradio app without knowing it, which says a lot about how far it has spread as an AI demo builder in the open-source world.
Your first app in five minutes
You need Python installed on your computer. Open a terminal (the text window where you type commands) and run:
Now create a file called app.py and paste this in:
Run it with python app.py. Gradio prints a local address, usually http://127.0.0.1:7860, and you can open it in your browser. "Local" means the page runs on your own machine and only you can see it for now. That page is your first working Gradio interface.
Here is what each part does:
▪ def greet(...) is your function. It takes a name and a number and returns a greeting with that many exclamation marks.
▪ fn=greet tells Gradio which function to run.
▪ inputs=[...] lists one input per function argument, in the same order. "text" becomes a text box, and the slider becomes a draggable bar from 1 to 5.
▪ outputs="text" shows the result in a text box.
▪ demo.launch() starts a small web server and opens the page.
That first step of this Gradio tutorial is a toy, but swapping in a real model takes only a few more lines. The example below uses the Hugging Face transformers library to score the mood of a sentence:
The "label" output is a component built for classification results. It shows the predicted category with a confidence bar, which is far easier for a manager to read than a raw number.
Interface, Blocks, or ChatInterface?
Most people start with Interface and move to Blocks once they need more control. The table below shows when each one makes sense.
Here is a short Blocks example. It lays out a two-column meeting notes tool with a button. The summarize function here just trims the text, so swap in your own model or API call.
The important line is btn.click(...). In Blocks, things happen through events. An event is something the user does, such as clicking a button, changing a text box, or uploading a file. You connect each event to a function and tell Gradio which components feed it and which ones show the result. Once that idea clicks, the rest of the library reads the same way, and the next part of any Gradio tutorial becomes much easier to follow.
A chat app is even shorter:
Your function receives the new message and the conversation so far. In a real app, you would pass both to a language model and return its answer.
The components you'll use most
These components cover most business and prototype work.
Getting your app in front of other people
Gradio gives you several ways to share an app with other people.
Temporary share links
Change the last line to demo.launch(share=True) and Gradio prints a public web address ending in gradio.live. Anyone with the link can use the app while your computer keeps running it. Gradio's documentation says its share servers only pass traffic through to your machine and do not store the data sent through your app.
There is a detail worth checking here. Gradio's current sharing guide says share links expire after one week. Older versions of the same guide said 72 hours. The limit has clearly changed over time, so check the documentation for the version you have installed before you promise a client that a link will still work on Friday.
Permanent hosting on Hugging Face Spaces
For a link that lasts, Gradio points people to Hugging Face Spaces, a hosting service for AI apps. You upload app.py and a list of required libraries, and Spaces runs the app for you. Running gradio deploy from your project folder walks you through it. Free hardware is enough for small models and API-based apps, but Spaces on free hardware go to sleep after a period of inactivity, so the first visitor after a quiet spell will wait while it wakes up.
Embedding, APIs, and AI agents
A hosted Gradio app can also live inside an existing website through a small HTML tag Gradio provides. More useful for developers, every Gradio app automatically gets an API. Look for the "Use via API" link in the page footer. It lists the endpoints and shows how to call them from Python (with the gradio_client package) or JavaScript (with @gradio/client). An API, short for application programming interface, is a way for one program to talk to another. In practice, the demo your sales team clicks through can also be called by your backend code.
Recent versions add one more option. Install the extra package with pip install "gradio[mcp]", launch with demo.launch(mcp_server=True), and Gradio exposes your functions as tools that AI assistants can call through the Model Context Protocol (MCP), an open standard for connecting assistants to outside tools. Gradio reads your function's docstring, the short description written at the top of a function, to tell the assistant what each tool does. The Hugging Face MCP course recommends including an "Args:" section that describes each parameter.
How teams outside engineering use it
Gradio grew up among researchers, but plenty of its users today are product, sales, and operations teams. Here are patterns that come up often.
▪ Client pitches. An agency shows a prospect a working version of the idea on a call, using the prospect's own sample data, instead of a slide with mockups. As an AI demo builder, this is where Gradio pays off fastest.
▪ Internal tools. A finance team gets a page where they drop in an invoice PDF and get the vendor name, date, and total back in a table. Nobody needs to learn the underlying model.
▪ Side-by-side model checks. Two models answer the same question in two panels, so a manager can judge quality before the company commits to one vendor.
▪ Feedback collection. Gradio's flagging feature lets testers mark bad outputs, which are saved for the team to review later.
▪ Content experiments. Writers test prompts for product descriptions or email subject lines through a simple web page without touching code.
What happens under the hood when traffic arrives
This is where the "in minutes" promise meets real life. A demo that works for you alone can behave very differently when 200 people open it after a LinkedIn post.
The queue
Every Gradio app has a built-in queue, a waiting line for requests. When someone clicks Submit, their request joins the line, and Gradio sends the result back using server-sent events, a method where the server keeps a connection open and pushes updates. Gradio's performance guide explains why this matters: an ordinary web request often times out in the browser after about a minute, while this approach doesn't, and it lets the page show each person a live estimate of their wait.
Why your app may feel slow even on a big server
Gradio's server has a pool of 40 worker threads by default, a number inherited from FastAPI, the web framework it is built on. That does not mean 40 people get served at once. By default, Gradio lets only one worker run any given event at a time. The setting is called default_concurrency_limit, and it starts at 1. If 30 people click the same button, they are served one after another.
Gradio picked that default on purpose. If your app runs a large model on a GPU, letting many requests hit it at once can run the machine out of memory and crash everything. If your function only calls an outside API such as OpenAI or Claude, the limit is far too cautious, and you can raise it a lot.
Two of these deserve a closer look. The max_size setting sounds like it would hurt users, and Gradio's own guide calls the effect a paradox: capping the line often makes the experience better, because people find out right away that the app is busy instead of watching a timer that never moves. Batching is often the bigger win for deep learning models. The guide says it is frequently faster than running workers in parallel, because a GPU can process four images in one pass almost as quickly as one.
Hardware matters too. The same guide says moving a deep learning model from CPU to GPU usually makes inference 10 to 50 times faster. If you upgrade, revisit your concurrency settings, because GPU memory is a separate and often smaller pool than regular memory.
Data gaps: when inputs are missing or messy
Real users leave fields empty, upload the wrong file type, and paste 40 pages into a box meant for a paragraph. Gradio passes whatever arrives straight to your function, so your function has to cope.
The most common gap is an empty input. If someone clicks Submit without uploading an image, your function receives None, Python's word for "nothing here". Without a check, the code crashes and the user sees a vague error. A cleaner pattern looks like this:
Raising gr.Error shows a friendly pop-up with your message, and the app keeps running for everyone else. Gradio also offers gr.Warning and gr.Info for softer messages that don't stop the function.
Other habits that help:
▪ Set a size limit on uploads with the max_file_size option in launch(), so one person can't fill your disk with a 4 GB video.
▪ Choose the input format your code expects. The Image component, for example, can hand your function a NumPy array, a PIL image, or a file path.
▪ Add examples with gr.Examples. Clickable sample inputs show people what good input looks like, which cuts down on bad input in the first place.
▪ Show uncertainty instead of hiding it. A Label output can show the top three categories with their scores, so a user can see when the model is torn between two answers.
Conflicting signals and real-time decisions
Interactive apps get messy signals from users. Someone clicks Submit three times because nothing seemed to happen. Someone edits the input while the model is still working on the old version. Two buttons update the same output panel. Gradio gives you tools for each case.
▪ Each event has a trigger_mode setting, so you can tell Gradio to ignore new clicks while a job is running or to keep only the latest request.
▪ The cancels option lets a Stop button end a running event, which matters for long text generation.
▪ If step two depends on step one, chain them with .then() so they run in order instead of racing each other.
▪ If two models give different answers, show both and let a person decide. A confidence threshold that sends unclear cases to a human reviewer is often worth more than a clever tie-breaker.
For real-time behavior, the most useful feature is streaming. If your function uses yield instead of return, Gradio updates the screen every time a new piece arrives. That is how chat apps display a language model's answer word by word instead of making people stare at a blank box for 15 seconds. Setting live=True on an Interface reruns the function whenever an input changes, which suits fast functions but can overwhelm slow ones. The gr.Timer component runs a function on a schedule, which helps for small dashboards that refresh every few seconds. Audio and image inputs can also stream from a microphone or webcam for apps that react as you speak or move. For long jobs, adding a gr.Progress argument to your function draws a progress bar so users know the app hasn't frozen.
Exceptions and edge cases worth testing before launch
These are the issues that tend to show up a day after the demo goes out.
▪ Shared global variables. A variable defined outside your functions is shared by every visitor. If you store one user's upload there, the next user may see it. Use gr.State for per-user data.
▪ Page refreshes. gr.State lives only as long as the browser tab. If users need settings to survive a refresh, recent versions offer gr.BrowserState, which stores values in the visitor's browser.
▪ Hidden API routes. Hiding a button doesn't stop someone from calling its function through the automatic API. Gradio 6 has settings to control which events are exposed, so review them for anything sensitive.
▪ Version upgrades. Gradio 6 made breaking changes. App-wide options such as the theme moved to launch(), the old tuple format for Chatbot history was dropped in favor of messages, and the show_api option was replaced by footer settings. The official migration guide suggests upgrading to version 5.50 first, since it prints a warning for everything that will break in 6. Pin your version in requirements.txt so a fresh install doesn't surprise you.
▪ Long-running jobs. Gradio's connection doesn't time out, but your hosting provider's proxy might. Test a worst-case input on the real host as well as on your laptop.
▪ Phones. A layout that looks fine on a laptop can stack awkwardly on a phone. Open your Gradio interface on a mobile browser before sharing it widely.
Security basics for teams
A shared link is a public door, so a few basics matter.
Before releasing Gradio 5 in October 2024, Hugging Face hired the security firm Trail of Bits to audit the code. The published findings included server settings that could let attackers steal access tokens, file uploads that could host malicious scripts, and components that could leak files from the server in simple setups. Hugging Face says all the identified issues were fixed and checked by Trail of Bits before Gradio 5.0 shipped. New advisories still appear on Gradio's GitHub page, including one about server-side request forgery through the gr.load() function. Server-side request forgery is a trick that makes your server fetch addresses an attacker chooses. The practical lesson is to keep Gradio updated.
On your side:
▪ Add a login. demo.launch(auth=("admin", "a-strong-password")) puts a username and password in front of the app. On Hugging Face Spaces, you can also add "Sign in with Hugging Face".
▪ Keep secrets out of code. Store API keys in environment variables or in your Space's secret settings, never in app.py.
▪ Treat share links as public. Anyone who gets the link can use your Gradio interface and anything it can reach on your machine.
How Gradio compares with the alternatives
Gradio is not the only way to put a web page on a Python script. Here is how it stacks up against the options teams usually weigh.
The short version: for a model demo, an AI assistant, or a quick internal tool, Gradio is usually the fastest route. For a data dashboard with lots of charts and filters, Streamlit or Dash may feel more natural. For a customer-facing product with your own brand, accounts, and payments, you'll probably end up with a custom front end. Many teams prototype in Gradio and rebuild later, which is a sensible path. As an AI demo builder, it earns its place by helping you learn what users want before you spend money on design.
When Gradio is the wrong choice
Gradio is generous with what it gives you for free, but it has limits.
▪ Themes and CSS go a long way, but a brand-heavy product needs more control over every screen than Gradio offers.
▪ Accounts, roles, billing, and multi-page flows are possible (Gradio 6 added a Navbar component for multi-page apps), yet a standard web framework handles them more cleanly.
▪ Contracts that guarantee response times need careful engineering around the single-process queue and in-memory sessions described earlier.
▪ Gradio assumes someone on the team can write and maintain Python code.
Hugging Face has also introduced a server mode that keeps Gradio's backend, including the queue, streaming, and MCP support, while letting developers write a completely custom front end. It narrows the gap, though it gives up the main reason people pick Gradio in the first place. A dedicated machine learning UI written by front-end developers still makes sense once a product has paying customers and a clear design.
Conclusion
The gap between "the model works on my laptop" and "the team can use it" used to take weeks. With Gradio, it often takes an afternoon. That changes how teams make decisions. A founder can test an idea with real customers before raising money for it. An operations lead can try an AI tool on real paperwork before buying a license.
The work that remains is the work this article focused on: handling empty and messy inputs, setting the queue for real traffic, protecting shared links, and knowing when a demo has turned into a product. If you take one thing from this Gradio tutorial, make it this: build the first version fast, put it in front of people, and let their reactions tell you what to build next.


