Stability AI Explained: Open-Source Generative Models

Stability AI Explained: Open-Source Generative Models

The most downloaded version of Stable Diffusion in 2026 isn't the newest one. It's a model from October 2022. A July 2026 analysis by HowAIWorks found that SD 1.5 and SDXL still pull the most downloads on Hugging Face, while the current flagship, SD 3.5 Large, gets under 2% of SD 1.5's numbers.

That says a lot about Stability AI. The company made its name by giving away the files behind an image generator, and people built so much on those early files that the newer models are still catching up. Among generative AI tools, that history makes Stability's products unusual. You can download them, run them on your own machine, change them, and build a business on them, within limits that many people misread.

If you're a founder weighing an image feature for your app, a developer choosing a model, or a marketing team tired of paying per image, those limits matter. They decide what you'll pay, what you're allowed to do, and where things will go wrong. This guide walks through what the company makes, how Stability AI models work in plain language, what "open" means under its license, and how the models behave once real traffic hits them.

What Stability AI actually is

Stability AI is a London-based company founded in 2019. Its big moment came in August 2022, when it released Stable Diffusion together with the CompVis research group at LMU Munich and the startup Runway. Anyone could download the model and run it on a home computer with a decent graphics card. At the time, the other well-known image generators, DALL-E and Midjourney, only ran on their makers' servers. That release moved image-focused open-source AI from research labs to ordinary laptops.

The years after that were rough. Founder Emad Mostaque stepped down as CEO in March 2024 amid reports of heavy losses. Prem Akkaraju, formerly the CEO of the visual effects studio Weta Digital, took over later that year and steered the business toward paying customers in film, music, games and advertising.

You can see that shift in the recent news. Stability signed co-development deals with Electronic Arts, Universal Music Group and Warner Music Group, launched an enterprise platform called Brand Studio in April 2026, and released Stable Audio 3.0 in May 2026. On August 25, 2026, TechCrunch reported a $76 million Series B that brought total funding to $232 million. The investors included Universal Music Group, Sony Music, Warner Music Group, EA and AMD Ventures.

So today the company sells two things at once: model files that anyone can download, and paid services built on those same models.

MARKET AND COMPANY NUMBERS AT A GLANCE

AI image generator market

Projected to reach USD 60.8 billion by 2030, growing about 38.2% a year (MarketsandMarkets, 2025 forecast).

Stability AI total funding

$232 million after the August 2026 Series B (TechCrunch, August 25, 2026). Some company databases still show older totals, such as the $181 million figure attributed to Tracxn in a June 2026 article, because they haven't caught up with the new round.

Downloads

Over 45 million on Hugging Face and about 214 million more across community sites like Civitai (Fast.io, June 2026). Treat these as rough, since counters include repeat downloads and mirrors.

Revenue

Stability doesn't publish it. One third-party estimate puts the 2026 run rate near $190 million (Fueler, July 2026), with no official number to check it against.

How these models make a picture

You don't need the math, but a rough mental picture helps when an image comes out wrong.

Think of a photo buried under TV static. Now think of someone who has studied millions of photos, each shown with more and more static added until nothing was left but noise. After enough practice, that person can look at a fuzzy mess and guess what it looked like one step cleaner. Repeat that guess 20 or 30 times and you get from pure static to a clear picture.

That's what a diffusion model does. It starts from random noise and removes a little at each step, steered by your text prompt. Three parts share the work:

▪       A text encoder reads your prompt and turns it into numbers the model can use. Older versions used one called CLIP. SD 3.5 adds a larger one called T5, which copes much better with long, detailed sentences.

▪       The denoiser guesses what each cleaner step should look like. In SD 1.5 and SDXL this part is a design called a U-Net. In the 3.5 family it's a transformer, the same broad type of design behind chatbots, which Stability calls MMDiT.

▪       A compressor called a VAE shrinks images into a small "latent" version so the model can work faster, then expands the result back to full size at the end. That's where the phrase "latent diffusion" comes from.

Three settings show up in almost every app built on these models. Steps control how many cleanup passes the model makes; more steps means slower and usually more detail, up to a point. Guidance scale (often labeled CFG) controls how strictly the model follows your prompt. Set it too low and the image wanders off topic. Set it too high and colors turn harsh and overcooked. The seed is the number that sets the starting noise. Keep the seed, prompt and settings the same and you'll usually get the same image back, which is what makes repeatable work possible.

The model family, one at a time

When someone says Stable Diffusion, they might mean any of several models. They differ in size, quality, hardware needs and license, so it pays to take them one by one.

SD 1.5 (2022)

This is the old workhorse. It has roughly a billion parameters (the adjustable numbers a model learns in training), works at 512 by 512 pixels, and runs on graphics cards with as little as 4GB of video memory, according to Fast.io's June 2026 review. Its faces and hands look dated now. What keeps it alive is the add-on library: thousands of community styles, character models and pose tools that, for niche uses, beat raw quality.

SDXL (2023)

SDXL moved up to 1024 by 1024 pixels with much better composition and lighting. It needs about 8GB of video memory. Several 2026 guides still call it the safest all-round pick, mainly because its add-ons are plentiful and its quirks are well known.

The SD 3.5 family (October 2024)

The current flagship line comes in three versions. SD 3.5 Large has 8.1 billion parameters, the best prompt-following in the family, and output up to about one megapixel; plan on 16GB or more of video memory. SD 3.5 Large Turbo is a "distilled" copy of Large. Distillation trains a model to reproduce a bigger model's results in fewer steps, and here it cuts the work to about four steps instead of thirty or so. SD 3.5 Medium has 2.5 billion parameters and needs about 9.9GB of video memory before counting the text encoders, per Stability's release post. Oddly, Medium supports a wider resolution range (0.25 to 2 megapixels) than Large.

The 3.5 models handle words inside images far better than earlier versions, which matters if you want a poster, a menu or a mock product label. Their weak spot is the add-on library. Far fewer community fine-tunes exist for 3.5 than for SDXL.

Beyond images

Stability AI models now cover more than pictures. Stable Audio 3.0, released in May 2026, is a family of music and sound models that Stability says was trained on fully licensed audio. Smaller versions are downloadable from Hugging Face, and a larger one runs through the API. Stable Audio Open Small, built with the chip designer Arm, is compact enough to run on a phone with no internet connection. For video and 3D there's Stable Video 4D 2.0, Stable Virtual Camera (a research preview) and SPAR3D, which turns one photo into a 3D object. These lines are younger, and most businesses meet them through enterprise deals.

Pro tip: Before picking a model, list the add-ons your workflow depends on. If you rely on a pose-control tool or a community style that only exists for SDXL, the "better" 3.5 model may cost you weeks of rebuilding.

Is it really open source?

This question trips up more teams than any other, and the honest answer is "partly."

In the strict sense, open source means anyone can use, change and share the software for any purpose. Most of what Stability releases is better described as "open weights." You get the trained model files and can run them wherever you like, but the license adds conditions, and the full training data isn't published.

The terms change from model to model. SD 1.5 and SDXL use CreativeML OpenRAIL licenses, which allow commercial use with no revenue ceiling but list banned uses such as harassment or illegal content. SD 3.5 and newer releases use the Stability AI Community License. It's free for research, for non-commercial use, and for commercial use by people or companies earning under $1 million a year. Past $1 million, you need a paid Enterprise License.

Two details in that license surprise people. The $1 million counts your total revenue from every source, not just money made with the model, so a consulting firm earning $1.2 million from unrelated work needs an enterprise license even for a small internal image tool. And Stability states that the license can be revoked if you break its terms.

So does Stability belong to the open-source AI movement? In spirit, yes. No other company did more to put a serious image generator in ordinary hands. In legal terms, a lawyer would call the newer models source-available with conditions. For planning purposes, the lawyer's view is the one that counts.

Key takeaways

✓   "Open" here mostly means you can download the model and run it yourself.

✓   SD 1.5 and SDXL have no revenue cap. The 3.5 models do.

✓   The $1 million line counts all company revenue, not only revenue from AI features.

✓   Research use stays free no matter how large your organization is.

Comparing the main image models

Here's how the five image models most teams choose between stack up. "Memory" means video memory on the graphics card, and the figures are typical minimums for comfortable local use. "Cap" is the revenue limit for free commercial use.

Model

Year

Size

Memory

License

Cap

Best suited to

SD 1.5

2022

About 1 billion

4GB

OpenRAIL-M

None

Cheap hardware, niche community styles

SDXL

2023

About 3.5 billion

8GB

OpenRAIL++-M

None

General work with mature add-ons

SD 3.5 Large

2024

8.1 billion

16GB or more

Community License

$1M

Best prompt-following, text in images

SD 3.5 Large Turbo

2024

8.1 billion

16GB or more

Community License

$1M

Fast previews in interactive apps

SD 3.5 Medium

2024

2.5 billion

About 10GB

Community License

$1M

Good quality on mid-range cards

The older models win on hardware and licensing; the 3.5 family wins on quality. That trade-off explains the download numbers at the top of this article.

Where the models run short on information

A diffusion model only knows what it saw in training. The early versions learned from LAION, a huge public collection of images and captions gathered from the web. The web has endless stock photos of American offices and golden retrievers. It has far fewer photos of a Kanchipuram silk saree on a loom, a specific industrial valve, or shop signs in Tamil script.

When the model hits one of these gaps, it doesn't warn you. It fills the hole with the closest average it knows. Ask for a traditional wedding from a region that was thin in the training data and you may get costumes from three cultures blended together, drawn with total confidence. Ask for your own product and you'll get a generic look-alike.

There are three practical ways to close those gaps:

▪       Train a LoRA. A LoRA is a small add-on file, often tens of megabytes, trained on your own pictures, sometimes as few as 20 to 50 good photos. It teaches the model a new product, face or style without retraining the whole thing.

▪       Start from a real image. Image-to-image mode begins from your photo rather than from pure noise, so the model edits what exists instead of inventing from scratch.

▪       Hand over the layout. ControlNet tools let you supply an outline, depth map or blurred sketch so the structure comes from you. Stability released Blur, Canny and Depth versions for SD 3.5 Large.

There's a second, quieter gap: what went into training. For SD 3.5, Stability describes the data as a mix of synthetic images and filtered public data, but it hasn't published a full list. If your legal team wants to trace where each training image came from, that answer doesn't exist. Stable Audio 3.0 is the exception, and it hints at where the company is heading for customers in music and film.

When the signals disagree

Conflicting signals show up at two levels: inside a single image request, and in the information you'll find about the company.

Inside one image request

Every image blends several inputs at once: your prompt, a negative prompt listing things to avoid, and possibly a reference image, a LoRA and a control map. When those inputs pull in different directions, the model doesn't choose one. It splits the difference, and the result is often worse than either option alone.

A few clashes come up again and again. Ask for "a minimalist empty desk" while using a LoRA trained on busy workspace photos, and you'll get a half-cluttered desk. Feed in a ControlNet pose of someone sitting with a prompt that says "running," and expect twisted limbs. Stack three or four LoRAs at full strength and each one drags the picture its own way; faces usually melt first. Dropping each LoRA's weight to around 0.6 or 0.7 tends to help.

Long prompts cause a sneakier clash in older versions. The CLIP encoder in SD 1.5 and SDXL reads 77 tokens at a time, which works out to roughly 50 to 60 words. Many apps quietly chop or split longer prompts, so a detail near the end may be ignored or weighted oddly. The T5 encoder in SD 3.5 handles much longer prompts, and that alone is a good reason to upgrade if your users write paragraphs.

Around the company

The outside information conflicts too, and you should know that before building a roadmap on it. In spring 2026, several review sites ran articles about a "Stable Diffusion 4" launch with 4096-pixel output. Yet a July 2026 model tracker reported that Stability still called SD 3.5 its most powerful image model, with no successor, and a tutorial updated the same month listed 3.5 as the latest official release. We couldn't find an SD4 announcement on Stability's own news pages while researching this piece. Plan around what you can download today.

The same goes for a figure you'll see everywhere: that models based on Stable Diffusion produce about 80% of all AI images. That number traces back to an Everypixel estimate from 2023, and 2026 articles still repeat it as current. With FLUX from Black Forest Labs (founded by former Stability researchers) and OpenAI's image models now widely used, the real share has almost certainly moved.

Pro tip: When you read a model comparison, check the version and settings behind it. A test run at 20 steps with default guidance can reverse its verdict at 40 steps with tuned guidance.

Making decisions in real time

When a user taps a button and waits, speed becomes a product decision.

A full SD 3.5 Large image at 28 to 40 steps can take several seconds on a strong GPU. In a chat-style interface, that feels slow. Teams usually handle it in one of four ways:

▪       Use Large Turbo for the first draft. Four steps make near-instant previews possible, and the full model can re-render only the image the user picks.

▪       Generate small, then upscale the keepers. Most users throw away most drafts, so there's no point paying full price for every one.

▪       Show the image forming. Displaying blurry in-between steps makes a six-second wait feel shorter.

▪       Use optimized builds. Stability and NVIDIA released TensorRT and FP8 versions of SD 3.5 that Stability says run twice as fast with 40% less memory on supported RTX cards. There are AMD-tuned versions too.

Every request needs a safety check on the prompt, the output, or both. Through Stability's API, that filtering happens on their side. When you self-host, it's yours to build, and you have to decide what happens when the filter fires. You can show an error, retry with a new seed, or pass the request to a human reviewer. An unlimited retry loop is a quick way to burn GPU money on a prompt that will never pass, so cap retries at two or three.

You also need a fallback. If your own server is full, do you make the user wait, switch to a smaller model, or route the request to the API? Settle that before launch, not during a 2 a.m. traffic spike.

Exceptions and edge cases worth testing

These models handle the average request well. The trouble lives in the corners, so test these before you ship:

▪       Hands and fingers. Better in 3.5 than in 1.5, still not fixed. Extra fingers show up most when hands are small or holding things.

▪       Words inside images. SD 3.5 can spell short phrases. Long sentences, fine print and non-Latin scripts still come out garbled much of the time. For anything that must be exact, like a price or a legal line, add the text afterward in a design tool.

▪       Unusual sizes. Each model learned at certain resolutions. Ask SD 1.5 for a very wide banner and you'll often get two copies of the subject side by side, because it's filling space it never learned how to fill.

▪       Crowds. Faces far from the camera tend to smear into the background.

▪       Watermarks and logos. Early versions of Stable Diffusion sometimes produced blurry stock-photo watermarks, since so many training images carried them. That behavior became part of Getty Images' UK case, covered below. Scan outputs for stray marks before any commercial use.

▪       False alarms from safety filters. Filters can block harmless requests such as medical diagrams, swimwear catalogs or classical paintings. If your business works in one of those areas, test the filter in week one.

▪       Same seed, slightly different image. A fixed seed usually reproduces a result, but switching GPU type, driver, precision or software library can shift the output a little. If you need exact repeatability for records or audits, save the image file itself, not just the settings.

How the system holds up under heavy load

One image on a laptop and 50,000 a day for customers are different jobs.

Memory is the first wall. SD 3.5 Large's weights alone take around 16GB at standard precision, before the text encoders. If a server accepts more requests at once than its video memory can hold, it doesn't slow down politely. It crashes with an out-of-memory error. The standard fix is a queue in front of the model, with a fixed number of workers per GPU.

Cold starts come next. Loading a large model into GPU memory can take tens of seconds. Serverless GPU platforms that shut down idle machines save money, but the first user after a quiet spell waits. One always-on warm machine is usually worth the cost.

Batching is a trade: four images in one pass cost less per image, but the user waits for the whole batch.

Style drifts at volume. Across thousands of images for one brand, small differences in color and mood add up. Teams that do this well lock the model version, LoRA version, sampler and settings, and they write all of it down. Changing one piece can shift a whole catalog's look.

Costs grow in a straight line on the hosted route. Per Fast.io's June 2026 review, Stable Image Ultra costs about 8 cents per image through Stability's API, and SD 3.5 Medium runs about 3.5 cents. At 50,000 images a day, choosing one over the other is a difference of more than $2,000 a day. Rate limits also apply, and at that scale you'll want a direct contract rather than standard tiers. Enterprises that already buy from Amazon or Microsoft can also reach SD 3.5 Large through Amazon Bedrock or Azure AI Foundry, which keeps billing and security reviews inside existing contracts.

One pressure point isn't technical at all. A startup that grows from $800,000 to $1.1 million in revenue crosses the Community License line, and from that point it needs an enterprise license for the 3.5 models. Put that milestone on the calendar before your finance team finds it the hard way. It's one of the few places where Stability AI models behave differently at scale for purely legal reasons.

Three ways to run it

Most teams choose between three routes, based on volume, privacy needs and engineering time.

Factor

Self-hosting on your own GPUs

Stability's own API

Cloud marketplaces (Bedrock, Azure)

Upfront cost

High: GPUs or rented servers

None

None

Cost per image

Falls as volume rises

Fixed per image, a few cents to about 8 cents

Fixed per image, set by the cloud provider

Data privacy

Prompts and images never leave your servers

Sent to Stability

Stays inside your cloud account

Control over models

Full: any version, LoRA or add-on

Limited to what the API offers

Limited to listed models

Setup effort

Weeks of engineering

Hours

Days, mostly paperwork and access setup

Safety filtering

You build it

Handled for you

Handled for you, plus provider controls

Best for

High volume, custom styles, strict privacy

Prototypes and moderate volume

Enterprises with existing cloud contracts

A common path is to start on the API, prove that users want the feature, then move to self-hosting once monthly image bills pass the cost of a couple of rented GPUs.

The legal picture in 2026

Getty Images sued Stability in the UK and the US in 2023, claiming its photos were used without permission to train Stable Diffusion. On November 4, 2025, the High Court in England ruled largely in Stability's favor. Getty had already dropped its main training claims during the trial, because the training didn't take place in the UK. The judge then found that the model's weights don't store copies of Getty's images, so the model isn't an "infringing copy" under UK law. Getty won only a narrow trademark point tied to watermarks produced by early versions.

That's not the final word. Hogan Lovells' litigation tracker lists the UK case as under appeal, and Getty's US lawsuit is still working its way through the courts. The UK judgment was also narrow and fact-specific.

What does this mean for you? The UK ruling lowered one risk, but it didn't remove others. For commercial work, review every output, avoid prompting with living artists' names or other companies' brands, and ask your lawyer whether you need contractual protection. Some enterprise deals include it. This section is a summary, not legal advice.

Who should use this, and who should look elsewhere

Startup founders building an image feature are a strong fit if they need control over the model, want customer data to stay private, or expect costs to drop as volume grows. The usual route is to start on the API and self-host later.

Developers get something closed products can't offer: the chance to see how generative AI tools work under the hood. Hugging Face's diffusers library and the node-based ComfyUI interface are the most common starting points, and both support every model in the table above.

Content writers and marketers usually shouldn't self-host. Hosted generative AI tools built on these models, or a plugin connected to Stability's API, cover blog headers, concept art and social posts. Avoid using them for pictures of real people or real news events.

Office teams can use them for internal mockups and slide visuals, after checking with IT.

Look elsewhere if you want the best photorealism straight out of the box with zero setup. Closed tools like Midjourney or OpenAI's image models may get you there faster. The same goes if you need proof that every training image was licensed; right now, only some of Stability's lines, such as Stable Audio 3.0, can say that.

A simple way to start

If you're still deciding, this sequence avoids most of the expensive mistakes:

Step 1.             Try the models in a hosted interface first. Stability's developer platform gives new users a small batch of free credits, enough to judge whether the quality fits your needs.

Step 2.             Write down 20 real prompts from your actual use case, not demo prompts, and run them through two models, such as SDXL and SD 3.5 Medium.

Step 3.             Check the license against your company's total revenue, including revenue expected over the next year.

Step 4.             Pick a route using the table above.

Step 5.             If the model can't draw your product or style, train a LoRA on 20 to 50 of your own images.

Step 6.             Lock your versions and settings in writing before you produce anything at volume.

Pro tip: Save the seed, model version and full settings with every image you publish. When a client asks for "the same thing, but blue" three months later, you'll be glad you did.

Final thoughts

Stability AI is two companies sharing one name: the scrappy lab that released Stable Diffusion in 2022 and let the internet build on it, and an enterprise vendor signing deals with record labels and game studios. Both shape what you get.

For most readers, the practical view is simple. The older models are free, flexible and supported by a huge community. The newer ones are better but come with a revenue cap and a thinner add-on library. Wherever the data is thin, the model will guess, so plan to fill the gaps with your own images. At scale, memory, cold starts and licensing will bite before image quality does.

Among today's generative AI tools, few give you this much control over where the model runs and what you do with it. With open-source AI image models, you're trading convenience for freedom, and knowing the terms of that trade is the whole game.

Nidhi Jain

Nidhi Jain

Nidhi is an exceptionally talented and creative content writer, bringing life to ideas through her words. With marketing knowledge and a deep understanding of various industries, she crafts captivating content that resonates with our audience. Her in-depth knowledge of trending tech and consumer affairs adds a unique perspective to her work, making it engaging and impactful.

Build Your Agile Team

We provide you with a top-performing extended team for all your development needs in any technology.

Hourly
$20
It Includes
Duration
Hourly Basis
Communication
Phone, Skype, Slack, Chat, Email
Hiring Period
25 Hours (MIN)
Project Trackers
Daily Reports, Basecamp, Jira, Redmime, etc
Methodology
Agile
Monthly
$2600
It Includes
Duration
160 Hours
Communication
Phone, Skype, Slack, Chat, Email
Hiring Period
1 Month
Project Trackers
Daily Reports, Basecamp, Jira, Redmime, etc
Methodology
Agile
Team
$13200
It Includes
Team Members
1 (PM), 1 (QA), 4 (Developers)
Communication
Phone, Skype, Slack, Chat, Email
Hiring Period
1 Month
Project Trackers
Daily Reports, Basecamp, Jira, Redmime, etc
Methodology
Agile

Frequently Asked Questions

Is Stable Diffusion free to use for my business?
It depends on the version and your revenue. SD 1.5 and SDXL can be used commercially with no revenue limit, as long as you follow the banned-use list in their licenses. SD 3.5 is free for businesses earning under $1 million a year in total revenue. Above that, you need an Enterprise License from Stability.
What computer do I need to run it myself?
The graphics card matters most, specifically its video memory. SD 1.5 runs on about 4GB, SDXL on about 8GB, and SD 3.5 Large wants 16GB or more. SD 3.5 Medium sits in the middle at around 10GB. If you don't have a card like that, the API or a hosted app is a better starting point.
How is Stability AI different from Midjourney or OpenAI's image tools?
Midjourney and OpenAI keep their models on their own servers, so you use them only through their apps or APIs. Stability AI models can be downloaded, run on your own hardware, and trained on your own images. That's the main reason developers treat it as the backbone of open-source AI image work.
Can I teach it to draw my own products?
Yes. The usual method is a LoRA, a small add-on file trained on 20 to 50 clear photos of your product from different angles. Training takes from under an hour to a few hours on a single good GPU, and many hosted services will do it for you. Expect a retrain or two.
Did the Getty lawsuit make it safe to use these images commercially?
Not entirely. The November 2025 UK ruling went largely in Stability's favor, but the case is under appeal and a separate US case is still open. The ruling also didn't decide whether AI training on copyrighted images is lawful in general. Review outputs carefully, avoid copying brands or living artists, and get legal advice for high-stakes commercial use.