Monitoring Microservices with Prometheus & Grafana

Monitoring Microservices with Prometheus & Grafana

A customer emails your support address at 9:40 on a Tuesday morning. Checkout keeps failing, she writes, and she has tried twice. Your team opens the admin panel and everything looks normal. The servers are running, the site loads, and nobody can say what's wrong, because nobody can see inside the part that broke.

That gap is what monitoring exists to close. In simple terms, it means collecting small measurements from your software all day long (how many requests arrived, how many failed, how long each one took) and keeping them so you can look back later. When it's set up well, the first warning comes from a graph instead of a customer.

This guide covers one of the most common ways to do it. Prometheus collects and stores the measurements, and Grafana turns them into dashboards people can read. You don't need an engineering background to follow along. Technical terms get a short explanation when they first appear, and the messy parts of real life get their own space: missing data, graphs that disagree, and what happens when a system grows faster than expected.

Why microservices are harder to watch

An older-style application is one big program. If it breaks, you look at that program. A microservices setup splits the work into many small programs that talk to each other over a network. One handles logins, another handles payments, another sends emails. A single click from a customer can pass through six of them.

That design has real benefits. A team can change one piece without touching the rest, and a busy piece can be given more capacity on its own. The price is that failures become harder to locate. The payment service might be slow only because the database behind the inventory service is slow. If you watch only the payment service, you'll spend an hour looking in the wrong place.

Guessing gets expensive. ITIC's 2024 Hourly Cost of Downtime Survey, which polled more than 1,000 firms worldwide, found that one hour of downtime costs over $300,000 for more than 90% of mid-size and large enterprises. ITIC also points out that small businesses carry the same risk, even though their dollar losses are a fraction of that.

Good system monitoring answers three questions. Is something broken right now? Where is it broken? Was it already getting worse before it broke? Without stored history, you can only answer the first one.

What Prometheus and Grafana do

Prometheus is a free, open-source tool that collects numbers from your software and stores each one with a timestamp. It began at SoundCloud in 2012, after the company found that its older metrics tools couldn't handle its container-based setup. It joined the Cloud Native Computing Foundation (CNCF) in 2016 and graduated in August 2018, the second project to do so after Kubernetes.

Grafana is a separate free tool that draws charts. Torkel Ödegaard released it in January 2014, and it started as a fork of Kibana. Grafana doesn't keep your measurements. It asks other systems for data each time a dashboard loads, and it can combine results from several sources on one chart.

The two work well together because each has a narrow job. Prometheus is the record keeper and Grafana is the window you look through. A third piece, Alertmanager, handles who gets notified and when. The table shows how they divide the work, which is the foundation of Prometheus monitoring on most teams.

Tool

Main job

What it holds

A question it helps answer

Prometheus

Collects and stores measurements, runs alert rules

Time-stamped numbers, kept 15 days by default

How many checkout requests failed in the last five minutes?

Grafana

Draws dashboards and can also send alerts

Dashboards and settings, but not your measurements

What did the payment service look like during last night's release?

Alertmanager

Groups, silences and routes alerts

Short-term alert state

Who should hear about this, and how?

One more term is worth knowing: observability. It means being able to understand what's happening inside a system from what it reports. Three kinds of reports matter. Metrics are numbers over time. Logs are written records of events. Traces follow the path of one request across services. Prometheus handles metrics only, which is a deliberate limit and shapes how you pair it with other tools.

Who uses it: a snapshot

Before picking among observability tools, it helps to see what other teams run. Grafana Labs publishes an annual survey of practitioners. One caveat: the company's own 2024 report says its open-source-friendly community likely skews the results, so treat the numbers as a view of that community rather than of every business.

Finding

Figure

Source

Organizations using Prometheus in production in some capacity

67% (another 19% investigating or building trials)

Grafana Labs survey, 2025 (1,255 responses)

Respondents investing in Prometheus

77%

Grafana Labs survey, 2026 (1,363 responses, 76 countries)

Organizations investing in both Prometheus and OpenTelemetry

65%

Grafana Labs survey, 2026

Alert fatigue named as the biggest obstacle to faster incident response

30%

Grafana Labs survey, 2026

Complexity and overhead named as top observability concern

38%

Grafana Labs survey, 2026

Organizations using SaaS for observability in some capacity

50% (up from 43% in 2025)

Grafana Labs survey, 2026

Mid-size and large enterprises where one hour of downtime costs over $300,000

Over 90%

ITIC Hourly Cost of Downtime Survey, 2024

OpenTelemetry, mentioned in the table, is a newer open standard for collecting metrics, logs and traces. Prometheus 3.0, released on November 14, 2024 as the project's first major release in seven years, can take in OpenTelemetry metrics directly. The two are used side by side more often than as rivals, though that receiver handles metrics only.

How a number travels from your code to a chart

1.      Your service publishes a page of numbers. Each service exposes a simple web page, usually at an address ending in /metrics, listing its current counts and timings. Libraries for most programming languages create this page. For software you can't change, such as a database or a Linux server, a small helper program called an exporter publishes the numbers instead.

2.      Prometheus visits each page on a schedule. This visit is called a scrape. Because Prometheus does the fetching (the "pull" model), it notices right away when a service stops answering. It also records a built-in value called up for every target, 1 for a good scrape and 0 for a failed one. You get a basic health check without writing anything.

3.      Prometheus stores what it collected. The data lands in a time series database on its own disk. The official documentation puts the average at 1 to 2 bytes per sample, which is small.

4.      You ask questions in PromQL. That's the Prometheus query language. A question like "what share of checkout requests failed over the last five minutes?" becomes one short expression.

5.      Rules watch the answers. If a rule's condition holds for long enough, Prometheus sends an alert to Alertmanager, which decides who hears about it.

6.      Grafana draws the charts. It runs queries against Prometheus and shows the results as graphs, tables and gauges. Most people in a company only ever see this last step.

Every number Prometheus stores is one of four types:

•         Counter: only goes up, such as total orders placed. Prometheus works out how fast it's growing, and a restart that resets it to zero is handled.

•         Gauge: moves up and down, such as memory in use or jobs waiting in a queue.

•         Histogram: sorts measurements into buckets, for example how many responses took under 0.1 seconds, under 0.5 seconds and under 1 second. Percentiles can be worked out from it, even across many copies of a service.

•         Summary: similar to a histogram, but the percentiles are calculated inside the app, which makes them harder to combine across several copies of a service.

Each number also carries labels, which are tags such as service="checkout" or status="500". Labels let you slice one measurement by service, version or region. They also cause the most trouble at scale, which comes up later.

Pro tip. In your first week, stop one test service on purpose and watch up drop to 0 in Prometheus. Seeing a failure appear on screen once makes the alert rules you write later easier to trust.

Setting it up for the first time

Start with one service and one dashboard. It's easier to learn the tools on something small, and a single working example is simple to copy to the next service.

The order of work looks like this. Run Prometheus (the project provides ready-made packages and container images, and its web page opens on port 9090 by default). Add a metrics library to one service so it publishes a /metrics page. Tell Prometheus where that page is. Run Grafana (port 3000 by default), add Prometheus as a data source by entering its address, then build or import a dashboard.

Telling Prometheus where to look takes a few lines in its configuration file. This example scrapes a service called orders every 15 seconds:

global:

  scrape_interval: 15s

scrape_configs:

  - job_name: "orders"

static_configs:

   - targets: ["orders:8080"]

A fixed list of addresses works for a handful of services. It stops working when containers appear and disappear every few minutes. For that case Prometheus supports service discovery: it asks the platform (Kubernetes, for example) which services exist right now and starts scraping new ones automatically. Teams on Kubernetes often install a bundle such as the kube-prometheus-stack Helm chart, which sets up Prometheus, Alertmanager, Grafana and a starter set of dashboards together.

A first round of Prometheus monitoring for a single service gives you request counts, error counts and response times. That's enough to answer the question from the opening scenario.

What to measure

The hard part isn't installation. It's choosing which numbers matter, because one service can publish thousands. Three short checklists have become standard, and they overlap.

•         The four golden signals, from Google's site reliability engineering book: latency, traffic, errors and saturation.

•         The RED method, from Tom Wilkie (now CTO of Grafana Labs): rate, errors and duration. It suits services that answer requests.

•         The USE method, from Brendan Gregg: utilization, saturation and errors. It suits machines and resources such as CPU, disk and memory.

For a checkout service, the common set looks like this:

Signal

What it means

Example for a checkout service

Rate

How many requests arrive

Orders submitted per minute

Errors

How many requests fail

Share of checkout requests that return a server error

Duration

How long requests take

Time to confirm payment, shown as the 95th and 99th percentile

Saturation

How full a resource is

Database connections in use out of the maximum allowed

Duration deserves care, because an average can hide the pain. Here's a made-up example. A service answers 99 requests in 100 milliseconds and one request in 8 seconds. The average works out to about 180 milliseconds, which looks fine, while one customer in a hundred waits 8 seconds. The 99th percentile is the time that 99% of requests beat, and it shows the problem straight away. That's the reason histograms are worth their small extra setup.

Business numbers belong here too. Orders per minute, sign-ups and failed payments can be measured the same way, and the 2026 Grafana Labs survey found that half of organizations now use their monitoring setup for business-related metrics. Basic system monitoring of CPU and memory still has a place, but it works best underneath the request-level numbers.

Building Grafana dashboards people read

A dashboard fails when nobody opens it. The usual cause is trying to show everything on one screen. A few habits help:

•         One dashboard per question. "Is checkout healthy?" gets its own screen, and "why is the database slow?" gets another.

•         Overview first. Request rate, error rate and response time go at the top, with detail underneath.

•         Use variables. Grafana can add drop-downs for service, environment or version, so one dashboard serves every service instead of 30 near-identical copies.

•         Mark deployments on the graphs. A thin vertical line at each release explains many sudden changes.

•         Let Grafana pick the rate window. Its built-in $__rate_interval variable chooses a window long enough for your scrape interval, which helps avoid empty or jumpy graphs.

The Grafana integration with Prometheus is a built-in data source, so connecting the two takes an address and a click. The more useful work is deciding who each dashboard is for. Developers want detail. A founder or an office manager wants three numbers and a color: orders per minute, error rate, and whether anything is red. Both views can come from the same Prometheus data.

Alerts that don't cry wolf

Dashboards help when someone is looking. Alerts are for when nobody is. The 2026 Grafana Labs survey found that alert fatigue was the biggest single obstacle to faster incident response, named by 30% of respondents and nearly double the next most common answer. Too many alerts teach people to ignore all of them. These settings keep the count down:

•         Alert on symptoms customers feel, such as error rate or slow checkout, instead of every possible cause such as high CPU on one machine.

•         Add a waiting period. Prometheus alert rules have a for setting, so a condition must hold for, say, five minutes before the alert fires. A one-minute blip then stays quiet. Grafana calls the same idea a pending period.

•         Let Alertmanager group alerts. The official documentation describes a case where hundreds of service instances can't reach one database and each sends its own alert. Grouping turns them into one notification.

•         Use inhibition and silences. Inhibition mutes smaller alerts when a bigger one is already firing. A silence mutes an alert for a set time, such as during planned maintenance.

•         Run more than one Alertmanager. It supports clustering for high availability. The documentation says not to put a load balancer between Prometheus and the Alertmanagers, and to point Prometheus at every instance instead.

Prometheus rules plus Alertmanager is one path. Grafana has its own alerting too, and a single Grafana rule can query several data sources. Either works, but give each alert one home. Two systems watching the same thing with slightly different settings produce duplicate pages and arguments about which one is right. A good Grafana integration can still show the current alert list on a dashboard, so people see warnings next to the graphs.

Pro tip. Write the first action into the alert text. "Checkout error rate above 5% for 5 minutes. Check recent deploys, then database connections." Someone half awake at 3 a.m. needs a starting point more than a diagnosis.

When the data has gaps

Real systems drop data. A network hiccup, a restart or a busy server can make a scrape fail and leave a hole in the graph. Knowing how Prometheus and Grafana treat holes saves you from false alarms and, worse, false calm.

A failed scrape

When a scrape fails, up drops to 0 for that target and its graphs show a gap. Prometheus also marks a series as stale when it stops appearing, so a dead service's numbers don't linger as if they were current.

A scrape interval that's too long

When you ask Prometheus for the latest value, it looks back up to five minutes by default. A series scraped less often than that vanishes from the answer between scrapes. A PromCon 2017 talk on staleness by Brian Brazil, a longtime Prometheus developer, puts the practical limit at around two minutes, since you have to allow for one failed scrape.

Short-lived jobs

A nightly backup job may finish before Prometheus ever visits it. Such jobs can push their results to a helper called the Pushgateway, which Prometheus then scrapes. The catch is that the Pushgateway deliberately has no expiry for pushed values, so a job that stopped running months ago can still show its last success. The official instrumentation guide says the key metric for a batch job is the time of its last success, which you can alert on when it gets too old.

Silence that looks like health

A graph with no line and a graph sitting at zero look alike at a glance, yet they mean different things. A new service that has never had an error may publish no error series at all, so a query dividing errors by requests returns nothing instead of 0%. The instrumentation guide suggests exporting a default value of 0 for any series you know will exist.

Grafana alert rules have to decide what "no data" means. You can set it to Alerting, Normal, Error or Keep Last State. Keep Last State quiets short data source problems, but Grafana's own documentation warns that it may not suit cases where strict monitoring is critical. A related setting controls how many checks in a row can come back empty before Grafana treats an alert as resolved. For a customer-facing service, an alert on missing data is usually worth having.

When the signals disagree

Two numbers that don't match are normal. The skill is knowing which one to trust for which decision.

What you see

Likely reason

What to do

up shows 1 but customers report failures

The service answers its metrics page even though something it depends on, such as the database, is broken

Alert on error rate and on a real test transaction, not only on up

Average response time looks fine, but users say it's slow

The average hides the slowest requests

Chart the 95th and 99th percentiles

The dashboard shows a spike but no alert fired

The problem didn't last through the waiting period, or the alert uses a different time window

Use the same saved query for both and review the waiting period

Two Prometheus servers show slightly different numbers

Each scraped at slightly different moments

Expected. Use one for dashboards and don't compare to the decimal

Metrics look healthy but logs show errors

Metrics summarize. A rare error can disappear inside a percentage

Use metrics to see that something is off, and logs to see why

The saved query mentioned in the third row is called a recording rule. Prometheus calculates it ahead of time and stores the result as a new series. Using one for both the dashboard and the alert keeps them consistent, and it makes dashboards load faster too.

Making decisions in real time

Monitoring feels instant, but it has delays built in, and each one is a trade-off. Take a typical alert with example settings. Prometheus scrapes every 15 seconds. A rule checks the data on its own schedule, often once a minute. The rule waits five minutes before firing. Alertmanager then pauses briefly to group related alerts before sending. Add it up and a real problem can take more than five minutes to reach a phone.

Shorter delays mean faster warnings and more false alarms. Longer delays mean calmer nights and slower response. Match the setting to the damage a delay would cause: a payments service might justify a two-minute wait, while an internal reporting tool can wait fifteen. Faster scraping also costs storage, since halving the interval doubles the samples.

Some decisions can be automated. Autoscalers on Kubernetes can add capacity based on a Prometheus query (KEDA is one tool that does this), and teams often compare error rates between the old and new version during a release to decide whether to roll back. Automation raises the stakes on data quality, because a gap or a stale number can trigger the wrong action. Give any automated decision a safe default for when data goes missing.

How it behaves under pressure and at scale

Cardinality, the quiet cost

Every unique combination of labels is its own time series. Take a request counter with a method label (4 values), a status label (5) and an endpoint label (20). That's 400 series for each copy of the service. Add a user ID label with 10,000 users and it becomes 4,000,000. These numbers are an illustration, but the multiplication is real. Memory and disk use in Prometheus follow the number of series, not just traffic. The official guidance says a metric whose label combinations could pass 100 deserves a second look. Detail at the level of single users belongs in logs or traces.

Storage arithmetic

The documentation gives a simple formula: retention time in seconds, times samples per second, times bytes per sample. Try 100,000 active series scraped every 15 seconds with the default 15-day retention. That's about 8.6 billion samples, or roughly 9 GB at 1 byte per sample and 17 GB at 2 bytes. The docs also suggest setting the retention size to at most 80 to 85% of the disk, because compaction needs temporary room.

One server is one server

Prometheus's local storage isn't clustered or replicated, so it's limited to one machine's capacity and durability, and the documentation says to manage it like any single-node database. Teams usually grow in steps:

Setup

Good for

Weak spot

One Prometheus server

Learning, small teams, a few dozen services

A lost disk means lost data, and one machine sets the limit

Two identical servers

Surviving one server failing

Double the cost, and their numbers differ slightly

Prometheus plus remote storage (Thanos, Mimir or similar)

Long history and many services queried in one place

More moving parts to run

A managed service, such as Grafana Cloud

Teams that don't want to run storage

Ongoing fees that grow with data volume

Grafana Labs announced Mimir 3.0 on November 5, 2025, describing it as a horizontally scalable, Prometheus-compatible metrics backend. Prometheus also has an agent mode, started with the --agent flag, that only collects data and forwards it elsewhere.

Watching the watcher

During an outage more people open dashboards, and heavy queries pile onto the same server that's collecting data. Keep refresh rates sensible and use recording rules for expensive queries. Alert on Prometheus's own numbers too, such as how long scrapes take. Many teams also add an alert that always fires and send it to an outside service. When it stops arriving, the alerting chain itself is down. Alertmanager is built to lean toward sending duplicates over missing an alert if cluster members lose contact, which is the right bias for a pager.

What it costs and who does the work

Prometheus and Grafana are free to download. The bill shows up as servers, storage and people's time. The 2026 Grafana Labs survey is useful here. Half of respondents use SaaS for observability in some capacity, up from 43% the year before, and cost was the top tool-selection criterion for the third year running at 65%, followed by ease of use at 49%. Teams that run everything themselves were most likely to name complexity as their top concern, while SaaS users more often pointed to cost.

For a small team, a sensible path is to start self-hosted on one server, learn what you actually need to measure, and then decide whether running it yourself is worth the hours. If you're comparing observability tools as a business owner, three questions help. Who is on call when the monitoring itself breaks? How long is history kept? How does the bill grow when we add services?

Mistakes teams make early

•         Tracking everything. Start with rate, errors and duration, and add more when a real question needs it.

•         Putting user IDs, session IDs or full URLs into labels.

•         Alerting on causes instead of symptoms.

•         Reading "no data" as "no problem."

•         Leaving all history on one local disk with no plan for backups or growth.

•         Building dashboards and alerts with no named owner. Treat system monitoring like any other product that someone is responsible for.

Key takeaways

•         Prometheus collects and stores numbers, Grafana shows them, and Alertmanager decides who is told.

•         Start with rate, errors and duration for one service, and watch the 95th and 99th percentiles instead of averages.

•         Missing data needs its own handling. An empty graph is not a healthy graph.

•         Labels multiply series. Keep them few and never use user IDs.

•         A single Prometheus server has hard limits. Plan for a second server or remote storage before you need it.

Conclusion

Monitoring microservices comes down to a loop. Collect a few meaningful numbers, store them cheaply, show them where people will look, and alert only when a person needs to act. Prometheus and Grafana cover that loop well, and both are free to try. Start with one service, learn how gaps and conflicting numbers behave on your own data, and add structure as you grow. Done this way, Prometheus monitoring turns into quiet background work, and the next 9:40 Tuesday email arrives after the alert instead of before it.

Nidhi Jain

Nidhi Jain

Nidhi is an exceptionally talented and creative content writer, bringing life to ideas through her words. With marketing knowledge and a deep understanding of various industries, she crafts captivating content that resonates with our audience. Her in-depth knowledge of trending tech and consumer affairs adds a unique perspective to her work, making it engaging and impactful.

Build Your Agile Team

We provide you with a top-performing extended team for all your development needs in any technology.

Hourly
$20
It Includes
Duration
Hourly Basis
Communication
Phone, Skype, Slack, Chat, Email
Hiring Period
25 Hours (MIN)
Project Trackers
Daily Reports, Basecamp, Jira, Redmime, etc
Methodology
Agile
Monthly
$2600
It Includes
Duration
160 Hours
Communication
Phone, Skype, Slack, Chat, Email
Hiring Period
1 Month
Project Trackers
Daily Reports, Basecamp, Jira, Redmime, etc
Methodology
Agile
Team
$13200
It Includes
Team Members
1 (PM), 1 (QA), 4 (Developers)
Communication
Phone, Skype, Slack, Chat, Email
Hiring Period
1 Month
Project Trackers
Daily Reports, Basecamp, Jira, Redmime, etc
Methodology
Agile

Frequently Asked Questions

Do I need Kubernetes to use Prometheus?
No. Prometheus works against ordinary servers and containers outside Kubernetes. You list the addresses to scrape in its configuration file or use another service discovery option. Its Kubernetes support is strong, which is why the two are often mentioned together, but it's optional.
Do I need both Prometheus and Grafana?
Prometheus has its own web page for running queries and viewing quick graphs, and version 3.0 gave it a rewritten interface. That's fine for an engineer investigating a problem. Most teams add Grafana for dashboards that other people can read, with variables, shared layouts and alerting. A Grafana integration takes little effort, so most teams run both.
How long does Prometheus keep data?
By default, 15 days. You can change it with a time limit, a size limit or both, and whichever is reached first removes the oldest data. For months or years of history, send data to a long-term storage system, since Prometheus's local storage isn't meant to be durable long-term storage by itself.
Can Prometheus replace logs and traces?
No. It stores numbers over time. It can tell you that error rates doubled at 2:10 p.m., but it can't show the failed request or the error message. Other observability tools handle that, such as Loki for logs and Tempo or Jaeger for traces.
How much does it cost?
The software is free and open source. You pay for the machines, the disk and the time to maintain it. The official documentation says Prometheus averages 1 to 2 bytes per sample, so storage stays modest until the number of series grows large. Managed services take over the maintenance and charge for usage instead.