Splunk APM and RUM: What They Add and When to Use Them

Table of Contents

Summarize the Content of the Blog

Key takeaways

APM and RUM watch opposite ends of the same request. APM sees your services; RUM sees your users' browsers. Neither alone tells the whole story for a web application.
A backend can be healthy while users have a slow experience, because front-end rendering, network, and third-party scripts sit outside APM's view. RUM catches that gap.
RUM measures the metrics that actually correlate with user experience: web vitals like Largest Contentful Paint and Cumulative Layout Shift [2].
Which to start with depends on where your pain is. Backend latency points to APM first; user complaints that backend metrics do not explain point to RUM first.

The two ends of a request

When a user loads your web application, a request travels from their browser, across the network, into your services, to a database, and back, and then the browser renders the result. That full path has two very different halves.

The server half is your infrastructure and application code: the services that receive the request, process it, and return a response. The client half is everything that happens in the user's browser: downloading assets, running JavaScript, rendering the page, and reacting to input.

APM watches the server half. RUM watches the client half. This is the cleanest way to understand the two, and it explains immediately why teams with customer-facing applications usually need both: a problem in either half degrades the experience, and each tool is blind to the other's half.

What APM sees

Splunk APM monitors traces and spans from your distributed applications, so you can troubleshoot issues affecting key business workflows and improve application performance [1].

In plain terms, APM follows a single request through every service it touches on the server side. When a checkout request is slow, APM shows you the trace: it entered the API gateway, called the cart service, which called the pricing service, which called the database, and the database call took 800 of the 900 milliseconds. Without APM, you know checkout is slow. With APM, you know exactly which service and which call to fix.

This is indispensable in a microservices architecture, where a single request can cross a dozen services and the slow one is impossible to find by inspecting each in isolation. APM makes the invisible path of a distributed request visible, which is the whole reason distributed systems need it. The broader platform this sits in is covered in Splunk Observability Services: Scope, Fit, and Cost.

What APM does not see: anything that happens in the user's browser after your server responds. As far as APM is concerned, once the response leaves your infrastructure, the request is done. That blind spot is exactly what RUM fills.

What RUM sees

Splunk RUM provides insight into the front-end user experience. It collects performance metrics, web vitals, errors, and other data so you can detect and troubleshoot problems, measure the health of your application, and assess the performance of the user experience [2].

RUM runs in the user's actual browser, so it captures what the user actually experienced, not what your servers think they delivered. It measures the metrics that correlate with real experience: web vitals such as Largest Contentful Paint (how long until the main content appears), Cumulative Layout Shift (how much the page jumps around as it loads), and front-end errors that never touch your backend.

This matters because the user does not experience your backend. They experience their browser. A page can be served quickly by a fast backend and still feel slow because a large image blocks rendering, a third-party analytics script hangs, or the user is on a poor mobile connection. None of that appears in APM. All of it appears in RUM.

RUM is where you find out about problems before the support tickets do, and it is the component that connects application performance to something the business cares about directly: whether customers have a good experience.

Why a fast backend can still feel slow

This is the scenario that justifies running both, and it is more common than teams expect.

Your APM dashboards are green. Response times are excellent. Every service is fast. And users are complaining that the site is slow. Without RUM, this is baffling, and teams waste days re-checking a backend that is genuinely fine.

The explanation is almost always in the client half that APM cannot see. The backend returns in 100 milliseconds, but the browser then spends two seconds downloading a bloated JavaScript bundle, waiting on a slow third-party tag, and rendering a layout that shifts twice. The user's experience is a two-second wait. Your backend metrics say 100 milliseconds. Both are true, and only RUM reconciles them.

This gap, between backend performance and perceived performance, is why a customer-facing application needs both APM and RUM. APM keeps your services fast. RUM tells you whether that speed actually reaches the user. Optimizing one without the other means you are either fixing a backend that was already fine or shipping a fast backend behind a slow experience.

Which to start with

You do not have to adopt both at once. Start where your pain is.

Start with APM if: your problem is backend latency, your architecture is distributed enough that finding the slow service by hand is impractical, or your incidents are about service performance rather than user experience.

Start with RUM if: users complain about speed that your backend metrics do not explain, front-end experience is commercially important, or you have no visibility at all into what users actually experience.

Add the second when the first reveals its blind spot, which it will. Teams that start with APM eventually hit the "backend is green, users complain" wall and reach for RUM. Teams that start with RUM eventually see a slow user experience traced to the backend and reach for APM. The sequence is driven by where you are blindest today, and the components are designed to be adopted incrementally rather than all at once.

bitsIO, a four-time Splunk Partner of the Year and Splunk Elite Partner, helps teams decide which observability components fit their architecture and sequence adoption through its Splunk Observability practice. Where cost is a factor in that sequencing, the levers are covered in Reducing Splunk Observability Costs Without Losing Visibility.

Frequently asked questions

APM (application performance monitoring) watches your backend, tracing a request across services to find which is slow [1]. RUM (real user monitoring) watches your frontend, capturing what real users experience in their browsers, including web vitals and errors [2]. APM sees your services; RUM sees your users.

It monitors traces and spans from distributed applications, letting you follow a single request across every service it touches and pinpoint which service and which call is slow [1]. This is essential in microservices architectures where the slow component is impossible to find by inspecting services in isolation.

RUM measures the front-end user experience: performance metrics, web vitals such as Largest Contentful Paint and Cumulative Layout Shift, front-end errors, and other data from the user's actual browser [2]. It captures what the user experienced, not what your servers think they delivered.

For customer-facing web applications, usually yes. A fast backend can still deliver a slow experience because of front-end rendering, network, or third-party scripts that APM cannot see. RUM catches that gap. Together they cover the full path from user to backend.

Because the slowness is in the client half that APM cannot see: a large JavaScript bundle, a slow third-party script, a shifting layout, or a poor connection. Your backend metrics are correct and so are the users. RUM reconciles them by measuring the actual browser experience.

Start where your pain is. Backend latency and hard-to-locate slow services point to APM first. User complaints that backend metrics do not explain point to RUM first. Add the second when the first reveals its blind spot, which it usually does.

Web vitals are metrics that correlate with real user experience, such as Largest Contentful Paint (time until main content appears) and Cumulative Layout Shift (how much the page moves as it loads). RUM captures these from real browsers so you can measure experience the way users actually feel it [2].

They are components of Splunk Observability Cloud and are adopted based on need rather than all at once. You can start with the component that addresses your pain and add others as blind spots appear. The full set of components and how they fit together is covered in the observability services guide.

‍

Unlock the Full Potential of Your Data

Boost Efficiency and Maximize ROI with bitsIO’s Advanced Solutions

Start Today – Optimize Your Splunk!