Performance profiling
Capturing where a Service spends its time or memory, down to the function.
Performance profiling
A profile shows where a running Service actually spends its time or memory, down to the function, drawn as a flame graph. SlideOps samples the running code for a few seconds. Nothing is restarted, nothing in your application changes, and nothing runs until you press the button.
This page covers capturing a profile of one Service. Reading a flame graph explains what you get back, and the pages after it cover comparing profiles, profiling continuously, whole-server profiles and the Profiles page.
Where it lives
Every Service has a Performance tab. It opens on what SlideOps found without changing anything:
- Ready to profile, or Cannot profile right now with the reason.
- The runtime it detected for each process, such as
CPython 3.12 × 4 processesorJVM. - The server's Linux kernel version and processor type.
- Which languages can be profiled?, the full list below.
Capability Services, such as a PostgreSQL and Redis pair, have no Performance tab. They are infrastructure rather than an application, so profile the application that uses them instead.
What a server needs
| Requirement | Why |
|---|---|
| Linux 5.10 or newer | The profiler runs in the kernel with eBPF. Every current Ubuntu, Debian, RHEL and Amazon Linux release qualifies. |
| An x86_64 or ARM64 processor | The profiler is built for both. |
| Root, or passwordless sudo, for the account SlideOps signs in with | Loading eBPF programs needs root. |
| A kernel not in lockdown mode | Lockdown forbids eBPF profiling outright. |
| The Service running | There is nothing to profile in a stopped Service. |
When any of these is missing, the Performance tab says which one, in those words, before you try.
Capturing a profile
The form asks two questions.
What do you want to know?
| Question | Use it for | Languages |
|---|---|---|
| What is using the CPU | High CPU, slow code, unexpected work. Only time spent running counts. | Every language below |
| Where time goes, including waiting | Slow requests. Also counts time waiting on the network, disk, locks and timers. | Every language below |
| What every thread is doing right now | A Service that is stuck or hung. A 5 second snapshot, one box per thread. | Every language below |
| What is waiting on locks | Contention and slow concurrent code: where threads wait for a lock or a condition. | Every language below |
| What is holding memory | Leaks and large memory use. | Go, Node.js, Java, Kotlin, Scala |
| Where memory is allocated | Garbage collection pressure and memory churn: every allocation during the capture, by bytes. | Go, Node.js, Java, Kotlin, Scala |
| Where memory grows | Each time the Service asks the system for more memory, and which code asked. | Every language below |
A question this Service cannot answer is shown greyed out with the reason. For a Python Service, for example, the two memory questions say that memory in use is read from the runtime's own profiler, which Python does not offer to an outside tool, and point you to Where memory grows instead.
Two memory notes are worth knowing:
- Go programs need to serve
net/http/pprofon a local port. SlideOps finds the port by itself, inside the container. - For Node.js and Java, What is holding memory shows what was allocated during the capture and is still held at the end, which is exactly where a leak shows. For Go it is all memory in use.
For how long?
10, 30 or 60 seconds; 30 is the default. Capture while the Service is doing the thing you want to understand: under load, or during the slow request. The threads snapshot always takes 5 seconds.
What SlideOps will do
Before you start, the form lists every step in plain words, for example:
- Install the SlideOps profiler at
/opt/slideops/profiler/slideops-profiler(about 23 MB, checksum verified). It stays there for the next capture. - Sample the 4 processes of
billingabout 99 times a second, using eBPF in the kernel. Nothing is restarted and nothing in the Service changes. - Expected overhead: about 1% of one CPU while sampling.
- Read the result back, name each function, then check the Service is still running.
Extra steps appear when a runtime needs them. For Java memory, SlideOps copies async-profiler (under 1 MB, checksum verified) into the container's /tmp for the capture and removes it afterwards. For Node.js 24 and newer, it opens V8's inspector inside the container for the capture and closes it again. V8 recompiles the application's optimised code once at the start, which slows it for a moment.
Press Capture for 30 seconds (or Take a 5 second snapshot). The card that replaces the form says which step it is on and how long it has run. You can leave the page: the capture keeps running, and its flame graph opens by itself when it finishes.
Only one capture runs on a server at a time. Starting a second while one runs is refused with "a capture is already running on this server".
Verification
After every capture SlideOps checks that the Service is still running and that the profile has enough samples and names to be useful. Anything worth knowing appears as a note at the top of the profile, for example:
- "Very few samples: the Service was almost idle. Capture again while it is under the load you want to understand."
- "Many native frames have no names, because the binaries were built without symbols." See Server profiles and debug symbols.
- A process a memory capture could not read, and why.
Languages
| Language | Versions proven | Notes |
|---|---|---|
| Python | 3.9 to 3.13 | |
| Node.js (JavaScript, TypeScript) | 18, 20, 22, 24 | Node.js 24 and newer are profiled through V8's inspector. |
| Java, Kotlin, Scala | Java 11, 17, 21; Kotlin 2 | |
| .NET (C#, F#) | 8, 9 | |
| Go | 1.23, 1.24 | Memory needs net/http/pprof. |
| Ruby | 3.3, 3.4, 4.0 | |
| PHP | 8.2, 8.3, 8.4 | |
| Rust, C, C++, Zig | any compiler | Names come from the binary's symbols. |
| Erlang, Elixir | OTP 27, Elixir 1.18 | Time in built-in functions written in C shows as its own stack. |
| Perl | 5.36 to 5.40 | Partial: measured, but not yet split by sub. |
| Deno | Partial: native frames only. | |
| Bun | Partial: native frames only. |
Every version in this table is proven by a test that profiles a real application in that language and checks the right function is named. A language is only listed once its test passes.
Finding out why
Three places elsewhere in SlideOps link straight here with the right question already chosen, and a note saying why you came:
| From | When | Opens with |
|---|---|---|
| Live usage on a Service's Overview | CPU at 80% of a core or more | What is using the CPU |
| Uptime, on the Service page or the Uptime page | The Service is down, its last check failed, or it took 2 seconds or more to answer | Where time goes, including waiting |
| A deploy in the Service's Activity | Any successful deploy | Before and after that deploy; see Comparing profiles |
Each link says Find out why, or Compare performance before and after for a deploy.
Past profiles
Below the form, Past profiles lists every finished capture, newest first, with its answer, length and languages. The latest 50 are kept per Service; older ones are removed as new ones arrive. Profiles taken by continuous profiling are kept for a day instead, and are not listed here; see Continuous profiling.
From the list you can open a profile, delete it, tick two to compare them, or use Compare with the one before for the previous profile of the same type.
Downloading
Download on a profile offers two formats:
- pprof file, for
go tool pprof, Pyroscope, Speedscope and most profilers. - Collapsed stacks, for
flamegraph.pland other flame graph tools.
Every download is recorded in the audit trail.
Safety and access
- Nothing is restarted, and no code or configuration of yours changes.
- Profiles hold function names, files and line numbers, never variables, arguments or data your Service handled.
- Profiles belong to their Workspace. Viewers can read them; capturing, deleting and changing settings need a role above Viewer.
- Every capture, download and deletion is recorded in the audit trail.
Removing the profiler
The profiler stays installed on a server so the next capture starts quickly. To remove it, open the server's Performance tab and choose Remove the profiler from this server. The next capture installs it again.