The two dashboards that disagree
Grafana published a demo last December that should be pinned above every platform team's desk. They took a plain Spring Boot REST service, instrumented it two ways at once, and then overloaded it. At 10:16 they cranked up the parallel request count until Tomcat's thread pool was exhausted and requests started queuing.
The eBPF dashboard showed p95 latency jump from about 0.9 seconds to roughly 3.3 seconds and stay there. The OpenTelemetry Java agent dashboard, watching the same service at the same moment, showed a flat green line at about 800 milliseconds. Same service. Same traffic. Same second. A 4x disagreement about how slow the thing was.
Neither dashboard is broken. The SDK measures the time a request is actively being serviced, which does not include the time it spent waiting in Tomcat's queue. eBPF, sitting down at the network and kernel layer, measures the full time from the client's perspective, queue included. The SDK is answering "how long did my code run." eBPF is answering "how long did the client wait." Those are different questions, and in an overload they have wildly different answers.
This is the thing to internalize before you roll out any zero-code observability tool: eBPF instrumentation and SDK instrumentation are not two implementations of the same measurement. They are two different vantage points, and the mistake almost everyone makes on the first pass is treating them as substitutes. The tell that they are not substitutes is hiding in plain sight: the flagship goal on OpenTelemetry's own 2026 roadmap for its eBPF tool is making it work alongside the SDKs, not replace them.
What OBI actually is, and why the pitch is seductive
OpenTelemetry eBPF Instrumentation, OBI, is the project formerly known as Grafana Beyla. Grafana Labs donated Beyla to OpenTelemetry in May 2025 under the OBI name, and Beyla now continues as Grafana's downstream distribution of the upstream project. It shipped its first OpenTelemetry release later that year, and the community set a 2026 roadmap whose headline item is a stable 1.0.
The pitch is genuinely strong. OBI uses eBPF to inspect application executables and the OS networking layer and produce trace spans plus Rate-Errors-Duration metrics for HTTP, HTTPS, and gRPC services, with no changes to application code, no libraries to install, and no restarts. It reads TLS-encrypted transactions without decrypting them. The language matrix is enormous because eBPF does not care about your language: Java, .NET, Go, Python, Ruby, Node.js, C, C++, Rust. The protocol list keeps growing, from HTTP and gRPC to MQTT, NATS, AMQP, and a pile of databases. You deploy it once as a Kubernetes DaemonSet and every service on the node becomes visible.
Read that back and you can see why a tired platform engineer falls in love. Every service instrumented, instantly, including the legacy Java app older than Java 8 and the Python service older than 3.9 that no SDK will ever cover. No pull requests. No begging six teams to add a dependency. If your job is "make the fleet observable by Q3," OBI looks like it just did the whole thing for free.
And for one specific layer of the problem, it did. The trap is assuming that layer is the whole problem.
Why eBPF sees the wire better than your code does
Start with what eBPF is genuinely better at, because the argument for running both falls apart if you pretend eBPF is just the shallow option. It is not shallow. It is differently positioned.
Because eBPF instrumentation operates at the network layer, request rates, error rates, latencies, and service-graph edges become available for every service uniformly, independent of language, runtime, or how the app was deployed. It is a platform feature, rolled out to every node, not an application feature that each team opts into. That uniformity is worth more than it sounds. Half the pain of an observability program is the long tail of services that never got instrumented, and eBPF erases that tail in one deploy.
The Tomcat example is the sharp end of this. eBPF metrics represent what clients actually experienced, queue time and all, which is exactly the number you want during an incident. The SDK's cleaner-looking 800ms is arguably the more misleading figure in that moment, because it quietly excludes the seconds your users spent waiting to be served at all.
eBPF also hands you service-graph metrics out of the box, without distributed tracing. The usual way to learn which service calls which is to generate that graph from spans, which forces you to run traces at a 100% sample rate and then pay a caching-and-memory tax in the pipeline to stitch client and server spans together. Because eBPF observes every request at the kernel level, it produces service-graph and span metrics directly, unsampled and unbiased, with none of that overhead. When the community adds a protocol like AMQP, it lights up for every language at once, instead of waiting for each SDK to reimplement it.
So eBPF is not the consolation prize. For baseline service health across the whole fleet, it is the better tool. The problem is that "baseline service health across the whole fleet" is the floor of observability, not the ceiling.
Why the SDK still has the answers
Now the other vantage point. The reason you instrument at all is usually not to draw a latency line. It is to answer why a request was slow or failed, and that answer lives inside the process, where eBPF cannot follow.
Two limits make this concrete. First, distributed tracing. eBPF supports it "to some extent," in Grafana's own careful phrasing, but the support is bounded: OBI offers production-ready tracing for Go, Node.js, Python, Ruby, and for frameworks that do not switch threads while handling a request. For a reactive Java application that hands a request across threads, eBPF still produces spans, but they represent single network calls and may end up stand-alone, not correlated into the right trace. The place your architecture is most likely to be complicated is precisely the place eBPF tracing gets least reliable.
Second, depth. An OpenTelemetry Java agent instruments the internal layers of your application. It gives you spans for the Spring Boot handler, the Hibernate transaction boundary, the individual SELECT and UPDATE queries, and it attaches events like a thrown exception with its full stack trace. That level of detail is explicitly beyond the scope of eBPF instrumentation, because those events never cross the network or a syscall boundary where a kernel probe could see them. When your 500 needs a root cause, the stack trace on the SDK span is the thing that ends the investigation. eBPF can tell you the request failed. It cannot tell you it threw a ConstraintViolationException on line 214.
Then there is intent, which neither eBPF nor any SDK auto-instrumentation can recover. Auto-instrumentation captures technical facts: status codes, durations, query text. It cannot know that this request was a checkout paid with store credit rather than a card. That distinction is business logic, and the only way to get it into your telemetry is to call the OpenTelemetry API yourself and attach a payment.method attribute to the span or bump a counter. Auto-instrumentation gets you to the door. Your own API calls are what put the answer in the room.
The roadmap is the argument
Here is the part that convinced me this is structural and not a temporary gap. If eBPF were on track to swallow the SDK, you would expect the people building the best eBPF tool to be racing toward feature parity so they could retire the agents. They are doing the opposite.
The flagship of OBI's 2026 roadmap is a stable 1.0. Sitting right next to it, as a named goal with its own sponsor, is "hybrid instrumentation": making OBI work with the OpenTelemetry SDKs rather than around them. The described work is telling. OBI is being built to wrap SDK-generated traces so timing stays accurate regardless of which layer produced the span, to keep labels consistent between eBPF and SDK telemetry, and to avoid emitting duplicate signals. In practice, OBI and Beyla already auto-detect an SDK: if the SDK pushes spans to an OTLP endpoint, the eBPF tool turns its own tracing off; if the SDK pushes semantic-convention metrics, eBPF stops emitting those too, and keeps providing only what it is uniquely good at, the span metrics, service-graph metrics, network metrics, and process metrics.
The coordination even has a config surface. When you run an SDK for tracing alongside eBPF for metrics, you set one resource attribute so your pipeline does not double-count:
# On the SDK-instrumented service: let eBPF own the span-derived metrics,
# and tell Tempo / the Collector not to also generate them from these spans.
OTEL_RESOURCE_ATTRIBUTES="span.metrics.skip=true"
You do not lose anything by doing this. eBPF already produces unsampled span and service-graph metrics at the kernel level, so you let it own the baseline, and you let the SDK do what it is uniquely good at: high-fidelity traces with library-level context for debugging, run at a sane head-sampling rate because you are no longer mining those traces for metrics. A project spending its 1.0 cycle teaching its tool when to defer to the other tool is not a project that thinks one tool wins.
Yes, but eBPF keeps closing the gap
The honest objection: OBI is moving fast, so maybe this boundary is just where the maturity happens to be in late 2026, with OBI at v0.13.0 and the 1.0 epic still working through its pre-release blocker list. The v0.10.0 release added language-agnostic traceparent propagation for gRPC over HTTP/2. Context propagation for .NET is in progress. There are experimental span links for Go channel handoffs. Give it a year, the argument goes, and the reactive-Java gap closes and the depth gap narrows.
Some of that will happen, and the gap will genuinely shrink. But the core limits are not on a maturity curve, they are set by where the instrument sits. A kernel-and-network probe observes bytes crossing boundaries. It cannot see a Hibernate transaction commit that never leaves the process, it cannot attach the exception object that was caught and rethrown, and it cannot know that a request represents a refund rather than a purchase, because none of that is on the wire. You could make eBPF arbitrarily clever and it would still be reconstructing in-process causality from the outside, which is exactly the information that in-process instrumentation has for free. There are also plainer limits that no roadmap has erased yet: eBPF instrumentation is Linux-only today, and service-graph metrics currently only resolve real service names on Kubernetes. This is complementarity by construction, not a queue of bugs waiting to be fixed.
What to do on Monday
Stop running the eBPF-versus-SDK debate in your platform channel. It is the wrong axis. Decide, per signal, which layer produces it.
Deploy eBPF as a DaemonSet for the baseline: RED metrics, service graphs, network and process metrics, for every service on every node, including the ones no team will ever instrument by hand. That is your breadth, and it is unbiased because it is unsampled. Keep SDK instrumentation on the services you actually debug, run traces at a reasonable head-sampling rate since you are no longer generating metrics from them, and let those traces carry the internal spans, the exception stack traces, and the trace-id-on-logs correlation that end investigations. Then instrument your business logic explicitly with the OpenTelemetry API, because the one attribute that turns a slow trace into a known cause is the one only you can name. Wire the seam with span.metrics.skip=true so nothing double-counts.
When your eBPF dashboard and your SDK dashboard disagree during the next incident, do not open a bug. Ask which question each one is answering. The 3.3 seconds is what your users felt. The 0.8 is what your code did. You need both numbers, and you need to know which is which.

