Every observability vendor gives you an agent or SDK that sends telemetry straight from your app to their backend. It works, the docs are good, and it's one less thing to run. A collector in the middle is extra infrastructure, so it needs a real reason to be there. Here's the reason.
What sending directly ties you to
When your apps talk straight to a vendor, that vendor ends up spread across every service you own. Their SDK is in your dependencies. Their attribute naming is in your instrumentation. Their sampling settings live in your app config and only change when the app is released.
None of this hurts at first. It hurts the day someone asks "can we try a different backend?" and the answer is "only after we re-instrument everything". By then switching costs more than it saves, so nobody switches. That's not really a tooling problem. It's a coupling problem, created years earlier by a decision nobody wrote down.
Direct to vendor
Each app carries each vendor's SDK and config. Switching vendors means touching every app.
Through a collector
Apps only speak OTLP. The vendor is one exporter config in one place.
What the collector changes
With a collector in the middle, apps just send OTLP and don't know anything else. The backend becomes exporter config in one component, which you can change, review and roll back on its own. Want to run a second backend side by side while you evaluate it? That's a config block, not a migration.
Processing moves out of the apps too. Scrubbing attributes, redacting PII, sampling and batching all happen in one component the platform team owns, instead of every service doing its own slightly different version.
Agent, gateway, or both
There are two ways to run it, and most real setups use both.
- Agent: one collector per node, usually a DaemonSet. It sits close to the workload, so it can add node and pod metadata and ride out short network blips. It's cheap, and it scales with your nodes.
- Gateway: a central deployment that all the agents forward to. Anything that needs the full picture goes here. Tail-based sampling is the big one, because a single node only ever sees part of a trace. It's also the one place to control egress and handle backend credentials, instead of doing that on every node.
The platform I built runs both. A DaemonSet in each workload cluster adds Kubernetes metadata and forwards to a gateway fleet in a separate observability cluster, which handles filtering, batching and sampling. Putting the gateway in its own cluster matters more than it sounds: an incident in a workload cluster can't take down the tools you need to debug it.
When you don't need it
The collector isn't free. It's another deployment to run, size, monitor and get paged for. And it sits between your services and your ability to see them, so if it breaks, your debugging breaks with it.
If you have one service, one team, and a vendor you're happy to stay with, sending directly is fine. The collector is worth it when your telemetry needs to outlive a vendor decision. Just be clear about which situation you're in, rather than adding it because it's the recommended architecture.