Case study
Centralizing metrics and traces from remote retail sites with vmagent and OpenTelemetry
Use case
A client needed centralized observability for POS systems and servers running in stores and warehouses.
The requirement sounded simple:
collect metrics and application traces from all remote locations in one place.
The systems were outside the Kubernetes cluster and distributed across separate local networks. We needed to collect host metrics, monitor connectivity between local devices, and receive application traces without making every remote site directly scrapeable from the central infrastructure.
The resulting design used node_exporter and vmagent at the edge, Prometheus Remote Write for infrastructure metrics, and OpenTelemetry Collector with VictoriaTraces for distributed tracing.
A client needed better visibility into systems running outside their central infrastructure.
Their application stack included POS systems and supporting servers installed across stores and warehouses. Each location had its own local network, while the main monitoring stack lived in a central Kubernetes cluster.
The network topology made the usual Prometheus model less convenient. Having the central Prometheus server directly scrape every remote machine would require reliable inbound connectivity to each site and additional routing, VPN, or firewall configuration.
We used a different model: collect locally and send telemetry outward.
Architecture
The central Kubernetes cluster already had Prometheus.
We added:
- VictoriaTraces for trace storage and querying;
- OpenTelemetry Collector as the OTLP ingestion layer.
At each remote location we installed:
node_exporterfor host metrics;vmagentfor local Prometheus scraping and Remote Write delivery.
The applications running at the remote sites could send traces using OTLP/HTTP to the central OpenTelemetry endpoint.
flowchart LR
subgraph EDGE["Store / warehouse network"]
POS["POS application"]
NODE["node_exporter"]
BLACKBOX["blackbox exporter"]
VM["vmagent"]
NODE --> VM
BLACKBOX --> VM
end
subgraph K8S["Central Kubernetes cluster"]
PROM["Prometheus"]
OTEL["OpenTelemetry Collector"]
VT["VictoriaTraces"]
VMS["VictoriaMetrics"]
VL["VictoriaLogs"]
OTEL --> VT
OTEL --> VMS
OTEL --> VL
end
VM -->|"Prometheus Remote Write / HTTPS"| PROM
POS -->|"OTLP/HTTP / HTTPS"| OTEL
This gave us two deliberately separate telemetry paths.
Infrastructure metrics followed:
Application traces followed:
Keeping the two paths separate also made troubleshooting easier. A problem with application tracing could not prevent basic host and network metrics from reaching Prometheus.
Collecting metrics locally with vmagent
vmagent ran close to the monitored systems and handled the normal Prometheus scrape process locally.
A simplified configuration looked like this:
| |
The node_exporter target provided normal Linux host metrics.
The second job used a blackbox exporter to check whether important devices on the local network were reachable. This was useful for POS and warehouse environments because a central system could otherwise see that a store was unavailable without knowing whether the problem was the Internet connection, the local server, or an individual device.
Add a site label early
One detail matters once several locations start sending the same metrics.
An instance such as:
| |
is not globally meaningful. Different stores can easily use identical RFC1918 address ranges.
A stable site label avoids collisions:
Queries can then use combinations such as:
instead of relying on an IP address to identify a physical location.
Pushing metrics to central Prometheus
The local agent was started with Remote Write enabled:
This changes the direction of the monitoring connection.
Instead of:
| |
we have:
| |
Only outbound connectivity from the site is required.
vmagent can also queue Remote Write data when the central endpoint is temporarily unreachable. That matters at retail and warehouse locations where WAN connectivity is usually less predictable than connectivity inside a datacenter.
For production installations, the queue should live on persistent storage:
| |
rather than an ephemeral directory if metrics need to survive an agent or machine restart.
Enabling Remote Write ingestion in Prometheus
Prometheus does not automatically accept Remote Write traffic on a normal installation.
The receiver must be enabled:
| |
After that, the endpoint becomes:
| |
The public endpoint in front of it can then be protected with TLS and the authentication mechanism used by the cluster.
For a limited number of remote sites this was a convenient way to reuse the Prometheus stack the client already operated.
For a much larger edge deployment, we would normally reconsider whether Prometheus itself should remain the Remote Write ingestion backend. A dedicated remote-write storage system such as VictoriaMetrics becomes more attractive as ingestion volume and retention requirements grow.
Adding distributed tracing
Host metrics answer questions such as:
- Is the POS server running?
- Is memory pressure increasing?
- Can a local device be reached?
- Did the site disappear from monitoring?
They cannot explain why an application request took several seconds or where time was spent between services.
For that we added distributed tracing.
The central Kubernetes cluster received OTLP data through an OpenTelemetry Collector.
A simplified receiver configuration was:
The externally exposed endpoint terminated TLS before forwarding OTLP traffic to the collector.
Applications therefore only needed an OTLP endpoint such as:
| |
while communication inside Kubernetes remained on the cluster network.
Exporting traces to VictoriaTraces
The collector already provided a convenient place to route different OpenTelemetry signals.
The relevant exporter configuration looked like this:
| |
The service pipelines then connected the OTLP receiver to the appropriate backend:
| |
The infrastructure metrics collected by vmagent still went to Prometheus.
The metrics pipeline shown here covered metrics produced by applications through OpenTelemetry. This distinction was useful because it allowed us to introduce application instrumentation without changing the existing Prometheus monitoring model.
Testing the OpenTelemetry path without an application
Before debugging application instrumentation, we wanted to prove that the telemetry pipeline itself worked:
A manually generated span was enough.
| |
A successful response looked like:
The generated span could then be located in VictoriaTraces using:
| |
This test was useful because it removed the application SDK from the equation.
If the synthetic span appeared in VictoriaTraces, the central OTLP path was working. Any remaining problem was on the application or external network side.
Common failure points
Several problems are worth checking before debugging the application itself.
Remote Write returns 404
Check whether Prometheus was started with:
| |
Without the receiver, /api/v1/write will not accept the vmagent stream.
Metrics arrive, but different stores overwrite each other’s identity
Check labels.
RFC1918 addresses are frequently reused between locations. Add an explicit site, location, or similar label before sending the data centrally.
vmagent loses buffered metrics after a reboot
Check:
| |
If the queue is stored under an ephemeral filesystem, the local retry buffer disappears with it.
Use persistent local storage when retaining metrics during WAN outages matters.
OTLP returns HTTP 200 but no trace appears
Check the complete collector pipeline.
Having an OTLP receiver is not enough. The trace signal must also be connected to an exporter:
Then verify the VictoriaTraces endpoint configured in the exporter.
Trace timestamps look wrong
Distributed tracing depends heavily on timestamps from different systems.
POS terminals, local servers, Kubernetes nodes, and application hosts should all have working time synchronization. Clock drift can make an otherwise valid trace difficult to interpret and can also cause problems for time-series ingestion.
Why the edge-push model worked well here
The main advantage was operational rather than architectural novelty.
We did not need the monitoring cluster to establish connections into every store and warehouse.
Each site only needed to know where to send telemetry:
Local Prometheus-style scraping could continue even during a temporary WAN problem, and vmagent handled delivery to the central system.
At the same time, OpenTelemetry gave application developers a standard interface for traces without coupling the applications directly to VictoriaTraces.
The resulting system kept responsibilities reasonably clear:
| |
Result
After the rollout, the client had one central observability point for systems that previously lived in isolated store and warehouse networks.
Prometheus could show host health and local network reachability for each site.
VictoriaTraces could show what happened inside application requests.
Most importantly, adding a new remote location no longer required the central monitoring system to gain inbound scrape access to that network. The site received the standard local monitoring components, a unique location label, and the two central ingestion endpoints.
That made the same pattern reusable for additional stores, warehouses, POS systems, and other infrastructure running outside Kubernetes.