all

Case study

Centralizing metrics and traces from remote retail sites with vmagent and OpenTelemetry

Use case

A client needed centralized observability for POS systems and servers running in stores and warehouses.

The requirement sounded simple:

collect metrics and application traces from all remote locations in one place.

The systems were outside the Kubernetes cluster and distributed across separate local networks. We needed to collect host metrics, monitor connectivity between local devices, and receive application traces without making every remote site directly scrapeable from the central infrastructure.

The resulting design used node_exporter and vmagent at the edge, Prometheus Remote Write for infrastructure metrics, and OpenTelemetry Collector with VictoriaTraces for distributed tracing.

A client needed better visibility into systems running outside their central infrastructure.

Their application stack included POS systems and supporting servers installed across stores and warehouses. Each location had its own local network, while the main monitoring stack lived in a central Kubernetes cluster.

The network topology made the usual Prometheus model less convenient. Having the central Prometheus server directly scrape every remote machine would require reliable inbound connectivity to each site and additional routing, VPN, or firewall configuration.

We used a different model: collect locally and send telemetry outward.

Architecture

The central Kubernetes cluster already had Prometheus.

We added:

  • VictoriaTraces for trace storage and querying;
  • OpenTelemetry Collector as the OTLP ingestion layer.

At each remote location we installed:

  • node_exporter for host metrics;
  • vmagent for local Prometheus scraping and Remote Write delivery.

The applications running at the remote sites could send traces using OTLP/HTTP to the central OpenTelemetry endpoint.

  flowchart LR
    subgraph EDGE["Store / warehouse network"]
        POS["POS application"]
        NODE["node_exporter"]
        BLACKBOX["blackbox exporter"]
        VM["vmagent"]

        NODE --> VM
        BLACKBOX --> VM
    end

    subgraph K8S["Central Kubernetes cluster"]
        PROM["Prometheus"]
        OTEL["OpenTelemetry Collector"]
        VT["VictoriaTraces"]
        VMS["VictoriaMetrics"]
        VL["VictoriaLogs"]

        OTEL --> VT
        OTEL --> VMS
        OTEL --> VL
    end

    VM -->|"Prometheus Remote Write / HTTPS"| PROM
    POS -->|"OTLP/HTTP / HTTPS"| OTEL

This gave us two deliberately separate telemetry paths.

Infrastructure metrics followed:

1
2
3
4
5
6
7
node_exporter / blackbox exporter
        ↓
      vmagent
        ↓
Prometheus Remote Write
        ↓
central Prometheus

Application traces followed:

1
2
3
4
5
6
7
application
     ↓
   OTLP
     ↓
OpenTelemetry Collector
     ↓
VictoriaTraces

Keeping the two paths separate also made troubleshooting easier. A problem with application tracing could not prevent basic host and network metrics from reaching Prometheus.

Collecting metrics locally with vmagent

vmagent ran close to the monitored systems and handled the normal Prometheus scrape process locally.

A simplified configuration looked like this:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
# promscrape.yml
---
global:
  scrape_interval: 1m

  external_labels:
    source: edge-vmagent
    environment: retail
    site: store-042

scrape_configs:
  - job_name: node_exporter

    static_configs:
      - targets:
          - 10.44.8.14:9201

  - job_name: icmp_checks
    scrape_interval: 1m

    metrics_path: /probe

    params:
      module:
        - icmp

    static_configs:
      - targets:
          - 10.44.8.179
          - 10.44.9.76
          - 10.44.9.81

    relabel_configs:
      - source_labels:
          - __address__
        target_label: __param_target

      - source_labels:
          - __param_target
        target_label: instance

      - target_label: __address__
        replacement: 10.44.8.21:9127

The node_exporter target provided normal Linux host metrics.

The second job used a blackbox exporter to check whether important devices on the local network were reachable. This was useful for POS and warehouse environments because a central system could otherwise see that a store was unavailable without knowing whether the problem was the Internet connection, the local server, or an individual device.

Add a site label early

One detail matters once several locations start sending the same metrics.

An instance such as:

1
10.44.8.14:9201

is not globally meaningful. Different stores can easily use identical RFC1918 address ranges.

A stable site label avoids collisions:

1
2
external_labels:
  site: store-042

Queries can then use combinations such as:

1
2
3
4
up{
  site="store-042",
  job="node_exporter"
}

instead of relying on an IP address to identify a physical location.

Pushing metrics to central Prometheus

The local agent was started with Remote Write enabled:

1
2
3
4
args:
  - -promscrape.config=/etc/vmagent/promscrape.yml
  - -remoteWrite.url=https://metrics-ingest.example.net/api/v1/write
  - -remoteWrite.tmpDataPath=/var/lib/vmagent/remotewrite

This changes the direction of the monitoring connection.

Instead of:

1
Prometheus → remote store → node_exporter

we have:

1
remote store → central Prometheus

Only outbound connectivity from the site is required.

vmagent can also queue Remote Write data when the central endpoint is temporarily unreachable. That matters at retail and warehouse locations where WAN connectivity is usually less predictable than connectivity inside a datacenter.

For production installations, the queue should live on persistent storage:

1
/var/lib/vmagent/remotewrite

rather than an ephemeral directory if metrics need to survive an agent or machine restart.

Enabling Remote Write ingestion in Prometheus

Prometheus does not automatically accept Remote Write traffic on a normal installation.

The receiver must be enabled:

1
--web.enable-remote-write-receiver

After that, the endpoint becomes:

1
/api/v1/write

The public endpoint in front of it can then be protected with TLS and the authentication mechanism used by the cluster.

For a limited number of remote sites this was a convenient way to reuse the Prometheus stack the client already operated.

For a much larger edge deployment, we would normally reconsider whether Prometheus itself should remain the Remote Write ingestion backend. A dedicated remote-write storage system such as VictoriaMetrics becomes more attractive as ingestion volume and retention requirements grow.

Adding distributed tracing

Host metrics answer questions such as:

  • Is the POS server running?
  • Is memory pressure increasing?
  • Can a local device be reached?
  • Did the site disappear from monitoring?

They cannot explain why an application request took several seconds or where time was spent between services.

For that we added distributed tracing.

The central Kubernetes cluster received OTLP data through an OpenTelemetry Collector.

A simplified receiver configuration was:

1
2
3
4
5
6
7
8
receivers:
  otlp:
    protocols:
      grpc:
        endpoint: 0.0.0.0:4321

      http:
        endpoint: 0.0.0.0:4320

The externally exposed endpoint terminated TLS before forwarding OTLP traffic to the collector.

Applications therefore only needed an OTLP endpoint such as:

1
https://otel-ingest.example.net:8443

while communication inside Kubernetes remained on the cluster network.

Exporting traces to VictoriaTraces

The collector already provided a convenient place to route different OpenTelemetry signals.

The relevant exporter configuration looked like this:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
exporters:
  otlphttp/victoria:
    compression: gzip
    encoding: proto

    metrics_endpoint: >-
      http://vm-single.telemetry.svc.cluster.local:8430/opentelemetry/v1/metrics

    logs_endpoint: >-
      http://vl-single.telemetry.svc.cluster.local:9430/insert/opentelemetry/v1/logs

    traces_endpoint: >-
      http://vt-single.telemetry.svc.cluster.local:10430/insert/opentelemetry/v1/traces

The service pipelines then connected the OTLP receiver to the appropriate backend:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
service:
  pipelines:
    traces:
      receivers:
        - otlp

      processors: []

      exporters:
        - otlphttp/victoria

    metrics:
      receivers:
        - otlp

      processors:
        - deltatocumulative

      exporters:
        - otlphttp/victoria

    logs:
      receivers:
        - otlp

      processors: []

      exporters:
        - otlphttp/victoria

The infrastructure metrics collected by vmagent still went to Prometheus.

The metrics pipeline shown here covered metrics produced by applications through OpenTelemetry. This distinction was useful because it allowed us to introduce application instrumentation without changing the existing Prometheus monitoring model.

Testing the OpenTelemetry path without an application

Before debugging application instrumentation, we wanted to prove that the telemetry pipeline itself worked:

1
2
3
4
5
OTLP request
    ↓
OpenTelemetry Collector
    ↓
VictoriaTraces

A manually generated span was enough.

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
START_NS=$(date -u +%s%N)
END_NS=$((START_NS + 1000000000))

curl -v \
  http://otel-gateway.telemetry.svc.cluster.local:4320/v1/traces \
  -H 'Content-Type: application/json' \
  --data-binary "{
    \"resourceSpans\": [
      {
        \"resource\": {
          \"attributes\": [
            {
              \"key\": \"service.name\",
              \"value\": {
                \"stringValue\": \"edge-curl-test\"
              }
            }
          ]
        },
        \"scopeSpans\": [
          {
            \"scope\": {
              \"name\": \"manual-test\"
            },
            \"spans\": [
              {
                \"traceId\": \"b27cc6b7af7943d2a2af33628159c4e1\",
                \"spanId\": \"29c7fd16fae24803\",
                \"name\": \"manual-otel-test-span\",
                \"kind\": 1,
                \"startTimeUnixNano\": \"$START_NS\",
                \"endTimeUnixNano\": \"$END_NS\",
                \"status\": {
                  \"code\": 1
                }
              }
            ]
          }
        ]
      }
    ]
  }"

A successful response looked like:

1
2
3
4
5
6
* Connected to otel-gateway.telemetry.svc.cluster.local
> POST /v1/traces HTTP/1.1
> Content-Type: application/json
>
< HTTP/1.1 200 OK
< Content-Type: application/json

The generated span could then be located in VictoriaTraces using:

1
service.name = edge-curl-test

This test was useful because it removed the application SDK from the equation.

If the synthetic span appeared in VictoriaTraces, the central OTLP path was working. Any remaining problem was on the application or external network side.

Common failure points

Several problems are worth checking before debugging the application itself.

Remote Write returns 404

Check whether Prometheus was started with:

1
--web.enable-remote-write-receiver

Without the receiver, /api/v1/write will not accept the vmagent stream.

Metrics arrive, but different stores overwrite each other’s identity

Check labels.

RFC1918 addresses are frequently reused between locations. Add an explicit site, location, or similar label before sending the data centrally.

vmagent loses buffered metrics after a reboot

Check:

1
-remoteWrite.tmpDataPath

If the queue is stored under an ephemeral filesystem, the local retry buffer disappears with it.

Use persistent local storage when retaining metrics during WAN outages matters.

OTLP returns HTTP 200 but no trace appears

Check the complete collector pipeline.

Having an OTLP receiver is not enough. The trace signal must also be connected to an exporter:

1
2
3
4
5
6
7
8
service:
  pipelines:
    traces:
      receivers:
        - otlp

      exporters:
        - otlphttp/victoria

Then verify the VictoriaTraces endpoint configured in the exporter.

Trace timestamps look wrong

Distributed tracing depends heavily on timestamps from different systems.

POS terminals, local servers, Kubernetes nodes, and application hosts should all have working time synchronization. Clock drift can make an otherwise valid trace difficult to interpret and can also cause problems for time-series ingestion.

Why the edge-push model worked well here

The main advantage was operational rather than architectural novelty.

We did not need the monitoring cluster to establish connections into every store and warehouse.

Each site only needed to know where to send telemetry:

1
2
metrics → central Remote Write endpoint
traces  → central OTLP endpoint

Local Prometheus-style scraping could continue even during a temporary WAN problem, and vmagent handled delivery to the central system.

At the same time, OpenTelemetry gave application developers a standard interface for traces without coupling the applications directly to VictoriaTraces.

The resulting system kept responsibilities reasonably clear:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
node_exporter
    host telemetry

blackbox exporter
    local connectivity checks

vmagent
    local scraping and metric forwarding

Prometheus
    infrastructure metrics

OpenTelemetry Collector
    application telemetry ingestion and routing

VictoriaTraces
    trace storage and search

Result

After the rollout, the client had one central observability point for systems that previously lived in isolated store and warehouse networks.

Prometheus could show host health and local network reachability for each site.

VictoriaTraces could show what happened inside application requests.

Most importantly, adding a new remote location no longer required the central monitoring system to gain inbound scrape access to that network. The site received the standard local monitoring components, a unique location label, and the two central ingestion endpoints.

That made the same pattern reusable for additional stores, warehouses, POS systems, and other infrastructure running outside Kubernetes.