Monitoring Idako Application Health with NetData

Getting started · monitoring

When you're collecting and logging critical industrial data, losing any of it usually isn't an option — which means you need to know, 24/7, that the application itself is running smoothly, and to find out immediately the moment something isn't. Idako 4.3.2 added exactly that: a comprehensive set of health metrics exposed through the /health endpoint in Prometheus format, the industry standard for monitoring — letting you plug in a monitoring agent to watch every component and get notified the instant something goes wrong.

This guide walks through wiring that endpoint into NetData: two of its built-in collectors read /health — one that works today with zero setup, another that unlocks per-component and per-server detail once you're on Idako 4.3.2 — and NetData's own alerting emails you the moment something breaks. Nothing new to run: if a NetData agent is already watching this host, this uses it as-is.

about 15 minutes Idako 4.3.2 or later, for full metrics a NetData agent already installed network access to your Idako instance's health port outbound email already working on the NetData host
1

Get the config files

Download the ready-to-copy config files and unzip them as a folder named idako-netdata — four small files that plug into a NetData install you already have, nothing built from scratch. Two are collector job files (where to look), two are matching alert rule files (what counts as a problem):

idako-netdata/ ├── README.md ├── go.d/ │ ├── httpcheck.conf │ └── prometheus.conf └── health.d/ ├── idako-httpcheck.conf └── idako-prometheus.conf
Two collectors, checking two different things. httpcheck (step 2) works against /health exactly as it is today, on any Idako version, and is what reliably catches the instance being completely unreachable. The prometheus collector (step 3) needs Idako 4.3.2's /health?format=prometheus and unlocks per-component status and per-server detail. Run both side by side — one isn't a replacement for the other.
2

Point httpcheck at Idako

Open go.d/httpcheck.conf — it polls /health with a plain HTTP request and checks for a 200 response containing "status":"OK". The only thing to adjust is the target address:

go.d/httpcheck.conf
update_every: 5

jobs:
  - name: idako_local
    url: http://idako-host:4880/health
    status_accepted: [200]
    response_match: '"status"\s*:\s*"OK"'
    timeout: 3

Replace idako-host and 4880 with your Idako instance's actual host and port. The job's name matters too: it becomes the chart family NetData groups this check under, and it's what the alert rules in step 4 match against by the idako_ prefix — keep that prefix if you rename it.

3

Add richer metrics via Prometheus

Idako 4.3.2 also exposes /health in Prometheus's own text format, via a query parameter — the same format the Idako + Grafana guide reads with a dedicated Prometheus server, read here instead by NetData's own prometheus collector. Where httpcheck only tells you up or down, this charts every subsystem individually: the OPC UA collector engine, the local buffer, the tsdb connection, and each configured server by name.

go.d/prometheus.conf
update_every: 10

jobs:
  - name: idako_local
    url: http://idako-host:4880/health?format=prometheus
    app: idako
    expected_prefix: idako_
    max_time_series: 500

Same host/port substitution as step 2. app: idako groups the resulting charts under a predictable name in the dashboard; expected_prefix is a guard rail against accidentally scraping something that isn't Idako.

Confirm the exact chart names once, after step 6. NetData derives each chart's internal name from the metric name and this app: setting. The alert rules in the next step already use the documented pattern (prometheus.idako.<metric>), but it's worth a one-time check once data is flowing: curl -s http://localhost:19999/api/v1/contexts | grep idako.
4

Install the alert rules

Two files, one per collector from steps 2–3, written as templates — a template attaches to every matching chart automatically, so adding a second Idako instance later (see "Monitoring more than one instance") needs no changes here.

Reachability — pairs with httpcheck

health.d/idako-httpcheck.conf
template: idako_health_unreachable
      on: httpcheck.status
 families: idako_*
   lookup: average -1m unaligned percentage of success
    units: %
    every: 10s
     warn: $this < 100
     crit: $this < 50
    delay: down 30s multiplier 1.5 max 2m
     info: percentage of /health checks that succeeded (HTTP 200 + status:OK) in the last minute
       to: sysadmin

template: idako_health_response_slow
      on: httpcheck.response_time
 families: idako_*
   lookup: average -1m unaligned
    units: ms
    every: 10s
     warn: $this > 2000
     crit: $this > 5000
    delay: down 1m multiplier 1.5 max 5m
     info: average response time of Idako's /health endpoint over the last minute
       to: sysadmin

families: idako_* scopes both templates to job names starting with idako_, so they won't fire on unrelated httpcheck jobs running on the same agent.

Component health — pairs with the Prometheus collector

health.d/idako-prometheus.conf
template: idako_up_prom
      on: prometheus.idako.idako_up
   lookup: average -2m unaligned
    units: ratio
    every: 10s
     warn: $this < 1
     crit: $this == 0
    delay: down 30s multiplier 1.5 max 2m
     info: Idako reports overall status OK
       to: sysadmin

template: idako_collector_up_prom
      on: prometheus.idako.idako_collector_up
   lookup: average -1m unaligned
    units: ratio
    every: 15s
     warn: $this < 1
    delay: down 30s multiplier 1.5 max 2m
     info: Idako's OPC UA collector engine is not running
       to: sysadmin

template: idako_collector_disconnected_servers_prom
      on: prometheus.idako.idako_collector_disconnected_servers
   lookup: min -1m unaligned
    units: servers
    every: 30s
     warn: $this > 0
    delay: down 30s multiplier 1.5 max 2m
     info: one or more configured, active OPC UA servers are disconnected
       to: sysadmin

template: idako_collector_server_down_prom
      on: prometheus.idako.idako_collector_server_up
   lookup: min -1m unaligned foreach *
    units: ratio
    every: 30s
     warn: $this < 1
    delay: down 30s multiplier 1.5 max 2m
     info: this OPC UA server connection is not OK
       to: sysadmin

template: idako_collector_bad_variables_prom
      on: prometheus.idako.idako_collector_bad_variables
   lookup: average -1m unaligned
    units: variables
    every: 30s
     warn: $this > 0
     crit: $this > 100
    delay: down 1m multiplier 1.5 max 5m
     info: variables reporting bad quality, summed across the whole fleet
       to: sysadmin

template: idako_buffer_up_prom
      on: prometheus.idako.idako_buffer_up
   lookup: average -1m unaligned
    units: ratio
    every: 15s
     warn: $this < 1
    delay: down 30s multiplier 1.5 max 2m
     info: Idako's local buffer subsystem is not running
       to: sysadmin

template: idako_buffer_values_lost_prom
      on: prometheus.idako.idako_buffer_values_lost
   lookup: sum -5m unaligned incremental
    units: values
    every: 1m
     warn: $this > 0
    delay: down 1m multiplier 1.5 max 5m
     info: values could not be stored in the local buffer and were lost in the last 5 minutes
       to: sysadmin

template: idako_tsdb_down_prom
      on: prometheus.idako.idako_tsdb_up
   lookup: average -1m unaligned
    units: ratio
    every: 15s
     warn: $this < 1
    delay: down 30s multiplier 1.5 max 2m
     info: Idako's connection to the time-series database is not OK
       to: sysadmin

template: idako_tsdb_batches_failed_prom
      on: prometheus.idako.idako_tsdb_batches_failed
   lookup: sum -5m unaligned incremental
    units: batches
    every: 1m
     warn: $this > 0
    delay: down 1m multiplier 1.5 max 5m
     info: batches of values failed to forward to the tsdb in the last 5 minutes
       to: sysadmin
RuleFires whenSeverity
Idako unreachablehttpcheck stops getting HTTP 200 + status:OKwarning → critical
/health responding slowlyaverage response time > 2s (warn) / 5s (crit)warning / critical
Idako reports not OKidako_up is 0critical
Collector engine downidako_collector_up is 0warning
OPC UA server disconnectedidako_collector_server_up is 0 — one alert per serverwarning
Bad variables detectedidako_collector_bad_variables > 0 (warn) / > 100 (crit)warning / critical
Local buffer downidako_buffer_up is 0warning
Values lostidako_buffer_values_lost > 0 in the last 5mwarning
TSDB downidako_tsdb_up is 0warning
TSDB batch failuresidako_tsdb_batches_failed > 0 in the last 5mwarning
Why two separate "is it there" checks? httpcheck.status always writes a real 0 the moment a request fails outright — connection refused, timeout, wrong status code. idako_up_prom only evaluates once NetData has parsed a fresh Prometheus scrape; if the /health call fails completely, there's no idako_up sample to evaluate at all — a gap, not a 0 — and NetData's alarm engine doesn't reliably treat a gap as a failure. Keep both rules active: idako_health_unreachable catches "completely gone," idako_up_prom catches "up, but internally unhealthy."
5

Set the alert recipient

Every rule above ends in to: sysadmin — that's a role, not an address. NetData maps roles to real recipients in one file that (like Grafana's SMTP settings) isn't part of what you downloaded, since it holds where alerts actually go rather than what counts as a problem:

shell — Linux
cd /etc/netdata
sudo ./edit-config health_alarm_notify.conf

Set (or confirm) these values:

health_alarm_notify.conf
SEND_EMAIL="YES"
EMAIL_SENDER="alerts@example.com"
DEFAULT_RECIPIENT_EMAIL="you@example.com"

DEFAULT_RECIPIENT_EMAIL is where every to: sysadmin alarm above goes, since sysadmin is NetData's built-in default role. To send Idako's alerts to a different address than everything else on this host, set role_recipients_email[sysadmin]="you@example.com" instead — it overrides the default for just that one role.

Needs a working mail transport on this host. health_alarm_notify.conf hands the message to the system's own sendmail-compatible command to actually deliver it. If this host doesn't already send outbound mail, point it at a relay or install a local MTA first — that's a host-level prerequisite this file can't provide on its own.
On Windows, the same file lives at C:\Program Files\Netdata\etc\netdata\health_alarm_notify.conf — edit it from the bundled MSYS2 shell (C:\Program Files\Netdata\msys2.exe) as Administrator, since edit-config is a bash script and Program Files needs elevation to write to.
6

Turn it on

Three mechanical steps: confirm both collectors are enabled, copy the four files from step 1 into NetData's live config directory, and restart.

Enable the collectors

Both ship with default_run: yes, so this is usually a check, not a change:

shell — Linux
cd /etc/netdata
sudo ./edit-config go.d.conf
# confirm under modules:
#   httpcheck: yes
#   prometheus: yes
On Windows, run the equivalent from the bundled MSYS2 shell as Administrator: cd /etc/netdata && ./edit-config go.d.conf.

Copy the files in

shell — Linux
sudo cp go.d/httpcheck.conf     /etc/netdata/go.d/
sudo cp go.d/prometheus.conf    /etc/netdata/go.d/
sudo cp health.d/idako-httpcheck.conf   /etc/netdata/health.d/
sudo cp health.d/idako-prometheus.conf  /etc/netdata/health.d/
On Windows (PowerShell, as Administrator), Copy-Item the same four files into C:\Program Files\Netdata\etc\netdata\go.d\ and ...\health.d\ respectively.

Restart

shell — Linux
sudo systemctl restart netdata
On Windows, Restart-Service Netdata from an elevated PowerShell (or Services → Netdata → Restart).
7

Confirm it's working

Open http://localhost:19999 (or your NetData host's address) — NetData's own dashboard, no separate login to set up.

Look for an idako section in the left-hand menu. httpcheck's status and response-time charts should appear within a few seconds of restarting; the Prometheus-collector charts (per-component status, per-server detail) appear once step 3's job has scraped at least once. It should look something like this:

The idako section of NetData's dashboard: httpcheck status and response-time charts alongside per-component status for the collector engine, local buffer, and tsdb, plus a connection chart per configured OPC UA server.
NetData's idako section, running against a healthy instance.
You should see charts for httpcheck status and response time, plus — once the prometheus collector has run — separate status charts for the collector engine, local buffer, and tsdb, and one per configured OPC UA server. If the prometheus charts are missing, re-check the target directly with curl -s http://idako-host:4880/health?format=prometheus and revisit the one-time context-name check from step 3.
Beyond one instance

Monitoring more than one Idako instance

Add one more job to each collector file — nothing else changes, since every alert rule from step 4 is a template (or already scoped with families: idako_*) that automatically covers every matching job:

go.d/httpcheck.conf (excerpt)
  - name: idako_site_b
    url: https://idako-site-b.example.com/health
    status_accepted: [200]
    response_match: '"status"\s*:\s*"OK"'
    timeout: 3
go.d/prometheus.conf (excerpt)
  - name: idako_site_b
    url: https://idako-site-b.example.com/health?format=prometheus
    app: idako
    expected_prefix: idako_
    max_time_series: 500

Restart NetData to pick up the new jobs — unlike a health.d-only change, adding a job needs a full restart, not just reload-health.

Before you trust it

Verifying the alerts actually work

A clean restart only proves NetData accepted the config, not that a real notification gets delivered. Two checks worth doing once, right after setup:

  • Send a test notification — NetData ships a self-test for exactly this: sudo /usr/libexec/netdata/plugins.d/alarm-notify.sh test sends a real test message through every configured method, sysadmin included. Confirms step 5's settings actually deliver, before any real alarm depends on them.
  • Trigger a real one — stop Idako (or block the port) and wait a minute or two. idako_health_unreachable should move to warning, then critical, and an email should follow shortly after. Bring Idako back and you should get a second, "recovered" notification once it clears.
Notifications not arriving? Start with the self-test above, not the health.d files — if alarm-notify.sh test doesn't deliver either, it's a mail-transport problem on this host, not a config problem with the rules themselves.
Next steps

Where to go from here

Want fuller dashboards and PromQL-based alerting?The Idako + Grafana guide reads this same /health?format=prometheus endpoint with a dedicated Prometheus server and a Grafana dashboard — heavier to run, more powerful to query.
Detailed charts before every instance is on 4.3.2NetData's pandas collector can read the plain JSON /health body as a bridge — Linux-only in practice; confirm python.d.plugin and a Python 3 + pandas/requests environment before relying on it.
Route alerts to more than emailhealth_alarm_notify.conf supports Slack, PagerDuty, and a long list of other methods alongside email — same to: sysadmin roles, more delivery options per role.
See exactly what Idako reportsThe full field-by-field reference for /health, including every metric this guide's charts and alerts are built from.
This talks to Idako over plain HTTP on its health port — no credentials, and nothing new to install if a NetData agent is already watching this host.

Monitoring Idako Application Health with Grafana

Getting started · monitoring

When you're collecting and logging critical industrial data, losing any of it usually isn't an option — which means you need to know, 24/7, that the application itself is running smoothly, and to find out immediately the moment something isn't. Idako 4.3.2 added exactly that: a comprehensive set of health metrics exposed through the /health endpoint in Prometheus format, the industry standard for monitoring — letting you plug in a tool like Grafana to watch every component and get notified the instant something goes wrong.

This guide walks through a self-hosted monitoring stack you run in Docker: Prometheus scrapes that health data, Grafana turns it into a live dashboard — and emails you the moment something breaks. No cloud account, no agents to install on the Idako host itself.

about 25 minutes Idako 4.3.2 or later Docker & Docker Compose network access to your Idako instance's health port an SMTP relay, for email alerts
1

Get the project folder

Download the complete configuration and unzip it as a folder named idako-monitoring — everything below lives in that one folder, so the whole stack can be started, stopped, and moved as a unit. You'll adjust a handful of values inside it (your Idako host/port, SMTP details, alert recipient) over the next few steps; nothing needs to be built from scratch. Here's what's inside:

idako-monitoring/ ├── docker-compose.yml ├── .env-example ├── .gitignore ├── prometheus/ │ └── prometheus.yml └── grafana/ └── provisioning/ ├── datasources/ │ └── prometheus.yml ├── dashboards/ │ ├── dashboards.yml │ └── idako-overview.json └── alerting/ ├── contactpoints.yaml ├── policies.yaml └── rules.yaml
Note the missing .env. That file holds your real SMTP password and isn't included in the download — step 4 has you create it yourself from .env-example, and .gitignore (also from step 4) keeps it out of version control.
2

Point Prometheus at Idako

Open prometheus/prometheus.yml — it tells Prometheus what to scrape and how often, requesting Idako's health data in Prometheus format via a query parameter. The only thing to adjust is the target address:

prometheus/prometheus.yml
global:
  scrape_interval: 15s

scrape_configs:
  - job_name: idako
    metrics_path: /health
    params:
      format: [prometheus]
    static_configs:
      - targets: ["idako-host:4880"]

Replace idako-host and 4880 in targets with your Idako instance's actual host and port — that's the only edit this file needs.

Running Idako on the same machine as Docker? On Linux, use host.docker.internal only if your Docker version maps it (add extra_hosts: ["host.docker.internal:host-gateway"] under the prometheus service in step 5 if needed) — on Docker Desktop for Windows/Mac it works out of the box.
3

Provision the Grafana dashboard

Instead of clicking through Grafana's UI to add a data source and build panels by hand, three files already sitting in grafana/provisioning/ do it automatically, the same way every time you start the stack. Here's what each one does — none of them need editing.

Connects Grafana to Prometheus

grafana/provisioning/datasources/prometheus.yml
apiVersion: 1

datasources:
  - name: Prometheus
    uid: prometheus
    type: prometheus
    access: proxy
    url: http://prometheus:9090
    isDefault: true

http://prometheus:9090 works because Docker Compose puts both containers on the same network and lets them reach each other by service name — there's nothing to substitute here. The explicit uid: prometheus matters more than it looks: the dashboard panels below and the alert rules in step 4 both reference the datasource by this exact id, so pinning it here keeps everything pointed at the same place, rather than relying on whatever id Grafana would otherwise generate on its own.

Tells Grafana where to find dashboards

grafana/provisioning/dashboards/dashboards.yml
apiVersion: 1

providers:
  - name: Idako
    folder: Idako
    type: file
    updateIntervalSeconds: 30
    options:
      path: /etc/grafana/provisioning/dashboards

The dashboard itself

grafana/provisioning/dashboards/idako-overview.json — eleven panels covering every subsystem: instance, collector, local buffer, and tsdb status; connected/disconnected server counts; buffer backlog; tsdb batch failures; the buffer pipeline (collected/forwarded/balance); collection & forwarding rate; and variable quality (total/good/bad). Nothing to change here either, unless you want to customize it later:

grafana/provisioning/dashboards/idako-overview.json
{
  "title": "Idako Overview",
  "uid": "idako-overview",
  "schemaVersion": 39,
  "version": 4,
  "editable": true,
  "timezone": "browser",
  "time": { "from": "now-6h", "to": "now" },
  "refresh": "5s",
  "panels": [
    {
      "title": "Instance Status",
      "type": "stat",
      "gridPos": { "x": 0, "y": 0, "w": 4, "h": 6 },
      "datasource": { "type": "prometheus", "uid": "prometheus" },
      "targets": [{ "expr": "up{job=\"idako\"}", "refId": "A" }],
      "fieldConfig": {
        "defaults": {
          "mappings": [{ "type": "value", "options": {
            "0": { "text": "DOWN", "color": "red" },
            "1": { "text": "OK", "color": "green" }
          }}],
          "thresholds": { "mode": "absolute", "steps": [
            { "color": "red", "value": null }, { "color": "green", "value": 1 }
          ]}
        }, "overrides": []
      },
      "options": { "reduceOptions": { "calcs": ["lastNotNull"] }, "colorMode": "background", "graphMode": "none" }
    },
    {
      "title": "Collector Status",
      "type": "stat",
      "gridPos": { "x": 4, "y": 0, "w": 4, "h": 6 },
      "datasource": { "type": "prometheus", "uid": "prometheus" },
      "targets": [{ "expr": "idako_collector_up", "refId": "A" }],
      "fieldConfig": {
        "defaults": {
          "mappings": [{ "type": "value", "options": {
            "0": { "text": "DOWN", "color": "red" },
            "1": { "text": "OK", "color": "green" }
          }}],
          "thresholds": { "mode": "absolute", "steps": [
            { "color": "red", "value": null }, { "color": "green", "value": 1 }
          ]}
        }, "overrides": []
      },
      "options": { "reduceOptions": { "calcs": ["lastNotNull"] }, "colorMode": "background", "graphMode": "none" }
    },
    {
      "title": "Local Buffer Status",
      "type": "stat",
      "gridPos": { "x": 8, "y": 0, "w": 4, "h": 6 },
      "datasource": { "type": "prometheus", "uid": "prometheus" },
      "targets": [{ "expr": "idako_buffer_up", "refId": "A" }],
      "fieldConfig": {
        "defaults": {
          "mappings": [{ "type": "value", "options": {
            "0": { "text": "DOWN", "color": "red" },
            "1": { "text": "OK", "color": "green" }
          }}],
          "thresholds": { "mode": "absolute", "steps": [
            { "color": "red", "value": null }, { "color": "green", "value": 1 }
          ]}
        }, "overrides": []
      },
      "options": { "reduceOptions": { "calcs": ["lastNotNull"] }, "colorMode": "background", "graphMode": "none" }
    },
    {
      "title": "TSDB Status",
      "type": "stat",
      "gridPos": { "x": 12, "y": 0, "w": 4, "h": 6 },
      "datasource": { "type": "prometheus", "uid": "prometheus" },
      "targets": [{ "expr": "idako_tsdb_up", "refId": "A" }],
      "fieldConfig": {
        "defaults": {
          "mappings": [{ "type": "value", "options": {
            "0": { "text": "DOWN", "color": "red" },
            "1": { "text": "OK", "color": "green" }
          }}],
          "thresholds": { "mode": "absolute", "steps": [
            { "color": "red", "value": null }, { "color": "green", "value": 1 }
          ]}
        }, "overrides": []
      },
      "options": { "reduceOptions": { "calcs": ["lastNotNull"] }, "colorMode": "background", "graphMode": "none" }
    },
    {
      "title": "Connected Servers",
      "type": "stat",
      "gridPos": { "x": 16, "y": 0, "w": 4, "h": 6 },
      "datasource": { "type": "prometheus", "uid": "prometheus" },
      "targets": [{ "expr": "idako_collector_connected_servers", "refId": "A" }],
      "fieldConfig": { "defaults": { "color": { "mode": "thresholds" },
        "thresholds": { "mode": "absolute", "steps": [{ "color": "blue", "value": null }] } }, "overrides": [] },
      "options": { "reduceOptions": { "calcs": ["lastNotNull"] }, "graphMode": "none" }
    },
    {
      "title": "Disconnected Servers",
      "type": "stat",
      "gridPos": { "x": 20, "y": 0, "w": 4, "h": 6 },
      "datasource": { "type": "prometheus", "uid": "prometheus" },
      "targets": [{ "expr": "idako_collector_disconnected_servers", "refId": "A" }],
      "fieldConfig": { "defaults": { "color": { "mode": "thresholds" },
        "thresholds": { "mode": "absolute", "steps": [
          { "color": "green", "value": null }, { "color": "red", "value": 1 }
        ]} }, "overrides": [] },
      "options": { "reduceOptions": { "calcs": ["lastNotNull"] }, "graphMode": "none" }
    },
    {
      "title": "Buffer Backlog (values waiting to forward)",
      "type": "timeseries",
      "gridPos": { "x": 0, "y": 6, "w": 12, "h": 8 },
      "datasource": { "type": "prometheus", "uid": "prometheus" },
      "targets": [{ "expr": "idako_buffer_values_balance", "refId": "A" }],
      "fieldConfig": { "defaults": { "custom": { "drawStyle": "line", "lineWidth": 2, "fillOpacity": 15 },
        "color": { "mode": "palette-classic" } }, "overrides": [] },
      "options": { "legend": { "displayMode": "list", "placement": "bottom" }, "tooltip": { "mode": "single" } }
    },
    {
      "title": "TSDB Batch Failures (cumulative)",
      "type": "timeseries",
      "gridPos": { "x": 12, "y": 6, "w": 12, "h": 8 },
      "datasource": { "type": "prometheus", "uid": "prometheus" },
      "targets": [{ "expr": "idako_tsdb_batches_failed", "refId": "A" }],
      "fieldConfig": { "defaults": { "custom": { "drawStyle": "line", "lineWidth": 2, "fillOpacity": 15 },
        "color": { "fixedColor": "red", "mode": "fixed" } }, "overrides": [] },
      "options": { "legend": { "displayMode": "list", "placement": "bottom" }, "tooltip": { "mode": "single" } }
    },
    {
      "title": "Buffer Pipeline (collected / forwarded / balance)",
      "type": "timeseries",
      "gridPos": { "x": 0, "y": 14, "w": 8, "h": 8 },
      "datasource": { "type": "prometheus", "uid": "prometheus" },
      "targets": [
        { "expr": "idako_buffer_values_stored", "legendFormat": "Collected", "refId": "A" },
        { "expr": "idako_buffer_values_forwarded", "legendFormat": "Forwarded", "refId": "B" },
        { "expr": "idako_buffer_values_balance", "legendFormat": "Balance", "refId": "C" }
      ],
      "fieldConfig": {
        "defaults": { "custom": { "drawStyle": "line", "lineWidth": 2, "fillOpacity": 10 },
          "color": { "mode": "palette-classic" } },
        "overrides": [
          { "matcher": { "id": "byName", "options": "Balance" },
            "properties": [{ "id": "color", "value": { "mode": "fixed", "fixedColor": "orange" } }] }
        ]
      },
      "options": { "legend": { "displayMode": "list", "placement": "bottom" }, "tooltip": { "mode": "multi" } }
    },
    {
      "title": "Collection & Forwarding Rate (values/sec)",
      "type": "timeseries",
      "gridPos": { "x": 8, "y": 14, "w": 8, "h": 8 },
      "datasource": { "type": "prometheus", "uid": "prometheus" },
      "targets": [
        { "expr": "idako_collector_rate", "legendFormat": "Collection rate", "refId": "A" },
        { "expr": "idako_tsdb_values_forwarded_rate", "legendFormat": "Forwarding rate", "refId": "B" }
      ],
      "fieldConfig": { "defaults": { "custom": { "drawStyle": "line", "lineWidth": 2, "fillOpacity": 10 },
        "color": { "mode": "palette-classic" }, "unit": "reqps" }, "overrides": [] },
      "options": { "legend": { "displayMode": "list", "placement": "bottom" }, "tooltip": { "mode": "multi" } }
    },
    {
      "title": "Variables by Quality (total / good / bad)",
      "type": "timeseries",
      "gridPos": { "x": 16, "y": 14, "w": 8, "h": 8 },
      "datasource": { "type": "prometheus", "uid": "prometheus" },
      "targets": [
        { "expr": "idako_collector_all_variables", "legendFormat": "Total", "refId": "A" },
        { "expr": "idako_collector_good_variables", "legendFormat": "Good", "refId": "B" },
        { "expr": "idako_collector_bad_variables", "legendFormat": "Bad", "refId": "C" }
      ],
      "fieldConfig": {
        "defaults": { "custom": { "drawStyle": "line", "lineWidth": 2, "fillOpacity": 10 },
          "color": { "mode": "palette-classic" } },
        "overrides": [
          { "matcher": { "id": "byName", "options": "Total" },
            "properties": [{ "id": "color", "value": { "mode": "fixed", "fixedColor": "blue" } }] },
          { "matcher": { "id": "byName", "options": "Good" },
            "properties": [{ "id": "color", "value": { "mode": "fixed", "fixedColor": "green" } }] },
          { "matcher": { "id": "byName", "options": "Bad" },
            "properties": [{ "id": "color", "value": { "mode": "fixed", "fixedColor": "red" } }] }
        ]
      },
      "options": { "legend": { "displayMode": "list", "placement": "bottom" }, "tooltip": { "mode": "multi" } }
    }
  ]
}
Multiple OPC UA servers? These panels use the fleet-wide counts, which need no changes as servers are added or removed. To chart a specific server by name, add a panel with idako_collector_server_up{server="Your Server Name"}.
4

Set up email alerts

Three pieces, provisioned the same way as the dashboard: where to send email from, who receives it, and what conditions trigger it.

SMTP credentials — the one file you actually create yourself

Grafana needs real SMTP relay details to send anything, and that includes a password — which is exactly why it isn't part of the download. .env-example is the safe-to-share template already sitting in your folder; .env is the real file you make from it:

.env-example
# Copy this file to .env before starting the stack:
#   cp .env-example .env
# Then fill in your real SMTP relay details below. .env is gitignored and
# stays local to this machine only — never commit real credentials to Git.

# Required for Grafana's email alerts to actually send. Use your
# organization's own SMTP relay, or a transactional email provider
# (SendGrid, Mailgun, Amazon SES, etc.)
GF_SMTP_ENABLED=true
GF_SMTP_HOST=smtp.example.com:587
GF_SMTP_USER=alerts@example.com
GF_SMTP_PASSWORD=your-smtp-password
GF_SMTP_FROM_ADDRESS=alerts@example.com
GF_SMTP_FROM_NAME=Idako Monitoring
shell
cp .env-example .env
# then edit .env with a text editor and fill in your real values
.gitignore
.env
Changed .env after the stack is already running? A plain docker compose restart grafana is not enough — Grafana's container keeps whatever environment it was originally created with. Use docker compose up -d --force-recreate grafana to actually pick up new values.

Who receives the alerts

This file controls where alert emails go. The only change needed is the placeholder address:

grafana/provisioning/alerting/contactpoints.yaml
apiVersion: 1

contactPoints:
  - orgId: 1
    name: idako-email
    receivers:
      - uid: idako-email-receiver
        type: email
        settings:
          addresses: you@example.com
          singleEmail: true

Replace you@example.com with your real recipient — a comma-separated list works if more than one person should be notified.

Routing — send everything to that one address

This tells Grafana to route every alert to the contact point above — nothing to change here:

grafana/provisioning/alerting/policies.yaml
apiVersion: 1

policies:
  - orgId: 1
    receiver: idako-email
    group_by: ["alertname"]

What triggers an alert

Six conditions, one per subsystem plus two data-quality checks. Every rule queries Prometheus directly and routes through the contact point above via the default policy:

grafana/provisioning/alerting/rules.yaml
apiVersion: 1

groups:
  - orgId: 1
    name: idako-alerts
    folder: Idako
    interval: 1m
    rules:
      - uid: idako-instance-not-ok
        title: Idako instance is not OK
        condition: C
        for: 2m
        labels:
          severity: critical
        annotations:
          summary: "Idako instance is not reachable (Prometheus scrape of the /health endpoint is failing)"
        noDataState: Alerting
        execErrState: Alerting
        data:
          - refId: A
            relativeTimeRange: { from: 300, to: 0 }
            datasourceUid: prometheus
            model:
              expr: up{job="idako"}
              instant: true
              intervalMs: 1000
              maxDataPoints: 43200
              refId: A
          - refId: C
            datasourceUid: "__expr__"
            model:
              type: threshold
              expression: A
              conditions:
                - evaluator:
                    type: lt
                    params: [1]
              refId: C

      - uid: idako-buffer-not-ok
        title: Local Buffer is down
        condition: C
        for: 2m
        labels:
          severity: critical
        annotations:
          summary: "Idako's local buffer subsystem is not running (idako_buffer_up is 0)"
        noDataState: Alerting
        execErrState: Alerting
        data:
          - refId: A
            relativeTimeRange: { from: 300, to: 0 }
            datasourceUid: prometheus
            model:
              expr: idako_buffer_up
              instant: true
              intervalMs: 1000
              maxDataPoints: 43200
              refId: A
          - refId: C
            datasourceUid: "__expr__"
            model:
              type: threshold
              expression: A
              conditions:
                - evaluator:
                    type: lt
                    params: [1]
              refId: C

      - uid: idako-tsdb-not-ok
        title: TSDB is down
        condition: C
        for: 2m
        labels:
          severity: critical
        annotations:
          summary: "Idako's connection to the time-series database is not OK (idako_tsdb_up is 0)"
        noDataState: Alerting
        execErrState: Alerting
        data:
          - refId: A
            relativeTimeRange: { from: 300, to: 0 }
            datasourceUid: prometheus
            model:
              expr: idako_tsdb_up
              instant: true
              intervalMs: 1000
              maxDataPoints: 43200
              refId: A
          - refId: C
            datasourceUid: "__expr__"
            model:
              type: threshold
              expression: A
              conditions:
                - evaluator:
                    type: lt
                    params: [1]
              refId: C

      - uid: idako-server-disconnected
        title: OPC UA server disconnected
        condition: C
        for: 2m
        labels:
          severity: warning
        annotations:
          summary: "OPC UA server {{ $labels.server }} ({{ $labels.endpoint }}) is disconnected"
        noDataState: OK
        execErrState: Alerting
        data:
          - refId: A
            relativeTimeRange: { from: 300, to: 0 }
            datasourceUid: prometheus
            model:
              expr: idako_collector_server_up
              instant: true
              intervalMs: 1000
              maxDataPoints: 43200
              refId: A
          - refId: C
            datasourceUid: "__expr__"
            model:
              type: threshold
              expression: A
              conditions:
                - evaluator:
                    type: lt
                    params: [1]
              refId: C

      - uid: idako-bad-variables
        title: Bad variables detected
        condition: C
        for: 2m
        labels:
          severity: warning
        annotations:
          summary: "{{ $values.A }} variable(s) are reporting bad quality"
        noDataState: Alerting
        execErrState: Alerting
        data:
          - refId: A
            relativeTimeRange: { from: 300, to: 0 }
            datasourceUid: prometheus
            model:
              expr: idako_collector_bad_variables
              instant: true
              intervalMs: 1000
              maxDataPoints: 43200
              refId: A
          - refId: C
            datasourceUid: "__expr__"
            model:
              type: threshold
              expression: A
              conditions:
                - evaluator:
                    type: gt
                    params: [0]
              refId: C

      - uid: idako-forwarding-behind-collection
        title: TSDB forwarding is falling behind collection
        condition: D
        for: 5m
        labels:
          severity: warning
        annotations:
          summary: "Forwarding rate is more than 10% slower than the collection rate — the buffer backlog is likely growing"
        noDataState: OK
        execErrState: Alerting
        data:
          - refId: A
            relativeTimeRange: { from: 300, to: 0 }
            datasourceUid: prometheus
            model:
              expr: idako_collector_rate
              instant: true
              intervalMs: 1000
              maxDataPoints: 43200
              refId: A
          - refId: B
            relativeTimeRange: { from: 300, to: 0 }
            datasourceUid: prometheus
            model:
              expr: idako_tsdb_values_forwarded_rate
              instant: true
              intervalMs: 1000
              maxDataPoints: 43200
              refId: B
          - refId: C
            datasourceUid: "__expr__"
            model:
              type: math
              expression: "($A - $B) / $A"
              refId: C
          - refId: D
            datasourceUid: "__expr__"
            model:
              type: threshold
              expression: C
              conditions:
                - evaluator:
                    type: gt
                    params: [0.1]
              refId: D
RuleFires whenSeverity
Idako instance is not OKup{job="idako"} is 0 for 2mcritical
Local Buffer is downidako_buffer_up is 0 for 2mcritical
TSDB is downidako_tsdb_up is 0 for 2mcritical
OPC UA server disconnectedidako_collector_server_up is 0 for 2m — one alert per server, named in the emailwarning
Bad variables detectedidako_collector_bad_variables > 0 for 2mwarning
TSDB forwarding is falling behind collectionforwarding rate < 90% of collection rate for 5mwarning
Why up{job="idako"} and not a custom Idako metric? This is Prometheus's own, built-in signal for "could I reach this target at all" — it gets a fresh value on every single scrape attempt, success or failure, so it detects a fully unreachable instance within one scrape interval. A metric Idako itself produces (like idako_up) simply stops updating when Idako is unreachable, which is a much slower and less reliable way to notice a total outage.
Both firing and resolved notifications are sent — that's Grafana's default (a per-contact-point "Disable resolved message" option turns the second one off, if you'd rather only hear about new problems). Firing waits out the rule's for duration to avoid paging on a blip; resolved notifications go out on the next evaluation after the condition clears, no equivalent delay.
5

The Docker Compose file

docker-compose.yml wires everything together: Prometheus reads its config from step 2, Grafana reads its provisioning from steps 3–4 and its SMTP settings from .env, and both get a named volume so data survives a restart. Nothing to change here — it's already set up.

docker-compose.yml
services:
  prometheus:
    image: prom/prometheus:latest
    container_name: idako-prometheus
    restart: unless-stopped
    volumes:
      - ./prometheus/prometheus.yml:/etc/prometheus/prometheus.yml:ro
      - prometheus-data:/prometheus
    ports:
      - "9090:9090"

  grafana:
    image: grafana/grafana:latest
    container_name: idako-grafana
    restart: unless-stopped
    depends_on:
      - prometheus
    volumes:
      - ./grafana/provisioning:/etc/grafana/provisioning:ro
      - grafana-data:/var/lib/grafana
    ports:
      - "3000:3000"
    env_file:
      - .env

volumes:
  prometheus-data:
  grafana-data:
.env must exist before this will start. env_file: - .env means Docker Compose refuses to start the grafana service if that file is missing — make sure step 4's cp .env-example .env happened first.
6

Start the stack

From inside the idako-monitoring folder:

shell
docker compose up -d

Confirm both containers are running:

shell
docker compose ps

Then confirm Prometheus can actually reach Idako — open http://localhost:9090/targets in a browser. The idako target should show State: UP. If it shows DOWN, the error message next to it almost always names the problem — usually the host/port in step 2, or a firewall between the Docker host and Idako.

7

Open the dashboard

Go to http://localhost:3000 to sign in.

Default credentials: admin / admin. Grafana ships with this login out of the box and prompts you to set a real password the moment you sign in — do that immediately, especially if port 3000 is reachable from beyond your own machine.

In the left menu, go to Dashboards — the Idako folder and the Idako Overview dashboard inside it were created automatically by the files from step 3. Open it — it should look like this:

The Idako Overview dashboard in Grafana: four green OK status tiles, a connected-server count, and six live charts showing buffer, tsdb, and variable-quality trends.
The Idako Overview dashboard, running against a healthy instance.
You should see green OK tiles for Instance, Collector, Local Buffer, and TSDB Status, a count of connected servers, and five live charts. If everything reads zero or empty, double-check the target is UP on the Prometheus targets page from step 6 first — an unreachable target is the most common cause.
Beyond one instance

Monitoring more than one Idako instance

Add one more entry to targets in prometheus/prometheus.yml — no changes needed anywhere else, including the dashboard and alert rules, since they already aggregate across whatever Prometheus is scraping:

prometheus/prometheus.yml (excerpt)
    static_configs:
      - targets:
          - "idako-host-1:4880"
          - "idako-host-2:4880"

Restart Prometheus to pick up the change: docker compose restart prometheus.

Before you trust it

Verifying the alerts actually work

A clean startup only proves Grafana accepted the SMTP settings, not that mail actually gets delivered. Two checks worth doing once, right after setup:

  • Send a test email — in Grafana, go to Alerting → Contact points, open idako-email, and use Test. This confirms your SMTP credentials and relay are correct in isolation, before any real alert depends on them.
  • Trigger a real one — stop Idako (or block the port) and wait a few minutes. "Idako instance is not OK" should reach Alerting → Active notifications within about 2–3 minutes, and land in your inbox shortly after. Bring Idako back and you should get a second, resolved email on the next evaluation.
SMTP authentication failing? The most common cause after a config change isn't a wrong password — it's a stale container. Confirm what's actually loaded with docker exec idako-grafana printenv | grep GF_SMTP and compare against your current .env; if they differ, that's the --force-recreate step from earlier being skipped, not a credentials problem.
Next steps

Where to go from here

More alert conditionsAdd another rule to the same rules.yaml group — e.g. a warning specifically for a growing idako_buffer_values_balance, before it turns into lost data.
Route by severityAdd a second contact point (e.g. a chat webhook) and a policy that matches on the severity: critical label already set on three of the six rules, so only the urgent ones page immediately.
Distinguish "down" from "metrics are broken"A blackbox_exporter container probing plain GET /health catches the narrower case where Idako is healthy but its Prometheus output specifically is malformed — up{job="idako"} alone can't tell those apart.
Keep history longer than Prometheus's defaultPrometheus's local storage is fine for weeks of data; for longer retention, point it at a remote-write target instead of changing anything on the Idako side.
See exactly what Idako reportsThe full field-by-field reference for /health, including every metric this dashboard and these alerts are built from.
This stack talks to Idako over plain HTTP on its health port — no credentials, no changes to Idako required. Everything here runs on infrastructure you control.