# Monitoring & Alerting — SNS Support **Entity:** SNS Support · **Status:** Buildable now ## What it is Centralized monitoring for all infrastructure — on-prem and cloud. Know when something is down before anyone reports it. This is what makes managed services possible. ## Stack | Component | Software | FOSS | Role | |-----------|----------|------|------| | Metrics collection | Prometheus | Yes | Scrapes targets, stores time-series | | Visualization | Grafana | Yes | Dashboards, alerting UI | | Node metrics | node_exporter | Yes | CPU, RAM, disk, network per host | | Container metrics | cAdvisor | Yes | Docker/LXC resource usage | | Uptime / HTTP checks | Uptime Kuma | Yes | External endpoint monitoring, status pages | | Alerting | Grafana Alerting | Yes | Routes alerts to email/Ntfy/Slack | | Notifications | Ntfy | Yes | Self-hosted push notifications (phone/desktop) | ## Architecture ``` ┌────────────┐ ┌────────────┐ ┌────────────┐ ┌────────────┐ │ Host A │ │ Host B │ │ Host C │ │ Linode │ │ node_exp │ │ node_exp │ │ node_exp │ │ node_exp │ └─────┬──────┘ └─────┬──────┘ └─────┬──────┘ └─────┬──────┘ │ │ │ │ └────────────────┼────────────────┼────────────────┘ ▼ ┌─────────────────┐ │ Prometheus │ ← VM on Proxmox │ (scrapes all) │ └────────┬────────┘ │ ┌────────┴────────┐ │ Grafana │ │ - Dashboards │ │ - Alert rules │ └────────┬────────┘ │ ┌────────┴────────┐ │ Ntfy │ ← Push notifications to phone └─────────────────┘ ┌─────────────────┐ │ Uptime Kuma │ ← Separate LXC, checks from outside │ (HTTP checks) │ └─────────────────┘ ``` ## What Gets Monitored | Category | Metrics | Alert threshold | |----------|---------|-----------------| | **Host health** | CPU, RAM, disk, load | Disk > 85%, RAM > 90%, load > 4x cores | | **Service up/down** | systemd unit status | Any critical service down > 2 min | | **Network** | Interface errors, bandwidth | Sustained > 80% capacity | | **HTTP endpoints** | Response time, status code | > 5s response or non-2xx | | **SSL certs** | Days until expiry | < 14 days (Caddy should auto-renew, this catches failures) | | **Backups** | Last successful PBS backup age | > 26 hours (missed nightly window) | | **Docker** | Container restarts, OOM kills | Any restart loop or OOM | | **WireGuard** | Peer last handshake age | > 5 minutes (tunnel down) | ## Build Steps 1. **Monitoring VM:** Clone Debian 12 template. 2 vCPU, 4GB RAM, 50GB disk. 2. **Prometheus + Grafana (Docker Compose):** ```yaml services: prometheus: image: prom/prometheus:latest volumes: - ./prometheus.yml:/etc/prometheus/prometheus.yml - prom-data:/prometheus command: --storage.tsdb.retention.time=90d grafana: image: grafana/grafana:latest volumes: [grafana-data:/var/lib/grafana] environment: GF_SECURITY_ADMIN_PASSWORD__FILE: /run/secrets/grafana_pw ``` 3. **node_exporter on every host:** ```bash apt install prometheus-node-exporter # Or binary install + systemd unit # Listens on :9100, scraped by Prometheus ``` 4. **Prometheus scrape config:** ```yaml # prometheus.yml scrape_configs: - job_name: nodes static_configs: - targets: - gitea.internal.sns:9100 - caddy.internal.sns:9100 - linode.wg.internal:9100 # ... all hosts ``` 5. **Uptime Kuma (separate LXC):** ```bash docker run -d --name uptime-kuma -p 3001:3001 \ -v uptime-kuma:/app/data louislam/uptime-kuma ``` Add monitors for all public endpoints + internal services. 6. **Ntfy (self-hosted push):** ```bash docker run -d --name ntfy -p 8080:80 \ -v ntfy-cache:/var/cache/ntfy binwiederhier/ntfy serve ``` Grafana alert contact point → Ntfy webhook → phone notification. 7. **Grafana dashboards:** Import community dashboards: - Node Exporter Full (ID: 1860) - Docker/cAdvisor (ID: 14282) - Custom: SNS Overview (service status grid) ## Security Posture - **Prometheus/Grafana:** Behind Authelia 2FA via Caddy. Not exposed publicly. - **node_exporter:** Binds to service VLAN IP only. No internet-facing metrics endpoints. - **Uptime Kuma:** Can optionally expose a public status page per-client (read-only). - **Ntfy:** Private topic names. No unauthenticated publish. - **Data retention:** 90 days for metrics. Enough for trend analysis without filling disks. ## Upgrade Path - **Thanos / Mimir:** When Prometheus needs long-term storage or multi-cluster federation. YAGNI until metrics volume overwhelms a single Prometheus instance. - **VictoriaMetrics:** Drop-in Prometheus replacement with better compression and performance if retention or scrape targets grow significantly. - **PagerDuty / Opsgenie:** When you have SLA contracts that need formal on-call rotation. Until then, Ntfy to your phone is sufficient.