- Merge Infrastructure + Secure + Systems → SNS Networking (Business #1) - Merge Web + Software + Cloud → SNS Digital (planned) - SNS Support unchanged (planned) - Add infra/ folder with 16 FOSS-first buildable designs - Update all agent knowledge, division briefs, legal structure - Restructure businesses/ from 7 to 3 operating folders
6.2 KiB
6.2 KiB
Monitoring & Alerting — SNS Support
Entity: SNS Support · Status: Buildable now
What it is
Centralized monitoring for all infrastructure — on-prem and cloud. Know when something is down before anyone reports it. This is what makes managed services possible.
Stack
| Component | Software | FOSS | Role |
|---|---|---|---|
| Metrics collection | Prometheus | Yes | Scrapes targets, stores time-series |
| Visualization | Grafana | Yes | Dashboards, alerting UI |
| Node metrics | node_exporter | Yes | CPU, RAM, disk, network per host |
| Container metrics | cAdvisor | Yes | Docker/LXC resource usage |
| Uptime / HTTP checks | Uptime Kuma | Yes | External endpoint monitoring, status pages |
| Alerting | Grafana Alerting | Yes | Routes alerts to email/Ntfy/Slack |
| Notifications | Ntfy | Yes | Self-hosted push notifications (phone/desktop) |
Architecture
┌────────────┐ ┌────────────┐ ┌────────────┐ ┌────────────┐
│ Host A │ │ Host B │ │ Host C │ │ Linode │
│ node_exp │ │ node_exp │ │ node_exp │ │ node_exp │
└─────┬──────┘ └─────┬──────┘ └─────┬──────┘ └─────┬──────┘
│ │ │ │
└────────────────┼────────────────┼────────────────┘
▼
┌─────────────────┐
│ Prometheus │ ← VM on Proxmox
│ (scrapes all) │
└────────┬────────┘
│
┌────────┴────────┐
│ Grafana │
│ - Dashboards │
│ - Alert rules │
└────────┬────────┘
│
┌────────┴────────┐
│ Ntfy │ ← Push notifications to phone
└─────────────────┘
┌─────────────────┐
│ Uptime Kuma │ ← Separate LXC, checks from outside
│ (HTTP checks) │
└─────────────────┘
What Gets Monitored
| Category | Metrics | Alert threshold |
|---|---|---|
| Host health | CPU, RAM, disk, load | Disk > 85%, RAM > 90%, load > 4x cores |
| Service up/down | systemd unit status | Any critical service down > 2 min |
| Network | Interface errors, bandwidth | Sustained > 80% capacity |
| HTTP endpoints | Response time, status code | > 5s response or non-2xx |
| SSL certs | Days until expiry | < 14 days (Caddy should auto-renew, this catches failures) |
| Backups | Last successful PBS backup age | > 26 hours (missed nightly window) |
| Docker | Container restarts, OOM kills | Any restart loop or OOM |
| WireGuard | Peer last handshake age | > 5 minutes (tunnel down) |
Build Steps
-
Monitoring VM: Clone Debian 12 template. 2 vCPU, 4GB RAM, 50GB disk.
-
Prometheus + Grafana (Docker Compose):
services: prometheus: image: prom/prometheus:latest volumes: - ./prometheus.yml:/etc/prometheus/prometheus.yml - prom-data:/prometheus command: --storage.tsdb.retention.time=90d grafana: image: grafana/grafana:latest volumes: [grafana-data:/var/lib/grafana] environment: GF_SECURITY_ADMIN_PASSWORD__FILE: /run/secrets/grafana_pw -
node_exporter on every host:
apt install prometheus-node-exporter # Or binary install + systemd unit # Listens on :9100, scraped by Prometheus -
Prometheus scrape config:
# prometheus.yml scrape_configs: - job_name: nodes static_configs: - targets: - gitea.internal.sns:9100 - caddy.internal.sns:9100 - linode.wg.internal:9100 # ... all hosts -
Uptime Kuma (separate LXC):
docker run -d --name uptime-kuma -p 3001:3001 \ -v uptime-kuma:/app/data louislam/uptime-kumaAdd monitors for all public endpoints + internal services.
-
Ntfy (self-hosted push):
docker run -d --name ntfy -p 8080:80 \ -v ntfy-cache:/var/cache/ntfy binwiederhier/ntfy serveGrafana alert contact point → Ntfy webhook → phone notification.
-
Grafana dashboards: Import community dashboards:
- Node Exporter Full (ID: 1860)
- Docker/cAdvisor (ID: 14282)
- Custom: SNS Overview (service status grid)
Security Posture
- Prometheus/Grafana: Behind Authelia 2FA via Caddy. Not exposed publicly.
- node_exporter: Binds to service VLAN IP only. No internet-facing metrics endpoints.
- Uptime Kuma: Can optionally expose a public status page per-client (read-only).
- Ntfy: Private topic names. No unauthenticated publish.
- Data retention: 90 days for metrics. Enough for trend analysis without filling disks.
Upgrade Path
- Thanos / Mimir: When Prometheus needs long-term storage or multi-cluster federation. YAGNI until metrics volume overwhelms a single Prometheus instance.
- VictoriaMetrics: Drop-in Prometheus replacement with better compression and performance if retention or scrape targets grow significantly.
- PagerDuty / Opsgenie: When you have SLA contracts that need formal on-call rotation. Until then, Ntfy to your phone is sufficient.