Running services, network usage, memory usage, bandwidth, disk I/O, successful logins, whether the thing is even alive, etc…
So far my only method has been “hope and pray”.
Running services, network usage, memory usage, bandwidth, disk I/O, successful logins, whether the thing is even alive, etc…
So far my only method has been “hope and pray”.
Uptime Kuma for outages, Prometheus for metrics, rsyslog to Loki backend for logs. Grafana for ingesting Prometheus and Loki data.