Fault monitoring: can we share ideas?

mgiammarco

Renowned Member
Feb 18, 2010
165
10
83
Hello,
I would like to monitor all my proxmox/ceph installations:
- check that backups start;
- check errors after backup run;
- check if ceph osd is down;
- check if some ha vm is in error state

I plan to use influxdb with telegraf or opendistro for elasticsearch or others.
I started with influxdb with telegraf:
- there is a proxmox plugin for telegraf;
- there is official support by ceph for influxdb and telegraf for sending metrics

After one day:
- ceph support is very undocumented/buggy/error prone
- proxmox and ceph metrics are not useful at all to quickly check for problems above

I now plan to send logs to elasticsearch then parse logs to get alerts.
Can we share some ideas/help/tricks to reach this goal?
Thanks,
Mario
 

About

The Proxmox community has been around for many years and offers help and support for Proxmox VE, Proxmox Backup Server, and Proxmox Mail Gateway.
We think our community is one of the best thanks to people like you!

Get your subscription!

The Proxmox team works very hard to make sure you are running the best software and getting stable updates and security enhancements, as well as quick enterprise support. Tens of thousands of happy customers have a Proxmox subscription. Get yours easily in our online shop.

Buy now!