Tidak Ada Deskripsi

Stu Bevan 7e667cdcaa Report a rebuilding array as a warning rather than OK 10 tahun lalu
.gitignore 29232e4f99 add .gitignore 12 tahun lalu
LICENSE 3451ba2e85 Initial commit 12 tahun lalu
README.md d723d6c97a support arguments in crashplan script 11 tahun lalu
check_bro.sh 106e6efc2d correct check_bro.sh to ignore 'Getting status' messages 11 tahun lalu
check_connections.sh 8cf91d3866 update permissions for new scripts 12 tahun lalu
check_crashplan_backup.py d723d6c97a support arguments in crashplan script 11 tahun lalu
check_enq.sh c899d341ec check_enq.sh: fix bug when more than two queues are specified 12 tahun lalu
check_file_growth.sh c725c68344 First commit, including plugins in repository 12 tahun lalu
check_filesystem_stat.sh ba5cf0778a add filesystem stat script 12 tahun lalu
check_load.sh be03b0fa21 This change seems more flexible and rely on the assumption that it comes after the last : 11 tahun lalu
check_memory.sh dcf85c1939 check_memory.sh: add plugin to check memory usage on Linux 12 tahun lalu
check_nfs_stale d51a15a89f add new plug-in check_nfs_stale 12 tahun lalu
check_ossec.sh 9c216e2d0e check_ossec.sh: add option to exclude multiple services 12 tahun lalu
check_osx_raid.sh 7e667cdcaa Report a rebuilding array as a warning rather than OK 10 tahun lalu
check_osx_smart.sh c725c68344 First commit, including plugins in repository 12 tahun lalu
check_osx_temp.sh c725c68344 First commit, including plugins in repository 12 tahun lalu
check_pps.sh d3df576bb8 check_pps.sh: warning and exit when sysfs and procfs are not present 12 tahun lalu
check_rsyslog.sh 8cf91d3866 update permissions for new scripts 12 tahun lalu
check_service.sh f9becaee43 Add sudo for linux + switch user variable 11 tahun lalu
check_traffic.sh b5e16c897f add check_traffic plugin example 11 tahun lalu
check_volume.sh c725c68344 First commit, including plugins in repository 12 tahun lalu
negate.sh 29569d9087 rename check_status_code.sh to negate.sh 12 tahun lalu

README.md

nagios-plugins

A collection of Nagios Plugins I've written for production unix environments.

A few of the scripts like check_service.sh and check_volume.sh are designed to be run in heterogenous unix environments and should work on Linux, OSX, AIX, and the BSD's provided a bash or bash-compatible shell to interpret them.

Each script has detailed usage presentable via the -h option and some of scripts include extended usage examples within the top commented section of the script.

One way to run the plugins requiring elevated privileges is to configure sudo on each monitored machine to allow the nagios user to execute the plug-ins in the plug-in directory as root:

$ visudo # Use visudo to edit the sudoers file , or :
$ echo ’Defaults:nagios !requiretty ’ >> /etc/sudoers
$ echo ’nagios ALL=(root) NOPASSWD:/usr/local/nagios/libexec/∗’ >> /etc/sudoers

If that is the case be sure to limit write permissions for the scripts so that one cannot simply update the scripts with malicious code.

Plugins:

Heterogenous Unix (Unices):

check_load.sh - Check a system's load (run queue) via ``uptime''.

check_service.sh - Check the status of a system service.

check_volume.sh - Check free space for a volume or partition.

check_file_growth.sh - Check whether a file is growing in size (e.g. Monitor for stale log files).

check_filesystem_stat.sh - Recursively checks for filesystem input/output errors by directory using stat.

check_traffic.sh - Check rate of traffic type by bpf using tcpdump for interface

negate.sh - Checks exit code of another program and returns a custom Nagios status code based on the result.

OSX only:

check_osx_raid.sh - Check RAID status of a disk. (Find degraded and failing arrays).

check_osx_smart.sh - Check S.M.A.R.T status of a disk. (Find failing disks).

check_osx_temp.sh - Check temperature of system components. (Find systems running hot).

Linux only:

check_connections.sh - Check number of connections/sockets in a given state. Requires iproute2. (Find 10000+ EST connections)

check_pps.sh - Check PPS, BPS, or percentage of line-rate on a networking interface (Find periods of heavy traffic)

Application specific:

check_ossec.sh - Perform multiple checks for a OSSEC server (e.g. Find a disconnected agent).

check_bro.sh - Perform multiple checks for a Bro cluster (e.g. Find stopped workers).

check_enq.sh - Check status of a printer queue on AIX (Find queues in DOWN state)

check_rsyslog.sh - Check for rsyslog disk queue buffers (Find when logs are buffered)

check_crashplan_backup.py - Check latest backup times for crashplan server (uses API), notifies if backups hasn't been completed in 48 hours by default

  1. Add credentials to text file which script will read: printf 'user = admin@company.com\npassword = ChangeMe\n' > /root/crashplan-credentials-for-nagios.txt
  2. Use a. Check all hosts: check_crashplan_backup.py /root/crashplan_creds.txt crashplan.company.com:4285 b. Check single host: check_crashplan_backup.py /root/crashplan_creds.txt crashplan.company.com:4285 server1.company.com
  3. Use sudo when exucuting script with nagios if you put credentials in a file not readable by nagios user
  4. You can edit the max_backup_time variable in the script to adjust the max number of days w/o backup for critical