How to Write a Linux Health Check Script (With Examples)

发布时间:2026/10/8 23:46:33
How to Write a Linux Health Check Script (With Examples) Are you looking to create custom health check scripts for your Linux systems? Need practical, step-by-step instructions with real-world examples? This comprehensive guide covers everything you need to know about creating effective health check scripts for Linux, including essential metrics, bash scripting techniques, alerting, logging, and best practices to ensure optimal system performance and reliability.Stop checking this by handZuzia watches uptime, SSL, ports, DNS, cron jobs and server metrics for every client server and pings you the moment something breaks. Free plan, no credit card.Start monitoring — its freeIntroduction to Linux Health ChecksLinux health checks are automated scripts or processes that systematically evaluate your servers operational status, performance metrics, and overall condition. These checks examine system resources, service status, performance indicators, and security status to ensure everything is functioning correctly and identify potential issues before they escalate into critical problems.Health check scripts transform server management from reactive troubleshooting to proactive maintenance. By automating regular health checks, you can continuously monitor system status, detect problems early, and respond quickly to issues before they impact users or cause downtime. Effective health checks help maintain optimal performance, prevent resource exhaustion, and ensure reliable service delivery.The goal of Linux health check scripts is to provide automated, continuous monitoring that detects issues immediately and enables rapid response. By creating custom health check scripts tailored to your specific needs, you can monitor exactly what matters most for your infrastructure, integrate with your existing tools, and maintain complete control over your monitoring approach.Essential Metrics to MonitorUnderstanding which metrics to include in health checks ensures comprehensive system monitoring.CPU Usage MetricsMonitor CPU performance to detect bottlenecks:CPU Utilization: Overall processor usage percentage. Should typically stay below 70-80% under normal load. Sustained high CPU usage indicates potential bottlenecks.Load Average: System load over 1, 5, and 15 minutes. Load average should be below the number of CPU cores for optimal performance.CPU Wait Time: Time CPU spends waiting for I/O operations. High wait times suggest disk or network bottlenecks rather than CPU limitations.Top CPU Processes: Identify which processes consume the most CPU resources to pinpoint resource-intensive applications.Memory Usage MetricsMonitor memory to prevent out-of-memory conditions:RAM Usage: Total and available memory. Should maintain at least 10-20% free memory for optimal performance. High memory usage can cause swapping and performance degradation.Swap Usage: Virtual memory usage on disk. High swap usage indicates insufficient RAM. Excessive swapping dramatically impacts performance.Memory Pressure: How close the system is to memory limits. Monitor trends to predict when upgrades are needed.Memory Leaks: Processes with continuously increasing memory consumption. Early detection prevents memory exhaustion.Disk Usage MetricsMonitor disk to prevent storage exhaustion:Disk Space Usage: Available storage capacity. Maintain at least 15-20% free disk space. Running out of disk space can cause service failures and data loss.Disk I/O Performance: Read/write operations per second and latency. High I/O rates or latency indicate potential bottlenecks.Inode Usage: File system metadata capacity. Running out of inodes prevents file creation even when disk space is available.I/O Wait Time: CPU time spent waiting for disk I/O operations. High I/O wait suggests disk bottlenecks.Network Performance MetricsMonitor network to ensure connectivity:Bandwidth Usage: Network traffic volume relative to capacity. High utilization may indicate attacks or saturation.Network Latency: Response times for network requests. Should be under 100ms for local networks.Packet Loss: Percentage of packets lost during transmission. Should be near 0%.Connection Count: Active network connections. Unusually high counts may indicate attacks or misconfigurations.Service Status MetricsMonitor critical services:Service Status: Verify all critical services are running and enabled at boot.Process Count: Ensure expected number of processes are running.Port Availability: Verify services are listening on correct ports.Response Times: Services should respond within acceptable timeframes.Creating Your First Health Check ScriptFollow these step-by-step instructions to create your first health check script.Step 1: Create Script FileCreate a new bash script file:# Create script file nano /usr/local/bin/health-check.sh # Or use your preferred editor vim /usr/local/bin/health-check.shStep 2: Add Script HeaderStart with a proper script header:#!/bin/bash # Linux System Health Check Script # Description: Monitors CPU, memory, disk, and network metrics # Author: Your Name # Date: $(date %Y-%m-%d) # Set script options set -euo pipefail # Configuration THRESHOLD_CPU80 THRESHOLD_MEM85 THRESHOLD_DISK90 LOG_FILE/var/log/health-check.logStep 3: Add Logging FunctionCreate a logging function for consistent log output:# Logging function log_message() { local level$1 shift echo [$(date %Y-%m-%d %H:%M:%S)] [$level] $* | tee -a $LOG_FILE }Step 4: Add CPU Check FunctionCreate function to check CPU usage:# Check CPU usage check_cpu() { local cpu_usage$(top -bn1 | grep Cpu(s) | sed s/.*, *\([0-9.]*\)%* id.*/\1/ | awk {print 100 - $1}) local load_avg$(uptime | awk -Fload average: {print $2} | awk {print $1} | sed s/,//) local cpu_cores$(nproc) log_message INFO CPU Usage: ${cpu_usage}%, Load Average: ${load_avg}, CPU Cores: ${cpu_cores} if (( $(echo $cpu_usage $THRESHOLD_CPU | bc -l) )); then log_message WARNING CPU usage is ${cpu_usage}% (threshold: ${THRESHOLD_CPU}%) return 1 fi return 0 }Step 5: Add Memory Check FunctionCreate function to check memory usage:# Check memory usage check_memory() { local mem_total$(free | grep Mem | awk {print $2}) local mem_used$(free | grep Mem | awk {print $3}) local mem_percent$(awk BEGIN {printf \%.2f\, ($mem_used/$mem_total)*100}) local swap_used$(free | grep Swap | awk {print $3}) local swap_total$(free | grep Swap | awk {print $2}) log_message INFO Memory Usage: ${mem_percent}%, Swap Used: ${swap_used}KB if (( $(echo $mem_percent $THRESHOLD_MEM | bc -l) )); then log_message WARNING Memory usage is ${mem_percent}% (threshold: ${THRESHOLD_MEM}%) return 1 fi return 0 }Step 6: Add Disk Check FunctionCreate function to check disk usage:# Check disk usage check_disk() { local disk_usage$(df -h / | awk NR2 {print $5} | sed s/%//) log_message INFO Disk Usage: ${disk_usage}% if [ $disk_usage -gt $THRESHOLD_DISK ]; then log_message WARNING Disk usage is ${disk_usage}% (threshold: ${THRESHOLD_DISK}%) return 1 fi return 0 }Step 7: Add Main FunctionCreate main function that runs all checks:# Main health check function main() { log_message INFO Starting health check... local exit_code0 check_cpu || exit_code1 check_memory || exit_code1 check_disk || exit_code1 if [ $exit_code -eq 0 ]; then log_message INFO Health check completed successfully else log_message ERROR Health check detected issues fi return $exit_code } # Run main function main $Step 8: Make Script ExecutableMake the script executable:chmod x /usr/local/bin/health-check.shStep 9: Test the ScriptTest your script:# Run script manually /usr/local/bin/health-check.sh # Check log output tail -f /var/log/health-check.logStep 10: Schedule with CronSchedule the script to run automatically:# Edit crontab crontab -e # Add line to run every 5 minutes */5 * * * * /usr/local/bin/health-check.shAdvanced Health Check TechniquesEnhance your health check scripts with advanced features.Adding Email AlertsAdd email notification when issues are detected:# Email alert function send_alert() { local subjectHealth Check Alert: $1 local message$2 local recipientadminexample.com echo $message | mail -s $subject $recipient } # Modify check functions to send alerts check_cpu() { # ... existing code ... if (( $(echo $cpu_usage $THRESHOLD_CPU | bc -l) )); then log_message WARNING CPU usage is ${cpu_usage}% send_alert High CPU Usage CPU usage is ${cpu_usage}% (threshold: ${THRESHOLD_CPU}%) return 1 fi return 0 }Adding Webhook NotificationsSend alerts via webhook for integration with monitoring systems:# Webhook alert function send_webhook() { local message$1 local webhook_urlhttps://your-webhook-url.com/alert local payload$(cat EOF { server: $(hostname), message: $message, timestamp: $(date -Iseconds) } EOF ) curl -X POST -H Content-Type: application/json \ -d $payload \ $webhook_url }Adding Service Status ChecksCheck critical service status:# Check service status check_services() { local services(nginx mysql ssh) local failed_services() for service in ${services[]}; do if ! systemctl is-active --quiet $service; then failed_services($service) log_message ERROR Service $service is not running fi done if [ ${#failed_services[]} -gt 0 ]; then log_message ERROR Failed services: ${failed_services[*]} return 1 fi return 0 }Adding Disk I/O MonitoringMonitor disk I/O performance:# Check disk I/O check_disk_io() { if ! command -v iostat /dev/null; then log_message WARNING iostat not available, skipping disk I/O check return 0 fi local disk_util$(iostat -x 1 1 | grep -v ^$ | tail -n 4 | awk {print $NF} | head -1 | sed s/%//) if [ -n $disk_util ] [ $disk_util -gt 80 ]; then log_message WARNING Disk utilization is ${disk_util}% return 1 fi return 0 }Adding Network Connectivity ChecksCheck network connectivity:# Check network connectivity check_network() { local test_hosts(8.8.8.8 google.com) local failed_hosts() for host in ${test_hosts[]}; do if ! ping -c 1 -W 2 $host /dev/null; then failed_hosts($host) log_message WARNING Cannot reach $host fi done if [ ${#failed_hosts[]} -gt 0 ]; then log_message ERROR Network connectivity issues to: ${failed_hosts[*]} return 1 fi return 0 }Comprehensive Health Check Script ExampleComplete example combining all features:#!/bin/bash # Comprehensive Linux Health Check Script set -euo pipefail # Configuration THRESHOLD_CPU80 THRESHOLD_MEM85 THRESHOLD_DISK90 LOG_FILE/var/log/health-check.log ALERT_EMAILadminexample.com WEBHOOK_URLhttps://your-webhook-url.com/alert # Logging function log_message() { local level$1 shift echo [$(date %Y-%m-%d %H:%M:%S)] [$level] $* | tee -a $LOG_FILE } # Alert functions send_email() { echo $2 | mail -s Health Check Alert: $1 $ALERT_EMAIL } send_webhook() { local payload$(cat EOF {server: $(hostname), alert: $1, message: $2, timestamp: $(date -Iseconds)} EOF ) curl -s -X POST -H Content-Type: application/json -d $payload $WEBHOOK_URL /dev/null || true } # Check functions (include all from above) check_cpu() { ... } check_memory() { ... } check_disk() { ... } check_services() { ... } check_disk_io() { ... } check_network() { ... } # Main function main() { log_message INFO Starting comprehensive health check... local exit_code0 local alerts() check_cpu || { exit_code1; alerts(High CPU usage); } check_memory || { exit_code1; alerts(High memory usage); } check_disk || { exit_code1; alerts(High disk usage); } check_services || { exit_code1; alerts(Service failures); } check_disk_io || { exit_code1; alerts(Disk I/O issues); } check_network || { exit_code1; alerts(Network issues); } if [ ${#alerts[]} -gt 0 ]; then local alert_messageHealth check detected issues: ${alerts[*]} log_message ERROR $alert_message send_email System Issues Detected $alert_message send_webhook System Issues $alert_message else log_message INFO All health checks passed fi return $exit_code } main $Tools and Software for Linux MonitoringWhile custom scripts provide flexibility, monitoring tools complement health check scripts.Automated Monitoring SolutionsZuzia.app- Cloud-based automated monitoring:Automatic Host Metrics collectionContinuous 24/7 monitoringHistorical data storageIntelligent alertingDashboard visualizationComplements custom scripts with comprehensive monitoringNagios- Enterprise monitoring solution:Extensive plugin ecosystemFlexible alerting systemWeb-based interfaceCan execute custom health check scriptsIntegrates with existing scriptsZabbix- Open-source enterprise monitoring:Comprehensive monitoring capabilitiesAuto-discovery featuresCustom script execution supportAdvanced alerting and visualizationCan run health check scripts as external checksPrometheus Grafana- Time-series monitoring:Powerful query languageHighly customizable dashboardsScript exporter for custom metricsCan collect metrics from health check scriptsExcellent for technical teamsCommand-Line ToolsUse these tools within health check scripts:top/htop: Process and CPU monitoringfree: Memory usage informationdf: Disk space usageiostat: Disk I/O statisticsvmstat: System-wide statisticssystemctl: Service status checkingss/netstat: Network connection monitoringBest Practices for Health Check ScriptsFollow these best practices to maintain effective health check scripts.Keep Scripts Simple and FocusedSingle responsibility: Each function should check one metricClear logic: Use straightforward conditional logicReadable code: Add comments explaining complex logicModular design: Break scripts into reusable functionsSimple scripts are easier to maintain, debug, and modify as requirements change.Use Proper Error HandlingSet error options: Useset -euo pipefailfor strict error handlingCheck command availability: Verify tools exist before using themHandle failures gracefully: Dont let one failed check stop othersReturn appropriate exit codes: Use exit codes to indicate success/failureProper error handling ensures scripts run reliably and provide accurate results.Implement LoggingStructured logging: Use consistent log format with timestampsLog levels: Use INFO, WARNING, ERROR levels appropriatelyLog rotation: Implement log rotation to prevent disk fillingCentralized logging: Consider sending logs to centralized systemGood logging enables troubleshooting and provides audit trail of system health.Set Appropriate ThresholdsBaseline first: Monitor for 1-2 weeks to establish baselinesAdjust thresholds: Fine-tune based on actual usage patternsDifferent thresholds: Use different thresholds for different serversDocument reasoning: Document why thresholds are set as they areAppropriate thresholds reduce false positives while ensuring critical issues are detected.Test Scripts ThoroughlyTest in non-production: Test scripts in development environments firstTest edge cases: Test with high/low values, missing tools, etc.Test failure scenarios: Verify alerts work when issues are detectedRegular review: Periodically review and update scriptsThorough testing ensures scripts work correctly in all scenarios.Document Your ScriptsHeader comments: Include purpose, author, date in script headerFunction documentation: Document what each function doesConfiguration documentation: Document threshold values and their reasoningUsage instructions: Document how to run and schedule scriptsDocumentation helps maintain scripts and enables others to understand and modify them.Integrate with Monitoring ToolsComplement, dont replace: Use scripts to complement automated monitoringExport metrics: Format output for monitoring tools to consumeUse monitoring APIs: Integrate with monitoring tool APIs when possibleLeverage best of both: Combine custom scripts with automated monitoringIntegration provides comprehensive monitoring combining custom checks with automated solutions.ConclusionCreating effective health check scripts for Linux systems enables proactive monitoring and rapid problem detection. By understanding essential metrics, following step-by-step scripting guides, implementing advanced techniques, and following best practices, you can create custom health check scripts that monitor exactly what matters for your infrastructure.Key TakeawaysStart simple: Begin with basic checks and gradually add complexityMonitor essentials: Focus on CPU, memory, disk, and network metricsImplement logging: Add proper logging for troubleshootingSet up alerts: Configure alerts for immediate notification of issuesTest thoroughly: Test scripts before deploying to productionDocument everything: Maintain documentation for future maintenanceIntegrate with tools: Combine custom scripts with automated monitoring solutionsNext StepsCreate basic script: Start with simple CPU, memory, and disk checksAdd logging: Implement structured loggingSet up alerts: Add email or webhook notificationsSchedule execution: Use cron to run scripts automaticallyTest and refine: Test scripts and adjust thresholds based on resultsIntegrate monitoring: Complement scripts with automated monitoring like Zuzia.appRemember, health check scripts are most effective when combined with comprehensive monitoring solutions. Use custom scripts for specific checks while leveraging automated monitoring for continuous coverage.For more information on Linux monitoring, explore related guides on server health checks, Linux resource monitoring, and server monitoring.Related guides, recipes, and problemsGuides:Understanding Server Health ChecksLinux Resource Monitoring Best PracticesServer Performance Monitoring Best PracticesServer Resource Monitoring Complete GuideComplete Guide to Server Monitoring on LinuxRecipes:How to Check System Health StatusMonitor System Services StatusMonitor Process Resource UsageProblems:High CPU Usage on Linux ServerHigh Memory Usage on Linux ServerServer Slow Performance IssuesFAQ: Common Questions About Health Check ScriptsWhat are health check scripts in Linux?Health check scripts are automated bash scripts that systematically evaluate your Linux servers operational status, performance metrics, and overall condition. They check system resources (CPU, memory, disk, network), service status, performance indicators, and security status to ensure everything is functioning correctly and identify potential issues before they escalate into critical problems.Health check scripts typically run on a schedule (via cron) and log results, send alerts when issues are detected, and provide visibility into system health. They complement automated monitoring solutions by providing custom checks tailored to your specific needs.Why are health checks important for Linux systems?Health checks are important because they:Prevent downtime: Detect problems early before they cause service outagesMaintain performance: Monitor performance metrics continuously to maintain optimal performanceEnable rapid response: Provide immediate alerts when issues are detectedSupport proactive maintenance: Transform management from reactive to proactiveOptimize resources: Identify resource bottlenecks and optimization opportunitiesWithout health checks, problems are discovered only after they impact users, leading to emergency fixes, increased costs, and potential data loss. Health checks enable proactive problem detection and rapid response.How can I automate health checks on my Linux server?Automate health checks by:Create health check script: Write bash script with checks for CPU, memory, disk, and servicesAdd logging: Implement logging to track health check resultsConfigure alerts: Add email or webhook notifications for issuesSchedule with cron: Usecrontab -eto schedule script execution (e.g., every 5 minutes)Test thoroughly: Test script in non-production environment firstMonitor logs: Review logs regularly to verify script is working correctlyExample cron entry:# Run health check every 5 minutes */5 * * * * /usr/local/bin/health-check.shFor comprehensive automated monitoring, complement custom scripts with automated solutions like Zuzia.app that provide continuous monitoring with minimal maintenance.

关于本文作者

来自尧图内容编辑团队

尧图内容编辑团队 内容团队

尧图内容编辑团队

本文由尧图网络内容编辑团队执笔。团队由资深项目经理、前端工程师与设计师组成,所有内容均来自亲手交付的真实项目,先讲清问题、再给出可落地的解法。尧图深耕北京网站建设十年,服务过京华建材集团、智造科技等各行业客户,把一线经验沉淀为可复用的行业观察。

  • 十年建站经验,覆盖建材、制造、服务、文创等
  • 项目经理把关选题与事实准确性
  • 工程师与设计师联合撰写专业细节
  • 统一编辑规范,保证文风与排版一致
  • 每月复盘转化数据,迭代选题方向

延伸阅读

相关资讯与近期热门内容

深度阅读推荐

建站决策前值得细读的三篇

网站改版的5个关键决策
2024-08-12

网站改版的5个关键决策

什么时候该改版、改到什么程度、如何避免流量掉光,京华建材集团改版复盘给出答案。

获取专属建站方案

看完文章,把您的行业与预算告诉我们,免费获取一份量身定制的官网建设方案与报价。

立即免费咨询