Files
xiaoxia-saas/docs/HEALTH-CHECKS.md
T
Xiaoxia AI 082ad20913 feat(health): add comprehensive health check endpoints
- Add /health endpoint for liveness probe (fast, no dependencies)
- Add /ready endpoint for readiness probe (checks database + redis)
- Add /startup endpoint for startup probe (checks migrations)
- Return 503 when not ready/started
- Detailed check results in response
- Include Kubernetes/Docker/Nginx configuration examples
- Add comprehensive health check documentation
- Include monitoring and alerting setup

Phase 4 Task 45/68 completed
2026-06-17 08:42:56 +08:00

6.7 KiB
Raw Blame History

健康检查和探针指南

📋 概述

小虾 SaaS 提供三种健康检查端点,用于不同的监控场景。


🔍 健康检查端点

1. /health - 存活检查(Liveness

用途: 检查应用是否存活

特点:

  • 快速响应(<10ms
  • 不检查依赖服务
  • 始终返回 200(除非崩溃)

请求:

curl http://localhost:8000/health

响应:

{
  "status": "healthy",
  "timestamp": "2026-06-17T08:41:00Z",
  "version": "1.0.0"
}

使用场景:

  • Kubernetes liveness probe
  • Docker health check
  • 负载均衡器健康检查

2. /ready - 就绪检查(Readiness

用途: 检查应用是否就绪接收流量

特点:

  • 检查所有依赖服务
  • 数据库连接
  • Redis 连接
  • 失败返回 503

请求:

curl http://localhost:8000/ready

成功响应(200):

{
  "status": "ready",
  "timestamp": "2026-06-17T08:41:00Z",
  "checks": {
    "database": {
      "status": "healthy",
      "type": "postgresql",
      "message": "Database connection successful"
    },
    "redis": {
      "status": "healthy",
      "type": "redis",
      "message": "Redis connection successful"
    }
  }
}

失败响应(503):

{
  "status": "not_ready",
  "timestamp": "2026-06-17T08:41:00Z",
  "checks": {
    "database": {
      "status": "unhealthy",
      "type": "postgresql",
      "message": "Database connection failed: timeout"
    },
    "redis": {
      "status": "healthy",
      "type": "redis",
      "message": "Redis connection successful"
    }
  }
}

使用场景:

  • Kubernetes readiness probe
  • 负载均衡器流量切换
  • 灰度发布健康检查

3. /startup - 启动检查(Startup

用途: 检查应用是否完成启动

特点:

  • 检查数据库迁移
  • 检查初始化状态
  • 失败返回 503

请求:

curl http://localhost:8000/startup

响应:

{
  "status": "started",
  "timestamp": "2026-06-17T08:41:00Z",
  "checks": {
    "database": {
      "status": "healthy",
      "type": "postgresql",
      "message": "Database connection successful"
    },
    "migrations": {
      "status": "healthy",
      "message": "Database migrations applied"
    }
  }
}

使用场景:

  • Kubernetes startup probe
  • 应用启动验证
  • 部署后检查

🐳 Docker 配置

Dockerfile

# 健康检查
HEALTHCHECK --interval=30s --timeout=3s --start-period=40s --retries=3 \
  CMD python -c "import urllib.request; urllib.request.urlopen('http://localhost:8000/health')" || exit 1

docker-compose.yml

services:
  api:
    image: xiaoxia-saas:latest
    healthcheck:
      test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
      interval: 30s
      timeout: 3s
      retries: 3
      start_period: 40s

☸️ Kubernetes 配置

Deployment

apiVersion: apps/v1
kind: Deployment
metadata:
  name: xiaoxia-api
spec:
  replicas: 3
  template:
    spec:
      containers:
      - name: api
        image: xiaoxia-saas:latest
        ports:
        - containerPort: 8000
        
        # 存活探针(应用崩溃时重启)
        livenessProbe:
          httpGet:
            path: /health
            port: 8000
          initialDelaySeconds: 15
          periodSeconds: 10
          timeoutSeconds: 3
          failureThreshold: 3
        
        # 就绪探针(依赖服务未就绪时不接收流量)
        readinessProbe:
          httpGet:
            path: /ready
            port: 8000
          initialDelaySeconds: 5
          periodSeconds: 5
          timeoutSeconds: 3
          failureThreshold: 2
        
        # 启动探针(启动慢的应用)
        startupProbe:
          httpGet:
            path: /startup
            port: 8000
          initialDelaySeconds: 0
          periodSeconds: 5
          timeoutSeconds: 3
          failureThreshold: 30  # 30 * 5s = 150s 最大启动时间

🔧 负载均衡器配置

Nginx

upstream xiaoxia_api {
    server 10.0.0.1:8000 max_fails=3 fail_timeout=30s;
    server 10.0.0.2:8000 max_fails=3 fail_timeout=30s;
}

server {
    location / {
        proxy_pass http://xiaoxia_api;
        
        # 健康检查
        health_check uri=/health interval=10s fails=3 passes=2;
    }
}

AWS ALB

TargetGroup:
  HealthCheckEnabled: true
  HealthCheckPath: /health
  HealthCheckIntervalSeconds: 30
  HealthCheckTimeoutSeconds: 5
  HealthyThresholdCount: 2
  UnhealthyThresholdCount: 3
  Matcher:
    HttpCode: '200'

📊 监控告警

Prometheus 配置

scrape_configs:
  - job_name: 'xiaoxia-health'
    metrics_path: '/health'
    scrape_interval: 30s
    static_configs:
      - targets: ['api-1:8000', 'api-2:8000']

告警规则

groups:
- name: xiaoxia_health
  rules:
  - alert: APINotHealthy
    expr: up{job="xiaoxia-health"} == 0
    for: 1m
    annotations:
      summary: "小虾 API 不健康"
      description: "{{ $labels.instance }} 连续 1 分钟不健康"
  
  - alert: APINotReady
    expr: probe_success{job="xiaoxia-ready"} == 0
    for: 2m
    annotations:
      summary: "小虾 API 未就绪"
      description: "{{ $labels.instance }} 连续 2 分钟未就绪"

🎯 最佳实践

1. 探针超时配置

liveness: 较长超时(10s),避免误杀
readiness: 较短超时(5s),快速摘除异常实例
startup: 最长超时(150s),给足启动时间

2. 失败阈值

liveness: 较高阈值(3次),避免频繁重启
readiness: 较低阈值(2次),快速摘除
startup: 很高阈值(30次),容忍启动慢

3. 检查间隔

liveness: 10-30s
readiness: 5-10s
startup: 5s

4. 依赖检查

  • liveness: 不检查依赖(避免雪崩)
  • readiness: 检查所有依赖
  • startup: 只检查必需依赖

🐛 故障排查

健康检查失败

# 1. 手动测试端点
curl -v http://localhost:8000/health
curl -v http://localhost:8000/ready
curl -v http://localhost:8000/startup

# 2. 查看日志
docker logs xiaoxia-api

# 3. 检查数据库连接
docker exec xiaoxia-api psql $DATABASE_URL -c "SELECT 1"

# 4. 检查 Redis 连接
docker exec xiaoxia-api redis-cli -u $REDIS_URL ping

Pod 频繁重启

# 查看 Pod 事件
kubectl describe pod xiaoxia-api-xxx

# 查看探针配置
kubectl get pod xiaoxia-api-xxx -o yaml | grep -A 10 livenessProbe

# 调整探针参数
kubectl patch deployment xiaoxia-api -p '
spec:
  template:
    spec:
      containers:
      - name: api
        livenessProbe:
          failureThreshold: 5
          timeoutSeconds: 5
'

最后更新: 2026-06-17