第10讲:AI 应用可观测性全景总结与生产实战

发布时间:2026/10/4 17:22:29
第10讲:AI 应用可观测性全景总结与生产实战 一、十讲内容全景回顾┌─────────────────────────────────────────────────────────────────────────┐ │ AI 应用可观测性从埋点到诊断 │ ├─────────────────────────────────────────────────────────────────────────┤ │ │ │ 第1讲 观测维度与指标体系 │ │ ┌─────────────────────────────────────────────────────────────────┐ │ │ │ RED 方法论: Rate / Error / Duration │ │ │ │ USE 方法论: Utilization / Saturation / Errors │ │ │ │ AI 专属指标: Token 消耗 / 模型延迟 / 幻觉率 / 决策置信度 │ │ │ └─────────────────────────────────────────────────────────────────┘ │ │ ↓ │ │ 第2讲 Span 设计与 Context Propagation │ │ ┌─────────────────────────────────────────────────────────────────┐ │ │ │ W3C Trace Context 标准 │ │ │ │ Span 生命周期: Start → Tag → Log → End │ │ │ │ 跨服务传播: HTTP Header / gRPC Metadata / MQ Header │ │ │ └─────────────────────────────────────────────────────────────────┘ │ │ ↓ │ │ 第3讲 LLM 调用监控 │ │ ┌─────────────────────────────────────────────────────────────────┐ │ │ │ Token 计量与计费 │ │ │ │ 模型定价表与成本分摊 │ │ │ │ 告警引擎: 阈值 / 复合 / 静默 │ │ │ └─────────────────────────────────────────────────────────────────┘ │ │ ↓ │ │ 第4讲 决策质量监控与 PSI 漂移检测 │ │ ┌─────────────────────────────────────────────────────────────────┐ │ │ │ 决策质量: 准确率 / 召回率 / 用户满意度 │ │ │ │ PSI 漂移检测: 特征分布变化监控 │ │ │ │ 自动回滚触发: 质量下降 → 回滚到上一版本 │ │ │ └─────────────────────────────────────────────────────────────────┘ │ │ ↓ │ │ 第5讲 Prompt 管理与 AB Test │ │ ┌─────────────────────────────────────────────────────────────────┐ │ │ │ Prompt 版本管理与灰度发布 │ │ │ │ ABTestEngine: 分流 / 指标对比 / 显著性检验 │ │ │ │ InjectionDetector: Prompt 注入攻击检测 │ │ │ └─────────────────────────────────────────────────────────────────┘ │ │ ↓ │ │ 第6讲 日志结构化与存储 │ │ ┌─────────────────────────────────────────────────────────────────┐ │ │ │ Event Schema: 统一日志格式 │ │ │ │ 高性能写入: 无锁 Ring Buffer / 批量压缩 │ │ │ │ 分层存储: ClickHouse(7d) → ES(30d) → S3(90d) │ │ │ └─────────────────────────────────────────────────────────────────┘ │ │ ↓ │ │ 第7讲 实时告警与自动化响应 │ │ ┌─────────────────────────────────────────────────────────────────┐ │ │ │ AlertLevel: P0 ~ P3 分级 │ │ │ │ RuleEngine: 阈值 / 复合 / 静默期 │ │ │ │ AutoResponder: 熔断 / 扩容 / 回滚 / 限流 │ │ │ │ StormSuppressor: 风暴抑制 │ │ │ └─────────────────────────────────────────────────────────────────┘ │ │ ↓ │ │ 第8讲 分布式链路追踪与根因分析 │ │ ┌─────────────────────────────────────────────────────────────────┐ │ │ │ Trace / Span / SpanContext 数据结构 │ │ │ │ 采样策略: HeadBased / TailBased / Priority │ │ │ │ 根因分析: 首个错误 Span / 瓶颈检测 / 异常检测 │ │ │ │ 服务拓扑图: 依赖关系可视化 │ │ │ └─────────────────────────────────────────────────────────────────┘ │ │ ↓ │ │ 第9讲 混沌工程与容灾演练 │ │ ┌─────────────────────────────────────────────────────────────────┐ │ │ │ 故障注入: 网络 / 服务 / 资源 / LLM │ │ │ │ 实验引擎: 稳态检查 → 注入 → 观察 → 回滚 │ │ │ │ 预置场景: LLM 降级 / 服务崩溃 / 资源耗尽 / 依赖故障 │ │ │ └─────────────────────────────────────────────────────────────────┘ │ │ ↓ │ │ 第10讲 全景总结与生产实战本讲 │ │ ┌─────────────────────────────────────────────────────────────────┐ │ │ │ 生产环境完整部署方案 │ │ │ │ Dashboard 设计: 从宏观到微观 │ │ │ │ SLA / SLO / SLI 体系 │ │ │ │ 运维 SOP: On-Call 手册 │ │ │ │ 未来演进方向 │ │ │ └─────────────────────────────────────────────────────────────────┘ │ └─────────────────────────────────────────────────────────────────────────┘二、生产环境完整部署方案2.1 整体架构┌─────────────────────────────────────────────────────────────────────┐ │ 用户请求 │ └──────────────────────────┬──────────────────────────────────────────┘ │ ┌──────────────────────────▼──────────────────────────────────────────┐ │ API Gateway (Envoy / Kong) │ │ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ │ │ │ Rate Limiter │ │ Auth Filter │ │ Trace ID Gen │ │ │ └──────────────┘ └──────────────┘ └──────────────┘ │ └──────────┬───────────────────────────────────────────────────────────┘ │ ┌──────────▼───────────────────────────────────────────────────────────┐ │ Service Mesh (Istio) │ │ ┌──────────────────────────────────────────────────────────────┐ │ │ │ Sidecar Proxy: 自动注入 Trace / 采集 Metrics / 流量管控 │ │ │ └──────────────────────────────────────────────────────────────┘ │ └──────────┬───────────────────────────────────────────────────────────┘ │ ┌──────────▼───────────────────────────────────────────────────────────┐ │ AI 应用服务集群 │ │ │ │ ┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────┐ │ │ │ NLP │ │ LLM │ │ Vector │ │ Reranker │ │ │ │ Service │ │ Proxy │ │ DB │ │ │ │ │ └──────────┘ └──────────┘ └──────────┘ └──────────┘ │ │ │ │ 每个 Pod 内置: │ │ ┌──────────────────────────────────────────────────────────────┐ │ │ │ ① OpenTelemetry SDK (Go / Python) │ │ │ │ ② Prometheus Metrics Exporter (端口 :9090) │ │ │ │ ③ Structured Logger (JSON 格式 → stdout) │ │ │ │ ④ Health Check Handler (/healthz / /readyz) │ │ │ └──────────────────────────────────────────────────────────────┘ │ └─────────────────────────────────────────────────────────────────────┘ │ ┌──────────────────────────▼──────────────────────────────────────────┐ │ 可观测性基础设施 │ │ │ │ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ │ │ │ OpenTelemetry│ │ Kafka │ │ Fluentd │ │ │ │ Collector │ │ (缓冲层) │ │ (日志采集) │ │ │ └──────┬───────┘ └──────┬───────┘ └──────┬───────┘ │ │ │ │ │ │ │ ┌──────▼─────────────────▼─────────────────▼───────┐ │ │ │ 数据处理层 │ │ │ │ ┌──────────┐ ┌──────────┐ ┌──────────────┐ │ │ │ │ │ Metrics │ │ Traces │ │ Logs │ │ │ │ │ │ Processor│ │ Processor│ │ Processor │ │ │ │ │ └────┬─────┘ └────┬─────┘ └──────┬───────┘ │ │ │ └───────┼─────────────┼───────────────┼───────────┘ │ │ │ │ │ │ │ ┌───────▼─────────────▼───────────────▼───────────┐ │ │ │ 存储层 │ │ │ │ ┌──────────┐ ┌──────────┐ ┌──────────────┐ │ │ │ │ │Victoria │ │Jaeger │ │ Loki / ES │ │ │ │ │ │Metrics │ │(Traces) │ │ (Logs) │ │ │ │ │ └──────────┘ └──────────┘ └──────────────┘ │ │ │ └──────────────────────────────────────────────────┘ │ │ │ │ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ │ │ │ Grafana │ │ AlertManager│ │ Incident │ │ │ │ (Dashboards)│ │ (告警管理) │ │ Management │ │ │ └──────────────┘ └──────────────┘ └──────────────┘ │ └─────────────────────────────────────────────────────────────────────┘2.2 Kubernetes 部署清单# observability-stack.yaml --- apiVersion: v1 kind: Namespace metadata: name: observability --- # OpenTelemetry Collector apiVersion: apps/v1 kind: Deployment metadata: name: otel-collector namespace: observability spec: replicas: 3 selector: matchLabels: app: otel-collector template: metadata: labels: app: otel-collector spec: containers: - name: otel-collector image: otel/opentelemetry-collector-contrib:0.120.0 args: - --config/etc/otel/config.yaml ports: - containerPort: 4317 # gRPC - containerPort: 4318 # HTTP volumeMounts: - name: config mountPath: /etc/otel resources: requests: cpu: 500m memory: 512Mi limits: cpu: 2 memory: 2Gi volumes: - name: config configMap: name: otel-config --- apiVersion: v1 kind: ConfigMap metadata: name: otel-config namespace: observability data: config.yaml: | receivers: otlp: protocols: grpc: endpoint: 0.0.0.0:4317 http: endpoint: 0.0.0.0:4318 processors: batch: timeout: 5s send_batch_size: 8192 memory_limiter: check_interval: 1s limit_mib: 1536 spike_limit_mib: 256 attributes: actions: - key: environment value: production action: upsert filter: error_mode: ignore traces: span: - attributes[http.target] /healthz exporters: prometheus: endpoint: 0.0.0.0:8889 namespace: ai_app otlp: endpoint: jaeger:4317 tls: insecure: true loki: endpoint: http://loki:3100/loki/api/v1/push tenant_id: ai-app service: pipelines: traces: receivers: [otlp] processors: [memory_limiter, batch] exporters: [otlp] metrics: receivers: [otlp] processors: [memory_limiter, filter, batch] exporters: [prometheus] logs: receivers: [otlp] processors: [memory_limiter, batch] exporters: [loki] --- # VictoriaMetrics apiVersion: apps/v1 kind: StatefulSet metadata: name: victoria-metrics namespace: observability spec: replicas: 2 selector: matchLabels: app: victoria-metrics serviceName: victoria-metrics template: metadata: labels: app: victoria-metrics spec: containers: - name: victoria-metrics image: victoriametrics/victoria-metrics:v1.108.0 args: - -storageDataPath/data - -retentionPeriod30d - -search.maxUniqueTimeseries1000000 ports: - containerPort: 8428 volumeMounts: - name: data mountPath: /data resources: requests: cpu: 1 memory: 2Gi limits: cpu: 4 memory: 8Gi volumeClaimTemplates: - metadata: name: data spec: accessModes: [ReadWriteOnce] storageClassName: ssd resources: requests: storage: 500Gi --- # Jaeger apiVersion: apps/v1 kind: Deployment metadata: name: jaeger namespace: observability spec: replicas: 2 selector: matchLabels: app: jaeger template: metadata: labels: app: jaeger spec: containers: - name: jaeger image: jaegertracing/all-in-one:1.63.0 env: - name: COLLECTOR_OTLP_ENABLED value: true - name: SPAN_STORAGE_TYPE value: elasticsearch - name: ES_SERVER_URLS value: http://elasticsearch:9200 ports: - containerPort: 16686 # UI - containerPort: 4317 # OTLP gRPC resources: requests: cpu: 500m memory: 1Gi limits: cpu: 2 memory: 4Gi --- # Grafana apiVersion: apps/v1 kind: Deployment metadata: name: grafana namespace: observability spec: replicas: 2 selector: matchLabels: app: grafana template: metadata: labels: app: grafana spec: containers: - name: grafana image: grafana/grafana:11.4.0 env: - name: GF_SECURITY_ADMIN_PASSWORD valueFrom: secretKeyRef: name: grafana-secret key: admin-password - name: GF_INSTALL_PLUGINS value: grafana-piechart-panel,grafana-worldmap-panel ports: - containerPort: 3000 volumeMounts: - name: dashboards mountPath: /etc/grafana/provisioning/dashboards - name: datasources mountPath: /etc/grafana/provisioning/datasources resources: requests: cpu: 500m memory: 512Mi limits: cpu: 1 memory: 1Gi volumes: - name: dashboards configMap: name: grafana-dashboards - name: datasources configMap: name: grafana-datasources --- apiVersion: v1 kind: ConfigMap metadata: name: grafana-datasources namespace: observability data: datasources.yaml: | apiVersion: 1 datasources: - name: VictoriaMetrics type: prometheus url: http://victoria-metrics:8428 access: proxy isDefault: true - name: Jaeger type: jaeger url: http://jaeger:16686 access: proxy - name: Loki type: loki url: http://loki:3100 access: proxy --- # AI 应用服务示例 apiVersion: apps/v1 kind: Deployment metadata: name: llm-proxy namespace: ai-app spec: replicas: 5 selector: matchLabels: app: llm-proxy template: metadata: labels: app: llm-proxy annotations: sidecar.istio.io/inject: true prometheus.io/scrape: true prometheus.io/port: 9090 spec: containers: - name: llm-proxy image: registry.ai-app/llm-proxy:v2.3.1 ports: - containerPort: 8080 name: http - containerPort: 9090 name: metrics env: - name: OTEL_EXPORTER_OTLP_ENDPOINT value: http://otel-collector.observability:4317 - name: OTEL_SERVICE_NAME value: llm-proxy - name: OTEL_TRACES_SAMPLER value: parentbased_traceidratio - name: OTEL_TRACES_SAMPLER_ARG value: 0.1 livenessProbe: httpGet: path: /healthz port: 8080 initialDelaySeconds: 10 periodSeconds: 15 readinessProbe: httpGet: path: /readyz port: 8080 initialDelaySeconds: 5 periodSeconds: 10 resources: requests: cpu: 500m memory: 512Mi limits: cpu: 2 memory: 2Gi volumeMounts: - name: config mountPath: /etc/app volumes: - name: config configMap: name: llm-proxy-config --- # HPA 自动扩缩容 apiVersion: autoscaling/v2 kind: HorizontalPodAutoscaler metadata: name: llm-proxy-hpa namespace: ai-app spec: scaleTargetRef: apiVersion: apps/v1 kind: Deployment name: llm-proxy minReplicas: 3 maxReplicas: 20 metrics: - type: Resource resource: name: cpu target: type: Utilization averageUtilization: 70 - type: Pods pods: metric: name: llm_proxy_queue_depth target: type: AverageValue averageValue: 100 behavior: scaleUp: stabilizationWindowSeconds: 60 policies: - type: Percent value: 100 periodSeconds: 15 scaleDown: stabilizationWindowSeconds: 300 policies: - type: Percent value: 10 periodSeconds: 60三、Dashboard 设计3.1 宏观大盘CEO 视角┌─────────────────────────────────────────────────────────────────────────┐ │ AI 应用可观测性 · 宏观大盘 [Last 24h] [Auto-refresh] │ ├─────────────────────────────────────────────────────────────────────────┤ │ │ │ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐│ │ │ 总请求量 │ │ 成功率 │ │ P99 延迟 │ │ 日活用户 ││ │ │ 1,234,567 │ │ 99.87% │ │ 892ms │ │ 89,123 ││ │ │ ▲ 12% vs 昨日 │ │ ▼ 0.05% │ │ ▲ 45ms │ │ ▲ 5% ││ │ └──────────────┘ └──────────────┘ └──────────────┘ └──────────────┘│ │ │ │ ┌─────────────────────────────────────────────────────────────────┐ │ │ │ 服务健康状态 │ │ │ │ ┌─────┐ ┌─────┐ ┌─────┐ ┌─────┐ ┌─────┐ ┌─────┐ │ │ │ │ │API │ │NLP │ │LLM │ │VecDB│ │Rer │ │Cache│ │ │ │ │ │GW │ │Svc │ │Proxy│ │ │ │anker│ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │200ms│ │50ms │ │1.2s │ │35ms │ │15ms │ │2ms │ │ │ │ │ └─────┘ └─────┘ └─────┘ └─────┘ └─────┘ └─────┘ │ │ │ └─────────────────────────────────────────────────────────────────┘ │ │ │ │ ┌─────────────────────────────────────────────────────────────────┐ │ │ │ 过去 24 小时 P99 延迟趋势 │ │ │ │ ▁▂▃▄▅▆▇█▇▆▅▄▃▂▁▂▃▄▅▆▇█▇▆▅▄▃▂▁▂▃▄▅▆▇█▇▆▅▄▃▂▁▂▃▄▅▆▇█▇▆▅▄▃▂▁ │ │ │ │ 00:00 04:00 08:00 12:00 16:00 20:00 现在 │ │ │ └─────────────────────────────────────────────────────────────────┘ │ │ │ │ ┌─────────────────────────────┐ ┌─────────────────────────────────┐ │ │ │ 错误分布 Top 5 │ │ Token 消耗趋势 │ │ │ │ ┌─────────────────────┐ │ │ ▁▂▃▄▅▆▇█▇▆▅▄▃▂▁ │ │ │ │ │ LLM Timeout 45% │ │ │ 今日: 12.3M tokens │ │ │ │ │ Rate Limit 22% │ │ │ 费用: $246.78 │ │ │ │ │ Vector DB 15% │ │ └─────────────────────────────────┘ │ │ │ │ Auth Fail 10% │ │ │ │ │ │ Other 8% │ │ │ │ │ └─────────────────────┘ │ │ │ └─────────────────────────────┘ └─────────────────────────────────┘ │ └─────────────────────────────────────────────────────────────────────────┘3.2 服务详细面板SRE 视角{ title: LLM Proxy 详细面板, panels: [ { title: QPS 延迟, type: timeseries, queries: [ rate(llm_proxy_requests_total[1m]), histogram_quantile(0.99, rate(llm_proxy_request_duration_seconds_bucket[5m])), histogram_quantile(0.95, rate(llm_proxy_request_duration_seconds_bucket[5m])), histogram_quantile(0.50, rate(llm_proxy_request_duration_seconds_bucket[5m])) ] }, { title: Token 消耗明细, type: stat, queries: [ sum(rate(llm_proxy_prompt_tokens_total[5m])), sum(rate(llm_proxy_completion_tokens_total[5m])), sum(rate(llm_proxy_cost_usd_total[5m])) ] }, { title: 模型调用分布, type: piechart, queries: [ count by(model) (llm_proxy_requests_total) ] }, { title: 错误原因分布, type: barchart, queries: [ count by(error_type) (llm_proxy_errors_total) ] }, { title: 熔断器状态, type: stat, queries: [ llm_proxy_circuit_breaker_state{state\open\}, llm_proxy_circuit_breaker_state{state\half-open\}, llm_proxy_circuit_breaker_state{state\closed\} ] }, { title: 最近 Trace 列表, type: table, datasource: Jaeger, queries: [ servicellm-proxy limit20 lookback1h ], columns: [TraceID, Duration, Spans, Services, Errors] } ] }3.3 业务指标面板产品经理视角{ title: AI 应用业务指标, panels: [ { title: 日活跃用户 (DAU), type: timeseries, query: sum(increase(ai_app_active_users_total[24h])) }, { title: 用户满意度评分, type: gauge, query: avg(ai_app_user_satisfaction_score), thresholds: { green: 4.0, yellow: 3.0, red: 2.0 } }, { title: 问答采纳率, type: timeseries, query: sum(rate(ai_app_answer_accepted_total[1h])) / sum(rate(ai_app_answers_total[1h])) }, { title: 模型版本分布, type: piechart, query: count by(model_version) (ai_app_model_deployments) }, { title: AB 实验效果对比, type: stat, queries: [ avg(ai_app_conversion_rate{experiment_groupcontrol}), avg(ai_app_conversion_rate{experiment_grouptreatment}) ] }, { title: Token 成本按部门, type: barchart, query: sum by(department) (ai_app_token_cost_total) } ] }四、SLA / SLO / SLI 体系4.1 定义SLA (Service Level Agreement) —— 对外承诺 └── SLO (Service Level Objective) —— 内部目标 └── SLI (Service Level Indicator) —— 可测量指标4.2 AI 应用 SLI 指标slis: # 可用性 availability: definition: 成功响应的请求占比 measurement: successful_requests / total_requests * 100 exclusion: 排除计划内维护窗口 # 延迟 latency_p99: definition: 最慢 1% 请求的响应时间 measurement: histogram_quantile(0.99, request_duration_seconds) exclusion: 排除超时已熔断的请求 # 吞吐量 throughput: definition: 每秒处理的请求数 measurement: rate(requests_total[1m]) # 新鲜度 freshness: definition: 数据从产生到可查询的时间 measurement: max(event_timestamp - ingestion_timestamp) # 正确性 correctness: definition: AI 回答被用户采纳的比例 measurement: accepted_answers / total_answers * 100 # 成本效率 cost_efficiency: definition: 每美元 token 产出的有效回答数 measurement: accepted_answers / total_cost_usd4.3 SLO 目标设定slo_targets: # Tier 1: 核心对话服务 tier_1: description: 用户直接感知的核心链路 targets: availability: 99.99% # 全年宕机 52min latency_p99: 2s # 99% 请求在 2s 内 correctness: 85% # 回答采纳率 85% burn_rate: warning: 10% / 7d # 7 天内消耗 10% 预算 critical: 20% / 1d # 1 天内消耗 20% 预算 consequences: - 触发 P0 告警 - 立即回滚最近变更 - 全员 On-Call # Tier 2: 辅助功能 tier_2: description: 增强体验但不阻塞核心功能 targets: availability: 99.9% latency_p99: 5s freshness: 30s burn_rate: warning: 20% / 7d critical: 50% / 1d consequences: - 触发 P1 告警 - 下个工作日修复 # Tier 3: 后台批处理 tier_3: description: 离线分析和模型训练 targets: availability: 99.0% throughput: 每日处理 1M 条 freshness: 1h consequences: - 触发 P2 告警 - 本周内修复4.4 错误预算管理error_budget: # 以 Tier 1 为例99.99% 可用性 每年 52min 错误预算 calculation: annual_uptime: 99.99% annual_downtime_budget: 52 minutes monthly_budget: 4.3 minutes weekly_budget: 1 minute consumption_tracking: - date: 2026-09-20 consumed: 12s remaining_weekly: 48s status: 正常 - date: 2026-09-21 consumed: 45s remaining_weekly: 15s status: 警告 - date: 2026-09-22 consumed: 72s remaining_weekly: -12s status: 超额 actions_on_exhaustion: yellow: - 暂停所有非紧急变更 - 启动根因分析 red: - 冻结所有变更 - 全员介入修复 - 通知 VP 级别五、运维 SOPOn-Call 手册5.1 On-Call 流程┌─────────────────────────────────────────────────────────────────┐ │ On-Call 响应流程 │ ├─────────────────────────────────────────────────────────────────┤ │ │ │ 收到告警 │ │ │ │ │ ▼ │ │ ┌──────────────────────┐ │ │ │ 1. ACKNOWLEDGE │ ← 5分钟内必须确认 │ │ │ 确认告警 │ │ │ └──────────┬───────────┘ │ │ │ │ │ ▼ │ │ ┌──────────────────────┐ │ │ │ 2. TRIAGE │ ← 判断级别 │ │ │ 分类定级 │ │ │ └──────────┬───────────┘ │ │ │ │ │ ┌──────┴──────┐ │ │ ▼ ▼ │ │ ┌────────┐ ┌────────────┐ │ │ │ P0/P1 │ │ P2/P3 │ │ │ │ 立即响应│ │ 工作时间处理│ │ │ └───┬────┘ └────────────┘ │ │ │ │ │ ▼ │ │ ┌──────────────────────┐ │ │ │ 3. DIAGNOSE │ ← 查看 Dashboard / Trace / Log │ │ │ 诊断根因 │ │ │ └──────────┬───────────┘ │ │ │ │ │ ▼ │ │ ┌──────────────────────┐ │ │ │ 4. MITIGATE │ ← 回滚 / 重启 / 扩容 / 降级 │ │ │ 止血恢复 │ │ │ └──────────┬───────────┘ │ │ │ │ │ ▼ │ │ ┌──────────────────────┐ │ │ │ 5. RESOLVE │ ← 确认指标恢复正常 │ │ │ 确认解决 │ │ │ └──────────┬───────────┘ │ │ │ │ │ ▼ │ │ ┌──────────────────────┐ │ │ │ 6. POSTMORTEM │ ← 24h 内提交事故事后分析 │ │ │ 事后复盘 │ │ │ └──────────────────────┘ │ └─────────────────────────────────────────────────────────────────┘5.2 常见故障处理手册incident_handbook: # 场景 1: LLM 响应超时飙升 scenario_llm_timeout: symptoms: - P99 延迟 5s - llm_proxy_timeout_errors 飙升 - 用户反馈回复很慢 diagnosis: step_1: 检查 LLM Provider 状态页 step_2: 查看 Jaeger 中 llm-proxy 的 Trace step_3: 检查 Token 消耗是否有异常突增 step_4: 检查模型是否被限流 mitigation: option_a: 切换到备用模型 (gpt-3.5-turbo → claude-haiku) option_b: 降低 max_tokens 限制 option_c: 启用本地缓存减少重复调用 option_d: 扩容 llm-proxy 实例 escalation: if_not_resolved_in: 15分钟 escalate_to: AI Platform Team Lead # 场景 2: Vector DB 不可用 scenario_vector_db_down: symptoms: - vector_db_errors 100% - RAG 功能完全不可用 - 告警: VectorDBConnectionFailed diagnosis: step_1: 检查 Vector DB Pod 状态 step_2: 查看 PVC 磁盘使用率 step_3: 检查网络策略是否误拦截 step_4: 查看 Vector DB 日志 mitigation: option_a: 重启 Vector DB Pod option_b: 扩容副本数 option_c: 降级为关键词搜索 (BM25) option_d: 切换只读副本 escalation: if_not_resolved_in: 10分钟 escalate_to: DBA Team # 场景 3: 模型漂移导致回答质量下降 scenario_model_drift: symptoms: - 用户满意度评分下降 10% - PSI 指标超出阈值 - 负面反馈增多 diagnosis: step_1: 查看 PSI Dashboard 确认漂移维度 step_2: 对比新旧模型版本的回答样本 step_3: 检查 Prompt 是否被意外修改 step_4: 检查上游数据分布是否变化 mitigation: option_a: 回滚到上一个稳定模型版本 option_b: 切换 AB 实验组到对照组 option_c: 临时增加后处理校验规则 escalation: if_not_resolved_in: 30分钟 escalate_to: ML Team Lead5.3 事后复盘模板# 事故事后复盘报告 ## 基本信息 - **事故编号**: INC-2026-0928-001 - **标题**: LLM 代理服务 P99 延迟飙升至 8s - **严重级别**: P0 - **日期**: 2026-09-28 - **持续时间**: 23 分钟 - **影响范围**: 所有依赖 LLM 的用户请求 ## 时间线 | 时间 | 事件 | |------|------| | 14:23 | 告警触发: P99 延迟 5s | | 14:24 | On-Call 工程师确认告警 | | 14:26 | 诊断: OpenAI API 响应变慢 | | 14:28 | 执行降级: 切换到备用模型 | | 14:35 | 指标恢复正常 | | 14:46 | OpenAI 发布状态更新确认故障 | | 15:00 | 切换回主模型 | ## 根因分析 - **直接原因**: OpenAI API 出现区域性延迟 - **根本原因**: 未配置多区域故障转移 - **促成因素**: 备用模型预热不足首次切换延迟较高 ## 改进措施 | 项目 | 负责人 | 截止日期 | |------|--------|----------| | 配置多区域 LLM Provider 故障转移 | infra-team | 2026-10-05 | | 备用模型保持 Warm Pool | ml-team | 2026-10-03 | | 添加 Provider 健康探测 | sre-team | 2026-10-01 | | 更新 On-Call 手册 LLM 降级章节 | docs-team | 2026-09-30 | ## 附件 - [Grafana Dashboard 截图] - [Jaeger Trace 链接] - [告警记录]六、未来演进方向6.1 AI for Observability (AIOps)aiops_roadmap: phase_1: 异常检测自动化 capabilities: - 基于历史数据的动态阈值 - 多维度的异常关联分析 - 季节性模式自动识别 example: 自动区分工作日/周末流量差异避免误告警 phase_2: 根因分析智能化 capabilities: - 因果推断: 从相关性到因果性 - 知识图谱: 服务依赖关系推理 - 自然语言根因描述 example: LLM 分析 Trace 后直接输出: llm-proxy 的 call_llm span 耗时 8.2s占整条 Trace 的 98%根因为 OpenAI API 区域性延迟 phase_3: 自愈系统 capabilities: - 故障自动诊断 自动修复 - 渐进式回滚: 1% → 5% → 20% → 100% - 混沌工程与自愈形成闭环 example: 检测到模型漂移 → 自动触发 AB 实验回滚 → 验证指标恢复 → 通知团队 phase_4: 预测性可观测性 capabilities: - 容量预测: 提前 24h 预测资源需求 - 故障预测: 基于模式识别的故障预警 - 成本预测: Token 消耗和费用预估 example: 预测今晚 20:00 会有流量高峰提前扩容 3 个副本6.2 技术演进趋势technology_trends: # eBPF 深度观测 ebpf: description: 无需修改代码即可观测内核态行为 use_cases: - 网络延迟的精确测量 - 文件 I/O 性能分析 - 系统调用追踪 tools: - Pixie - Cilium Tetragon - Parca (连续分析) # OpenTelemetry 标准化 opentelemetry: description: 可观测性的行业标准统一 Metrics/Traces/Logs trends: - Profiling Signal 加入 OTel - eBPF 与 OTel 融合 - OTel Operator 简化部署 # 低成本存储 cost_effective_storage: description: 海量观测数据的存储成本优化 strategies: - 列式存储 高压缩比 - 对象存储作为冷备 - 采样 聚合降低数据量 tools: - Grafana Mimir - Thanos - ClickHouse # 持续验证 continuous_validation: description: 可观测性本身也需要被观测 practices: - SLI 覆盖率检查 - 告警质量评分 - Dashboard 使用率统计七、总结从埋点到诊断的完整链路┌─────────────────────────────────────────────────────────────────────────┐ │ AI 应用可观测性从埋点到诊断 │ │ │ │ 埋点层 (Instrumentation) │ │ ├── 代码埋点: OpenTelemetry SDK │ │ ├── 自动埋点: Istio Sidecar / eBPF │ │ └── 业务埋点: 自定义 Metrics / Events │ │ │ │ │ ▼ │ │ 采集层 (Collection) │ │ ├── Metrics: Prometheus Exporter (端口 :9090) │ │ ├── Traces: OTLP Exporter → OpenTelemetry Collector │ │ └── Logs│ └── Logs: Structured JSON → stdout → Fluentd │ │ │ ▼ │ 传输层 (Transport) │ ├── 缓冲: Kafka (削峰填谷) │ ├── 采样: Tail-Based Sampling (错误全采 / 慢请求半采 / 正常 1%) │ └── 压缩: gzip / snappy (减少带宽) │ │ │ ▼ │ 存储层 (Storage) │ ├── 热存储 (7天): ClickHouse / VictoriaMetrics │ ├── 温存储 (30天): Elasticsearch / Jaeger │ └── 冷存储 (90天): S3 / GCS (Parquet 格式) │ │ │ ▼ │ 分析层 (Analysis) │ ├── 指标分析: PromQL 查询 / 聚合 / 预测 │ ├── 链路分析: Trace 查询 / 服务拓扑 / 根因定位 │ └── 日志分析: 全文检索 / 模式识别 / 聚类 │ │ │ ▼ │ 展示层 (Visualization) │ ├── Grafana Dashboards: 宏观大盘 / 服务面板 / 业务面板 │ ├── Jaeger UI: Trace 详情 / 火焰图 / 比较视图 │ └── 自定义 Portal: 业务指标 / 成本报表 / SLA 看板 │ │ │ ▼ │ 行动层 (Action) │ ├── 告警: AlertManager → 电话 / Slack / 邮件 │ ├── 自动化: AutoResponder → 熔断 / 扩容 / 回滚 / 限流 │ └── 混沌: Chaos Engine → 故障注入 / 韧性验证 / 改进闭环 │ └─────────────────────────────────────────────────────────────────────────┘八、附录常用命令速查8.1 排查问题三板斧# 1. 看指标 —— 哪里慢了 # P99 延迟最高的服务 Top 5 topk(5, histogram_quantile(0.99, sum(rate(http_request_duration_seconds_bucket[5m])) by (service, le))) # 2. 看链路 —— 为什么慢 # 查询最近 1 小时最慢的 10 条 Trace # 在 Jaeger UI 中执行: # Service: llm-proxy # Operation: call_llm # Min Duration: 1s # Lookback: 1h # Limit: 10 # 3. 看日志 —— 报了什么错 # 查询特定 TraceID 的所有日志 {appllm-proxy} | trace_ida1b2c3d4e5f68.2 常用 PromQL 查询# 服务可用性 (排除 503 熔断) sum(rate(http_requests_total{status!~5..}[5m])) / sum(rate(http_requests_total[5m])) * 100 # P99 延迟趋势 histogram_quantile(0.99, sum(rate(http_request_duration_seconds_bucket[5m])) by (le)) # 错误率 (5xx / total) sum(rate(http_requests_total{status~5..}[5m])) / sum(rate(http_requests_total[5m])) * 100 # Token 消耗速率 sum(rate(llm_token_total[5m])) by (model) # 每秒成本 sum(rate(llm_cost_usd_total[5m])) # 熔断器打开次数 increase(circuit_breaker_open_total[1h]) # 最慢的 5 个端点 topk(5, avg by(endpoint) (http_request_duration_seconds_sum / http_request_duration_seconds_count)) # 各服务 Span 数量 count by(service.name) (spans{})8.3 常用 kubectl 命令# 查看所有 Pod 状态 kubectl get pods -n ai-app -o wide # 查看 Pod 日志 kubectl logs -n ai-app deployment/llm-proxy --tail100 -f # 查看 Pod 资源使用 kubectl top pod -n ai-app # 查看 HPA 状态 kubectl get hpa -n ai-app # 查看事件 kubectl get events -n ai-app --sort-by.lastTimestamp # 端口转发 (本地访问 Jaeger) kubectl port-forward -n observability svc/jaeger 16686:16686 # 进入 Pod 调试 kubectl exec -it -n ai-app pod/llm-proxy-xxx -- sh # 查看 ConfigMap kubectl get configmap -n ai-app llm-proxy-config -o yaml # 滚动重启 kubectl rollout restart -n ai-app deployment/llm-proxy # 查看滚动状态 kubectl rollout status -n ai-app deployment/llm-proxy九、推荐阅读与工具9.1 必读书籍书名作者推荐理由《Site Reliability Engineering》Google SRE TeamSRE 圣经定义了现代可观测性理念《The Art of Monitoring》James Turnbull从零搭建监控体系的实操指南《Distributed Tracing in Practice》Austin Parker 等链路追踪的权威实践手册《Chaos Engineering》Casey Rosenthal 等混沌工程的奠基之作《Observability Engineering》Charity Majors 等可观测性三大支柱的系统讲解9.2 开源工具推荐类别工具说明Metrics​VictoriaMetrics高性能时序数据库兼容 PromQLTraces​Jaeger / Tempo分布式链路追踪Logs​Loki / Quickwit低成本日志存储Profiling​Parca / Pyroscope持续性能分析eBPF​Pixie / Cilium零侵入内核观测混沌​Chaos Mesh / LitmusK8s 原生混沌工程告警​AlertManager / Keep告警管理和抑制Dashboard​Grafana / Perses可视化面板9.3 在线工具工具用途地址数字转大写金额转换zz365.top/daxieJSON 格式化Trace 数据美化zz365.top/json时间戳转换Unix 时间 ↔ 可读时间zz365.top/timestampCron 表达式告警周期配置zz365.top/cronBase64 编解码Token 解码调试zz365.top/base64十、结语至此《AI 应用可观测性从埋点到诊断》十讲全部结束。我们从最基础的观测维度出发一步步构建了完整的可观测性体系第1讲 知道要看什么指标体系 第2讲 知道怎么串起来Trace 第3-5讲 知道 AI 特有的坑LLM / 漂移 / Prompt 第6讲 知道怎么存日志结构 第7讲 知道怎么反应告警响应 第8讲 知道怎么查根因链路分析 第9讲 知道怎么验证韧性混沌工程 第10讲 知道怎么落地生产实战记住三条铁律没有度量就没有改进​ —— 先埋点再优化没有 Trace 的日志就是孤岛​ —— 永远带着 TraceID没有自动化的告警就是噪音​ —— 告警必须能触发行动祝你的 AI 应用永远稳定、高效、可观测 开发之余的小工具推荐设计混沌实验场景时经常需要计算各种时间窗口和延迟参数。zz365.top 的在线计算器可以快速进行毫秒/秒/分钟的单位换算和百分比计算帮助你在配置故障参数时更精确。所有计算纯前端完成不需要联网。

关于本文作者

来自尧图内容编辑团队

尧图内容编辑团队 内容团队

尧图内容编辑团队

本文由尧图网络内容编辑团队执笔。团队由资深项目经理、前端工程师与设计师组成,所有内容均来自亲手交付的真实项目,先讲清问题、再给出可落地的解法。尧图深耕北京网站建设十年,服务过京华建材集团、智造科技等各行业客户,把一线经验沉淀为可复用的行业观察。

  • 十年建站经验,覆盖建材、制造、服务、文创等
  • 项目经理把关选题与事实准确性
  • 工程师与设计师联合撰写专业细节
  • 统一编辑规范,保证文风与排版一致
  • 每月复盘转化数据,迭代选题方向

延伸阅读

相关资讯与近期热门内容

深度阅读推荐

建站决策前值得细读的三篇

网站改版的5个关键决策
2024-08-12

网站改版的5个关键决策

什么时候该改版、改到什么程度、如何避免流量掉光,京华建材集团改版复盘给出答案。

获取专属建站方案

看完文章,把您的行业与预算告诉我们,免费获取一份量身定制的官网建设方案与报价。

立即免费咨询