관측 · 알림 파이프라인 (공통)

관측 · 알림 파이프라인 (공통) A workflow diagram generated by Archify. 01 / 앱 (ns=wf) 02 / VictoriaMetrics 수집 03 / 저장 · 평가 04 / 라우팅 05 / 발송 스크랩 적재 평가 · 알림 메트릭 노출 · /actuator/prometheus · 앱 (ns=wf) › 스크랩 메트릭 노출 /actuator/prometheus 스크랩 대상 · ServiceMonitor 30s · VictoriaMetrics 수집 › 스크랩 · VMServiceScrape 스크랩 대상 ServiceMonitor 30s VMServiceScrape VMAgent · remoteWrite · VictoriaMetrics 수집 › 적재 VMAgent remoteWrite VMSingle · 보존 1개월+ · 저장 · 평가 › 평가 · 알림 VMSingle 보존 1개월+ VMAlert · PrometheusRule 평가 · 저장 · 평가 › 평가 · 알림 · for 복원 VMAlert PrometheusRule 평가 for 복원 Alertmanager · repeat 4h → 앱별 24h · 라우팅 › 평가 · 알림 Alertmanager repeat 4h → 앱별 24h Telegram · send_resolved · 발송 › 평가 · 알림 Telegram send_resolved firing Legend Agent logic Policy Tool action Context / trace Cloud service External system

단일 소스

  • • kube-prometheus-stack의 Prometheus는 비활성 — VM이 단일 소스(2026-07-20)
  • • 기존 ServiceMonitor·PrometheusRule은 VMServiceScrape·VMRule로 변환돼 그대로 쓰인다
  • • 보존 1개월 이상

노이즈 억제

  • • Watchdog 알림은 라우팅에서 버린다
  • • 앱별 sub-route로 repeat 24h, 주말 mute
  • • alertname·symbol로 그룹핑해 같은 종목 반복 발송을 줄인다