레이블이 pv인 게시물을 표시합니다. 모든 게시물 표시
레이블이 pv인 게시물을 표시합니다. 모든 게시물 표시

prometheus stall

특정 클러스터1 prometheus 조회가 간헐적으로 stall(멈춤,지연)되는 현상이 발생했다.
버스트(동시에 여러 접속 시도)시에 주로 발생한다.
같은 구성의 k8s 클러스터2 prometheus 는 지연없이 조회 된다.

참고로 클러스터는 2개의 노드에 각각 prometheus pod 1개씩 떠있고, 데이터(메트릭)이 유입되면 각각 똑같은 데이터를 저장하게 된다.

버스트요청 - stress test 로 10개 쿼리 동시 요청 x n 번(round)시 반응
문제가 되는 클러스터1 - prometheus 일부 요청들 지연 응답 또는 연결 취소 발생
문제가 없는 클러스터2 - prometheus 모든 요청 정상 응답

체크 - 쿼리가 무거워서?
무거운 쿼리 1동시 요청은 괜찮았지만 vector(1) 8동시 요청을 하면 23%확률로 스톨된다.

체크 - 클러스터들간의 데이터량(pod) 80 vs 40) 차이나서?
클러스터 양쪽 메트릭을 사이즈 차이가 크지 않다.

체크 - 요청 부하(GC / 스로틀 / 메모리 / 디스크 등)?
파드/노드 지표 12h 그래프를 봐도 cpu, 메모리가 튀지는 않는다.

체크 - 존/랙 업링크등의 문제가 있나?
같은 존/랙 떠있을것으로 추정되는 노드의 grafana 연결은 지연이 없다.

체크 - 진입 노드 외부 구간
문제가 있는 노드1를 통해서 grafana 조회시 지연이 없다.

체크 - LB(VIP)1 경유시 
스톨 발생

체크 - 새로운 LB2 를 생성하고 서비스는 기존 prometheus 로 연결
스톨 발생

체크 - LB 미경유 nodeport 로 바로 조회
스톨 발생

체크 - cilium 패킷 drop 이 발생하나?
cilium 측정(cilium hubbl ui 서비스 portfowrading)시 패킷 드랍 없음

[원인]
클러스터1,2 모두 --web.max-connections 512 (디폴트)로 되어 있다.
문제가 되는 클러스터1 이 클러스터2 200 개 정도에 비해 2배정도 많은 400 개 정도다.

현재 prometheus 연결 수는 config-reload container 를 통해 다음 명령으로 확인
kubectl -context ysoftman1-context -n ysoftman-monitoring-stack \
 exec ysoftman-monitoring-stack-prometheus-1 -c config-reloader -- \
 sh -c 'awk "\$2 ~ /:2382\$/ && \$4==\"01\" {c++} END{print c+0}" /proc/net/tcp /proc/net/tcp6'

또는 다음 prometheus 매트릭/쿼리로 확인할 수 있다.
net_conntrack_listener_conn_accepted_total{listener_name="http"} - net_conntrack_listener_conn_closed_total

클러스터1,2 둘다 batch 로 promql 조회를 하고 있는데 이게 클러스트의 pod 의 비례해서 상대적으로 많은 pod 를 가진 클러스터의 트래픽이 2배가 됐다.
512까지 여유가 없는 상태에서 버스트 요청시 클러스터2 보다 상대적으로 accept 가 빨리 찬다.

1. prometheus 가 들고 있는(accept 된) 연결이 512 에 도달 -> prometheus LimitListener 가 accept 호출 자체를 연결 하나가 닫혀 슬롯이 반환될 때까지 멈춘다. 
2. 그 동안 들어오는 새 연결은 커널이 알아서 3-way handshake 를 완료하고 accept 대기열에 쌓아둔다. 클라이언트 입장에선 connect 성공(tcp= 15ms), 요청까지 전송
3. 하지만 prometheus가 accept 을 안 하니 클라이언트는 first-byte 무한 대기한다. 서버가 끊는 게 아니라 클라이언트가 중단하게 된다.
4. 기존 연결이 닫혀 슬롯이 나면 대기열에서 순서대로 accept 재개되고 밀렸던 연결들이 한꺼번에 처리 된다. 1~7초 동기 방출되는 prometheus 쿼리 자체는 15ms 로 금방 처리된다.

이미 accept 돼 있던 연결(scrape 로 다른 클러스터에 긁어 가고 있는 경우 keep-alive 재사용)은 이 제한과 무관하게 계속 동작한다.
그래서 서버 히스토그램 기존 트래픽은 내내 정상으로 보였고, 신규 연결만 지연으로 관측됐다.

[해결방법]
max-connections 을 늘리자. 디폴트 512는 너무 작고 2048 정도로 늘려도 무리가 없다.
helm chart 로 설정
kube-prometheus-stack:
 prometheusSpec:
  prometheus:
      web:
        maxConnections: 2048

(수동 적용시) argocd syncPolicy.automated (selfHeal 포함) 를 제거해서 auto-sync 끈다.
kubectl -n ysoftman-argocd patch application ysoftman-monitoring-stack \
 --type merge -p '{"spec":{"syncPolicy":{"automated":null}}}'

(수동 적용시) prometheus cr 변경하면 operator 감지해서 적용
kubectl -n ysoftman-monitoring-stack patch prometheus ysoftman-monitoring-stack-prometheus \
 -type merge -p '{"spec":{"web":{"maxConnections":2048}}}'

pod 에 적용 모습
spec:
  containers:
  - args:
    - --web.max-connections=2048

--web.max-connections 는 커맨드라인 플래그라 reload operator가 StatefulSet 템플릿을 바꾸면서 롤링 재시작 된다.(replicas: 2 로 쓰고 있어서 한 대씩 앞 pod 이 ready 된 후 다음 pod 진행)
StatefulSet volumeClaimTemplate 기반이라 pod 삭제와 무관하게 pvc/pv(local-path-retain)는 그대로 남고, 새 pod 가 같은 pvc 에 다시 붙는다.
local-path 특성상 pv 가 노드에 고정돼 있어 pod 도 같은 노드로 다시 스케줄된다.
pod1대 재시작 중 나머지 replica 가 조회/수집을 계속 받아서 전체적으로 메트릭 유실은 없다.

max-connections 을 늘린 후로 스톨이 발생하지 않는다.ㅎ

install harbor

# chart 저장소로 chartmuseum 을 사용중이였는데 dashboard ui 가 없어 찾아 보니 chartmuseum-ui 가 있다.
# 그런데 chartmuseum-ui 는 helm 차트가 없고 현재 차트 업로드가 안되고 조회만 된다고 한다.
# harbor(https://github.com/goharbor/harbor) 는 CNCF 졸업한 프로젝트로 
# chart 외 Open Container Initiative(OCI) 표준을 따르는 컨테이너 이미지 및 기타 아티팩트(helm chart 파일)를 저장하고 
# 프로젝트별 구분 및 RBAC(롤 기반 접근제어)
# trivy(https://github.com/aquasecurity/trivy)를 통한 취약점 파악이 가능하다.
# 그리고 db(postgresql)를 사용해 아티팩트 사용에 대한 히스토리 및 push, pull 카운트도 알 수 있다.
# 물론 UI 도 있고, chartmuseum(3k) 보다 스타수도 26k로 많다.... 해서 harbor 를 설치해보자.

# harbor 2.x 에서는 chart 저장소가 chartmuseum 에서 oci 레지스트리로 변경되었다.
# harbor 2.8 (Apr 17, 2023) 부터 chartmuseum 을 지원하지 않는다.
# chart 는 image(artifact)와 동일하게 관리된다.

# harbor helm chart 다운로드
helm repo add harbor https://helm.goharbor.io
helm repo update
helm fetch harbor/harbor
tar zxvf harbor-1.17.1.tgz
cd harbor

# values.yaml 변경
# ingress, clusterIP, nodePort, loadBalancer (대소문자 구분) 중 하나를 선택하면 나머지 설정들은 스킵된다.
# nodePort 를 사용하면 기존 ClusterIP 타입의 서비스외  NodePort 타입의 harbor 서비스가 더 생성되고 여기서 기존 서비스들로 분배된다.(configmap > nginx.conf 참고)
expose.type: ingress
expose.tls.enabled: false

# harbor 접근할 웹 주소
expose.ingress.hosts.core: https://harbor.ysoftman.test

# docker 이미지나 helm chart 를 push/pull 할때 cli 명령에서 사용될 서버 주소
externalurl: https://harbor.ysoftman.test

# pv,pvc 기본으로 Container Storage Interface (CSI) 로 생성된다.
# pv,pvc 를 사용하면 파드를 삭제했다가 pvc 를 사용하는 새로운 파드를 다시 띄우면, k8s 는 기존 데이터가 남아있는 동일한 노드로 파드를 스케줄링(NodeAffinity)하여 이전 데이터를 그대로 이어받아 사용할 수 있다.
# kubectl get storageclasses 로 csi 가 연결된 스토리지 provisioner 확인
# csi 종류
# rancher.io/local-path 는 k8s 클러스터의 각 노드에 있는 로컬 스토리지를 Persistent Volume(PV)으로 사용해 emptyDir (Ephemeral Storage) 처럼 pod 삭제시 데이터가 날라가는것 방지할 수 있다.
# cinder.csi 는 openstack 환경에서 cinder API 호출하여 새로운 cinder 볼륨을 사용
# reclaimpolicy(pod 삭제시 pv등 정리 방법)는 delete 인 local-path 를 사용
persistence.persistentVolumeClaim.xxxxx.storageClass: "local-path-delete"

# 별도의 nfs 사용할 경우 pv, pvc 리소스를 추가하자.
# pv.yaml
apiVersion: v1
kind: PersistentVolume
metadata:
  name: ysoftman-harbor
spec:
  accessModes:
    - ReadWriteMany
  capacity:
    storage: 10Gi
  nfs:
    path: /xxxxx/ysoftman/harbor
    server: nfs.ysoftman.zzz
  persistentVolumeReclaimPolicy: Retain

# pvc.yaml
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: ysoftman-harbor
  namespace: harbor
spec:
  accessModes:
    - ReadWriteMany
  resources:
    requests:
      storage: 10Gi
  storageClassName: ""
  volumeName: ysoftman-harbor

# values.yaml 설정에서 existingClaim 명시
persistence.persistentVolumeClaim.xxxxx.storageClass: ""
persistence.persistentVolumeClaim.xxxxx.existingClaim: "ysoftman-harbor"

# harbor database(postgreSQL) 데이터 저장소로 nfs 사용시 권한 문제와 디렉터리 비어있지 않음등의 문제가 있다면 로컬 스토리지를 사용하자.
persistence.persistentVolumeClaim.database.storageClass: "local-path-delete"

# image/chart 저장소 타입 filesystem, azure, gcs, s3, swift 중 선택
persistence.imageChartStorage.type: filesystem

# image/chart filesystem 선택시 저장 경로
persistence.imageChartStorage.filesystem.rootdirectory: /storage

# 배포
helm upgrade --install harbor . \
--namespace harbor \
--create-namespace \
--values values.yaml

# 이제 admin / Harbor12345 로그인해서 메뉴에서 admin 암호를 변경(대소,특수,포함 8자리 이상등의 조건)한다.
https://harbor.ysoftman.test

#####

# nodePort 와 ingress 둘 다 사용하기
# ingress 상태로 띄우고
expose.type: ingress

# ingress 리소스를 templates/ysoftman_ingress.yaml 파일로 백업
kubectl get ing -o yaml > templates/ysoftman_ingress.yaml

# nodePort 로 다시 배포(helm upgrade)하면 nodePort 서비스가 추가된다.
expose.type: nodePort

# templates/ysoftman_ingress.yaml > backend service 를 NodePort 서비스명으로 변경 후 다시 배포(helm upgrade)

#####

# harbor database(postgresql)이 nfs로 저장이 안되고, 외부 gcs(google cloud storage), s3, 등을 사용할 수 없다면 덤프파일로 백업하자.

# database container 에 접속해서 psql 로 다음과 같이 사용자명과 db명을 파악하자.

# 로컬에서 접속할 수 있도록 port-foward
kubectl port-forward svc/harbor-database 5432:5432 -n harbor

# pg cli 툴 설치
# postgresql server 와 맞는 버전을 설치해야 한다.
# 14 버전 대신 15 버전으로 설치
brew unlink postgresql@14
brew install postgresql@15

# registry db 덤프
pg_dump -h localhost -p 5432 -U postgres -d registry > harbor_backup_20250807.sql

# 덤프 파일로 복구
psql -h localhost -p 5432 -U postgres -d registry < harbor_backup_20250807.sql

#####

# helm 3.7 부터는 oci registry 연동 시 https 만 지원하고 http 사용지 동작하지 않는다.
# oci 방식으로 저장된 차트는 url 로 다운로드가 할 수 없다.
# harbor ui 에서 올라간 차트 이미지에 대해 helm pull, docker pull 등의 커맨드를 클립보드로 복사하는 기능이 있다.

# harbor registry 로그인
helm registry login -u admin -p Harbor12345 https://harbor.ysoftman.test

# harbor 2.8 chartrepo 방식이 사라져 chartrepo 엔드포인트는 사용할 수 없고 oci 프로토콜로만 사용해야 한다.
# 차트 파일 library 프로젝트에 올리기
helm push ysoftman-chart-0.0.1.tgz oci://harbor.ysoftman.test/library

# library 프로젝트의 ysoftman 차트 다운로드
helm pull oci://harbor.ysoftman.test/library/ysoftman-chart --version 0.0.1

# docker 이미지 library 프로젝트로 태깅
docker tag aaa/ysoftman-image:dev harbor.ysoftman.test/library/ysoftman-image:dev

# docker 이미지 library 프로젝트에 올리기
docker push harbor.ysoftman.test/library/ysoftman:dev

# library 프로젝트의 ysoftman:dev 이미지 다운로드
docker pull harbor.ysoftman.test/library/ysoftman:dev

# harbor 버전 확인(로그인 메뉴 > about)
https://harbor.ysoftman.test/api/v2.0/systeminfo

Prev