← Back to Home

🚨 Self-Service Team Alerting

How five teams share one Kubernetes cluster and one Prometheus yet each owns its own alerts — using ArgoCD for GitOps, a platform team-monitoring chart, and team-owned PrometheusRules routed by the team label.

1. The Scenario: 5 Teams, 1 Cluster

Organization X runs one Kubernetes cluster shared by five teams — A through E. Each team owns a GitLab group with 20–25 projects, its own ArgoCD repo, and its own namespace:

TeamGitLab groupArgoCD repoNamespace
Agroup-agitlab.com/orgx/group-a/argocdns-a
Bgroup-bgitlab.com/orgx/group-b/argocdns-b
Cgroup-cgitlab.com/orgx/group-c/argocdns-c
Dgroup-dgitlab.com/orgx/group-d/argocdns-d
Egroup-egitlab.com/orgx/group-e/argocdns-e

There is one more namespace — platform. It is owned by the platform team (DevOps + platform engineers) and holds the shared infrastructure the other teams depend on: Prometheus, Alertmanager, the Kafka cluster, and the gateways.

The platform team also publishes a common Helm chart that every team uses to ship their services. That chart already creates Deployments, HTTPRoutes, and — importantly — supports ServiceMonitors, so a team can opt their service into Prometheus scraping just by declaring a monitor.

Now we want the same self-service experience for alerts. The goal: each team manages its own alert rules in its own ArgoCD repo, through a dedicated app called <team-a-monitoring>, without ever touching the platform's Prometheus or Alertmanager.

Team repo group-a/argocd owns the alert rules
↓
ArgoCD auto-discovers the team-a-monitoring folder
↓
ns-a a PrometheusRule CR is created
↓
Prometheus in platform, picks it up via ruleSelector
↓
Alertmanager routes by the team label
↓
#team-a-alerts Slack / PagerDuty / email
💡 The one-sentence idea: the platform team owns the alerting plumbing (Prometheus, Alertmanager, routing, the chart); each team owns the alerting content (the PrometheusRules) in its own repo and namespace. GitOps makes that content reviewed, versioned, and rollback-able.

2. Who Owns What

The single most important decision is the ownership boundary. If teams can't touch the shared Prometheus, and the platform team can't keep up with every team's alert, the line has to be drawn in a place that gives each side autonomy without giving either side a footgun.

ConcernPlatform teamTeams A–E
Prometheus + Alertmanager (CRDs, config, upgrades)Owns & operatesConsumes
Routing rules (which receiver gets which alert)OwnsRequests changes
Notification channels (Slack, PagerDuty, email)Owns the integrationsManages channel membership
The team-monitoring chartOwns & versionsConsumes as a dependency
Alert rule content (PrometheusRule)Defines conventionsOwns, in their repo
ServiceMonitor definitionsProvides the chart supportOwns, per service

This split is why the design scales: the platform team changes the chart or the routing once, and every team benefits. Teams change their rules as often as they like, in their own review flow, without a platform-ticket bottleneck.

⚠️ Anti-pattern to avoid: a single shared alerts.yaml that every team edits, in a repo only the platform team can approve. It becomes a merge-conflict hot spot and a review bottleneck — exactly the opposite of self-service.

3. The Repo Layout

Your central ArgoCD is already wired to treat each folder in a team's ArgoCD repo as a Helm app. The convention is:

argo-repo/
└── <app-1>/
    ├── Chart.yaml          # the chart for this app
    ├── values.yaml         # default values
    ├── dev/values.yaml     # env-specific overrides
    ├── prod/values.yaml
    └── templates/          # chart templates

That means we do not need to write a single Application or ApplicationSet. To add alerting to Team A, the team simply adds one more folder that follows the same convention:

team-a-argocd/                 # gitlab.com/orgx/group-a/argocd
├── api/                      # service app (Deployment + HTTPRoute + ServiceMonitor)
│   ├── Chart.yaml
│   ├── values.yaml
│   └── templates/...
├── worker/                   # another service app
│   └── ...
└── team-a-monitoring/        # ← NEW: the team's alert-rules app
    ├── Chart.yaml            # depends on the platform `team-monitoring` chart
    ├── values.yaml           # team: a, namespace: ns-a, rules: [...]
    ├── dev/values.yaml       # env-specific alert thresholds
    └── prod/values.yaml

The key properties of this layout:

4. The Platform team-monitoring Chart

The platform team publishes a tiny, reusable Helm chart that turns a team's values.yaml into well-formed PrometheusRule (and optionally ServiceMonitor) objects. This is where all the conventions live once: required labels, default severity, naming, and namespace placement.

Because the chart is the single enforcement point, a team can get it wrong locally but the rendered output will still carry the right team, environment, and default severity labels.

Chart.yaml

apiVersion: v2
name: team-monitoring
description: Renders team-owned PrometheusRule and ServiceMonitor objects
type: application
version: 1.5.0
appVersion: "1.0"

values.yaml (platform defaults)

# Platform-provided defaults. Teams override these in their own app.
team: ""            # REQUIRED: a | b | c | d | e
namespace: ""       # REQUIRED: the team's namespace, e.g. ns-a
environment: prod   # dev | staging | prod

# Extra labels injected into every object, in addition to `team` and
# `role: alert-rules` which the chart always adds itself.
labels:
  managed-by: argo-cd

# One entry per service/app. Each entry becomes a PrometheusRule.
rules: []
#  - name: api
#    rules:
#      - alert: ApiHighErrorRate
#        expr: >-
#          sum(rate(http_requests_total{job="api", code=~"5.."}[5m])) by (service)
#          / sum(rate(http_requests_total{job="api"}[5m])) by (service) > 0.05
#        for: 5m
#        labels:
#          severity: critical
#        annotations:
#          summary: "API 5xx error rate is above 5%"
#          description: "{{ $labels.service }} has a 5xx rate of {{ $value | humanizePercentage }}."
#          runbook_url: "https://wiki.example.com/runbooks/api-high-error-rate"

serviceMonitors: []

templates/_helpers.tpl

{{- define "team-monitoring.labels" -}}
app.kubernetes.io/name: team-monitoring
app.kubernetes.io/managed-by: {{ .Release.Service }}
team: {{ required "team is required" .Values.team | quote }}
{{- end -}}

templates/prometheusrule.yaml

The heart of the chart. For every entry under rules, it emits a PrometheusRule in the team's namespace and — crucially — injects team, environment, and a default severity into every individual rule's labels:

{{- $team := required "team is required" .Values.team -}}
{{- $ns := required "namespace is required" .Values.namespace -}}
{{- $env := .Values.environment | default "prod" -}}
{{- range .Values.rules }}
---
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
  name: {{ printf "%s-%s" $team .name | trunc 63 | trimSuffix "-" }}
  namespace: {{ $ns }}
  labels:
    role: alert-rules
    team: {{ $team | quote }}
    {{- with $.Values.labels }}
    {{- toYaml . | nindent 4 }}
    {{- end }}
spec:
  groups:
    - name: {{ printf "%s-%s" $team .name }}
      rules:
        {{- range .rules }}
        - alert: {{ .alert | quote }}
          expr: {{ .expr | quote }}
          {{- with .for }}
          for: {{ . }}
          {{- end }}
          labels:
            team: {{ $team | quote }}
            environment: {{ $env | quote }}
            severity: {{ dig "severity" "warning" .labels | quote }}
            {{- range $k, $v := dig "labels" (dict) . }}
            {{- if and (ne $k "severity") (ne $k "team") (ne $k "environment") }}
            {{ $k }}: {{ $v | quote }}
            {{- end }}
            {{- end }}
          {{- with .annotations }}
          annotations:
            {{- toYaml . | nindent 12 }}
          {{- end }}
        {{- end }}
{{- end }}

templates/servicemonitor.yaml

Teams usually declare their ServiceMonitors in the service app (the common chart already supports it), but the monitoring app can also own them so a team can add scraping without touching the service chart:

{{- range .Values.serviceMonitors }}
---
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
  name: {{ $.Values.team }}-{{ .name }}
  namespace: {{ $.Values.namespace }}
  labels:
    team: {{ $.Values.team | quote }}
    {{- with $.Values.labels }}
    {{- toYaml . | nindent 4 }}
    {{- end }}
spec:
  selector:
    matchLabels: {{- toYaml .selector | nindent 6 }}
  endpoints: {{- toYaml .endpoints | nindent 4 }}
{{- end }}
💡 Conventions are code: because the chart injects team and a default severity, a rule that forgets to set severity still lands as warning and still routes to the right team — no silent black hole.

5. A Team's Monitoring App

Team A creates a folder team-a-monitoring/ with a two-line Chart.yaml that depends on the platform chart, and a values.yaml that is only their rules.

team-a-monitoring/Chart.yaml

apiVersion: v2
name: team-a-monitoring
description: Team A alert rules
type: application
version: 1.0.0
dependencies:
  - name: team-monitoring
    version: 1.5.0
    repository: oci://registry.example.com/platform/charts

team-a-monitoring/values.yaml

team: a
namespace: ns-a
environment: prod

rules:
  - name: api
    rules:
      - alert: ApiHighErrorRate
        expr: >-
          sum(rate(http_requests_total{job="api", code=~"5.."}[5m])) by (service)
          / sum(rate(http_requests_total{job="api"}[5m])) by (service) > 0.05
        for: 5m
        labels:
          severity: critical
        annotations:
          summary: "API 5xx error rate is above 5%"
          description: "{{ $labels.service }} has a 5xx rate of {{ $value | humanizePercentage }}."
          runbook_url: "https://wiki.example.com/runbooks/api-high-error-rate"

      - alert: ApiP99LatencyHigh
        expr: >-
          histogram_quantile(0.99,
            sum(rate(http_request_duration_seconds_bucket{job="api"}[5m])) by (le, service))
          > 0.5
        for: 10m
        labels:
          severity: warning
        annotations:
          summary: "API p99 latency is above 500ms"
          runbook_url: "https://wiki.example.com/runbooks/api-latency"

  - name: worker
    rules:
      - alert: WorkerQueueDepth
        expr: sum(worker_queue_depth{job="worker"}) by (queue) > 1000
        for: 15m
        labels:
          severity: warning
        annotations:
          summary: "Worker queue depth is above 1000"
          runbook_url: "https://wiki.example.com/runbooks/worker-queue"

Try it yourself — pick the knobs and see the exact values.yaml snippet a team would paste into their rules: list:

Interactive: build an alert rule

# choose values above, then click "Generate rule"

💡 For 20–25 projects: keep one rules entry per service (not one giant rule list), and let dev/values.yaml relax thresholds while prod/values.yaml keeps them strict. A CODEOWNERS file in the repo can route review of each service's rules to its owning squad.

6. Prometheus & Alertmanager Wiring

Teams only create PrometheusRule objects. The platform side does two things once to make those objects actually work.

a) Prometheus discovers team rules

The platform's Prometheus CRD is told to load any PrometheusRule labelled role: alert-rules, from any namespace labelled monitoring: team:

apiVersion: monitoring.coreos.com/v1
kind: Prometheus
metadata:
  name: platform
  namespace: platform
spec:
  ruleSelector:
    matchLabels:
      role: alert-rules
  ruleNamespaceSelector:
    matchLabels:
      monitoring: team

Each team namespace gets the monitoring: team label once (a platform-controlled namespace operation), which keeps the blast radius explicit:

kubectl label namespace ns-a monitoring=team
kubectl label namespace ns-b monitoring=team
# ... ns-c, ns-d, ns-e

b) Alertmanager routes by the team label

Routing stays centralized and platform-owned. Every alert already carries a team label (injected by the chart), so the routing tree is just a per-team match plus a catch-all for infra alerts that have no team label:

route:
  group_by: ["alertname", "team"]
  group_wait: 30s
  group_interval: 5m
  repeat_interval: 4h
  receiver: platform-default
  routes:
    - match:
        team: a
      receiver: team-a-alerts
    - match:
        team: b
      receiver: team-b-alerts
    - match:
        team: c
      receiver: team-c-alerts
    - match:
        team: d
      receiver: team-d-alerts
    - match:
        team: e
      receiver: team-e-alerts
    - match:
        severity: critical
      receiver: platform-oncall

receivers:
  - name: platform-default
    slack_configs:
      - channel: "#platform-alerts"
  - name: team-a-alerts
    slack_configs:
      - channel: "#team-a-alerts"
  - name: team-b-alerts
    slack_configs:
      - channel: "#team-b-alerts"
  - name: team-c-alerts
    slack_configs:
      - channel: "#team-c-alerts"
  - name: team-d-alerts
    slack_configs:
      - channel: "#team-d-alerts"
  - name: team-e-alerts
    slack_configs:
      - channel: "#team-e-alerts"
  - name: platform-oncall
    pagerduty_configs:
      - service_key: "<PD_INTEGRATION_KEY>"
    slack_configs:
      - channel: "#platform-oncall"

Watch how a fired alert moves through that tree:

Interactive: route a fired alert

Select an alert above to see which receiver it reaches and why.

💡 Why central routing? because the platform team owns the Slack/PagerDuty integrations and the on-call escalation policy. Teams own the rules; the platform owns who gets paged. Add continue: true to the severity: critical route if you want critical team alerts to also page the platform on-call.

7. Alert Lifecycle

From a line of YAML in a team repo to a Slack message, an alert travels this path. Walk it step by step:

Interactive: follow an alert from commit to notification

  1. An engineer adds a rule under team-a-monitoring/values.yaml and opens a merge request.
  2. A teammate reviews it; CI runs helm lint and helm template to catch template and YAML errors.
  3. The MR merges to the default branch of gitlab.com/orgx/group-a/argocd.
  4. Central ArgoCD detects the change to the team-a-monitoring folder and refreshes that app.
  5. ArgoCD renders the chart and applies a PrometheusRule named a-api in ns-a.
  6. Prometheus (in platform) matches role: alert-rules and starts evaluating the rule every 30s.
  7. The expression stays true past the for: 5m threshold, so the alert fires and is sent to Alertmanager.
  8. Alertmanager matches team: a, routes to team-a-alerts, and posts to #team-a-alerts.
⚠️ The silent-failure trap: if a team's rule is missing the role: alert-rules label or the namespace is missing the monitoring: team label, Prometheus simply ignores it — no error. That's why the chart injects the label and why verification (next section) is a hard step, not an afterthought.

8. Guardrails & Best Practices

Self-service works only with rails. Here are the ones worth putting in place from day one.

RBAC: teams touch their own namespace only

The platform should scope each team's write access to its own namespace so a bad rule in ns-a can never rewrite Prometheus in platform or another team's rules:

apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
  name: alert-rule-author
  namespace: ns-a
rules:
  - apiGroups: ["monitoring.coreos.com"]
    resources: ["prometheusrules", "servicemonitors"]
    verbs: ["get", "list", "watch", "create", "update", "patch", "delete"]

On the ArgoCD side, each app's AppProject should already be restricting the destination to the team's namespace — the folder-per-app convention pairs naturally with that scoping.

CI: fail fast on malformed rules

# .gitlab-ci.yml (team repo)
lint-alerts:
  image: dtzar/helm-kubectl
  script:
    - helm dependency update team-a-monitoring
    - helm lint team-a-monitoring
    - helm template team-a-monitoring team-a-monitoring --namespace ns-a

Conventions that keep the noise down

💡 Summary: platform owns the plumbing, teams own the content, the chart owns the conventions, ArgoCD owns the delivery, and the team label owns the routing. That four-way split is what makes five teams on one cluster feel like five independent clusters.

9. Deploy It

Here is the rollout, in order. The platform steps happen once; the team step happens per team.

Platform team (once)

# 1. Publish the chart to your OCI registry
helm package team-monitoring
helm push team-monitoring-1.5.0.tgz oci://registry.example.com/platform/charts

# 2. Configure Prometheus to discover team rules
kubectl apply -f prometheus.yaml   # ruleSelector + ruleNamespaceSelector

# 3. Label team namespaces as eligible
kubectl label namespace ns-a monitoring=team
kubectl label namespace ns-b monitoring=team
kubectl label namespace ns-c monitoring=team
kubectl label namespace ns-d monitoring=team
kubectl label namespace ns-e monitoring=team

# 4. Apply the Alertmanager routing config
kubectl apply -f alertmanager-config.yaml

Each team (per team)

# 1. Add the monitoring app to the team repo
mkdir team-a-monitoring
#   ... add Chart.yaml + values.yaml as shown above ...

# 2. Preview locally before committing
helm dependency update team-a-monitoring
helm template team-a-monitoring team-a-monitoring --namespace ns-a

# 3. Commit and push — central ArgoCD discovers and syncs it automatically

Verification

# The rule object exists in the team namespace
kubectl get prometheusrules -n ns-a
kubectl get prometheusrule a-api -n ns-a -o yaml

# Prometheus actually loaded it (UI: Status → Rules, or via the API)
kubectl -n platform port-forward svc/prometheus-platform 9090:9090
curl -s localhost:9090/api/v1/rules | jq '.data.groups[] | select(.name | contains("a-api"))'

# Alertmanager has the routing tree
kubectl -n platform get secret alertmanager-platform -o jsonpath='{.data.alertmanager\.yaml}' | base64 -d

# Optionally fire a synthetic alert and confirm it lands in #team-a-alerts
💡 Roll out incrementally: start with one team and one warning-only rule, confirm the end-to-end path (rule → Prometheus → Alertmanager → channel), then open the floodgates. The whole point is that every subsequent team and every subsequent rule is just another merge request.
Copied to clipboard!