How five teams share one Kubernetes cluster and one Prometheus yet each owns its own alerts — using ArgoCD for GitOps, a platform team-monitoring chart, and team-owned PrometheusRules routed by the team label.
Organization X runs one Kubernetes cluster shared by five teams — A through E. Each team owns a GitLab group with 20–25 projects, its own ArgoCD repo, and its own namespace:
| Team | GitLab group | ArgoCD repo | Namespace |
|---|---|---|---|
| A | group-a | gitlab.com/orgx/group-a/argocd | ns-a |
| B | group-b | gitlab.com/orgx/group-b/argocd | ns-b |
| C | group-c | gitlab.com/orgx/group-c/argocd | ns-c |
| D | group-d | gitlab.com/orgx/group-d/argocd | ns-d |
| E | group-e | gitlab.com/orgx/group-e/argocd | ns-e |
There is one more namespace — platform. It is owned by the platform team (DevOps + platform engineers) and holds the shared infrastructure the other teams depend on: Prometheus, Alertmanager, the Kafka cluster, and the gateways.
The platform team also publishes a common Helm chart that every team uses to ship their services. That chart already creates Deployments, HTTPRoutes, and — importantly — supports ServiceMonitors, so a team can opt their service into Prometheus scraping just by declaring a monitor.
Now we want the same self-service experience for alerts. The goal: each team manages its own alert rules in its own ArgoCD repo, through a dedicated app called <team-a-monitoring>, without ever touching the platform's Prometheus or Alertmanager.
group-a/argocd owns the alert rulesteam-a-monitoring folderns-a a PrometheusRule CR is createdplatform, picks it up via ruleSelectorteam labelPrometheusRules) in its own repo and namespace. GitOps makes that content reviewed, versioned, and rollback-able.The single most important decision is the ownership boundary. If teams can't touch the shared Prometheus, and the platform team can't keep up with every team's alert, the line has to be drawn in a place that gives each side autonomy without giving either side a footgun.
| Concern | Platform team | Teams A–E |
|---|---|---|
| Prometheus + Alertmanager (CRDs, config, upgrades) | Owns & operates | Consumes |
| Routing rules (which receiver gets which alert) | Owns | Requests changes |
| Notification channels (Slack, PagerDuty, email) | Owns the integrations | Manages channel membership |
The team-monitoring chart | Owns & versions | Consumes as a dependency |
Alert rule content (PrometheusRule) | Defines conventions | Owns, in their repo |
| ServiceMonitor definitions | Provides the chart support | Owns, per service |
This split is why the design scales: the platform team changes the chart or the routing once, and every team benefits. Teams change their rules as often as they like, in their own review flow, without a platform-ticket bottleneck.
alerts.yaml that every team edits, in a repo only the platform team can approve. It becomes a merge-conflict hot spot and a review bottleneck — exactly the opposite of self-service.Your central ArgoCD is already wired to treat each folder in a team's ArgoCD repo as a Helm app. The convention is:
argo-repo/
└── <app-1>/
├── Chart.yaml # the chart for this app
├── values.yaml # default values
├── dev/values.yaml # env-specific overrides
├── prod/values.yaml
└── templates/ # chart templates
That means we do not need to write a single Application or ApplicationSet. To add alerting to Team A, the team simply adds one more folder that follows the same convention:
team-a-argocd/ # gitlab.com/orgx/group-a/argocd
├── api/ # service app (Deployment + HTTPRoute + ServiceMonitor)
│ ├── Chart.yaml
│ ├── values.yaml
│ └── templates/...
├── worker/ # another service app
│ └── ...
└── team-a-monitoring/ # ← NEW: the team's alert-rules app
├── Chart.yaml # depends on the platform `team-monitoring` chart
├── values.yaml # team: a, namespace: ns-a, rules: [...]
├── dev/values.yaml # env-specific alert thresholds
└── prod/values.yaml
The key properties of this layout:
team-a-monitoring is about alerts only — it does not redeploy the services themselves.dev/ and prod/ values convention works for thresholds (a stricter SLO in prod than in dev).team-monitoring ChartThe platform team publishes a tiny, reusable Helm chart that turns a team's values.yaml into well-formed PrometheusRule (and optionally ServiceMonitor) objects. This is where all the conventions live once: required labels, default severity, naming, and namespace placement.
Because the chart is the single enforcement point, a team can get it wrong locally but the rendered output will still carry the right team, environment, and default severity labels.
apiVersion: v2
name: team-monitoring
description: Renders team-owned PrometheusRule and ServiceMonitor objects
type: application
version: 1.5.0
appVersion: "1.0"
# Platform-provided defaults. Teams override these in their own app.
team: "" # REQUIRED: a | b | c | d | e
namespace: "" # REQUIRED: the team's namespace, e.g. ns-a
environment: prod # dev | staging | prod
# Extra labels injected into every object, in addition to `team` and
# `role: alert-rules` which the chart always adds itself.
labels:
managed-by: argo-cd
# One entry per service/app. Each entry becomes a PrometheusRule.
rules: []
# - name: api
# rules:
# - alert: ApiHighErrorRate
# expr: >-
# sum(rate(http_requests_total{job="api", code=~"5.."}[5m])) by (service)
# / sum(rate(http_requests_total{job="api"}[5m])) by (service) > 0.05
# for: 5m
# labels:
# severity: critical
# annotations:
# summary: "API 5xx error rate is above 5%"
# description: "{{ $labels.service }} has a 5xx rate of {{ $value | humanizePercentage }}."
# runbook_url: "https://wiki.example.com/runbooks/api-high-error-rate"
serviceMonitors: []
{{- define "team-monitoring.labels" -}}
app.kubernetes.io/name: team-monitoring
app.kubernetes.io/managed-by: {{ .Release.Service }}
team: {{ required "team is required" .Values.team | quote }}
{{- end -}}
The heart of the chart. For every entry under rules, it emits a PrometheusRule in the team's namespace and — crucially — injects team, environment, and a default severity into every individual rule's labels:
{{- $team := required "team is required" .Values.team -}}
{{- $ns := required "namespace is required" .Values.namespace -}}
{{- $env := .Values.environment | default "prod" -}}
{{- range .Values.rules }}
---
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: {{ printf "%s-%s" $team .name | trunc 63 | trimSuffix "-" }}
namespace: {{ $ns }}
labels:
role: alert-rules
team: {{ $team | quote }}
{{- with $.Values.labels }}
{{- toYaml . | nindent 4 }}
{{- end }}
spec:
groups:
- name: {{ printf "%s-%s" $team .name }}
rules:
{{- range .rules }}
- alert: {{ .alert | quote }}
expr: {{ .expr | quote }}
{{- with .for }}
for: {{ . }}
{{- end }}
labels:
team: {{ $team | quote }}
environment: {{ $env | quote }}
severity: {{ dig "severity" "warning" .labels | quote }}
{{- range $k, $v := dig "labels" (dict) . }}
{{- if and (ne $k "severity") (ne $k "team") (ne $k "environment") }}
{{ $k }}: {{ $v | quote }}
{{- end }}
{{- end }}
{{- with .annotations }}
annotations:
{{- toYaml . | nindent 12 }}
{{- end }}
{{- end }}
{{- end }}
Teams usually declare their ServiceMonitors in the service app (the common chart already supports it), but the monitoring app can also own them so a team can add scraping without touching the service chart:
{{- range .Values.serviceMonitors }}
---
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
name: {{ $.Values.team }}-{{ .name }}
namespace: {{ $.Values.namespace }}
labels:
team: {{ $.Values.team | quote }}
{{- with $.Values.labels }}
{{- toYaml . | nindent 4 }}
{{- end }}
spec:
selector:
matchLabels: {{- toYaml .selector | nindent 6 }}
endpoints: {{- toYaml .endpoints | nindent 4 }}
{{- end }}
team and a default severity, a rule that forgets to set severity still lands as warning and still routes to the right team — no silent black hole.Team A creates a folder team-a-monitoring/ with a two-line Chart.yaml that depends on the platform chart, and a values.yaml that is only their rules.
apiVersion: v2
name: team-a-monitoring
description: Team A alert rules
type: application
version: 1.0.0
dependencies:
- name: team-monitoring
version: 1.5.0
repository: oci://registry.example.com/platform/charts
team: a
namespace: ns-a
environment: prod
rules:
- name: api
rules:
- alert: ApiHighErrorRate
expr: >-
sum(rate(http_requests_total{job="api", code=~"5.."}[5m])) by (service)
/ sum(rate(http_requests_total{job="api"}[5m])) by (service) > 0.05
for: 5m
labels:
severity: critical
annotations:
summary: "API 5xx error rate is above 5%"
description: "{{ $labels.service }} has a 5xx rate of {{ $value | humanizePercentage }}."
runbook_url: "https://wiki.example.com/runbooks/api-high-error-rate"
- alert: ApiP99LatencyHigh
expr: >-
histogram_quantile(0.99,
sum(rate(http_request_duration_seconds_bucket{job="api"}[5m])) by (le, service))
> 0.5
for: 10m
labels:
severity: warning
annotations:
summary: "API p99 latency is above 500ms"
runbook_url: "https://wiki.example.com/runbooks/api-latency"
- name: worker
rules:
- alert: WorkerQueueDepth
expr: sum(worker_queue_depth{job="worker"}) by (queue) > 1000
for: 15m
labels:
severity: warning
annotations:
summary: "Worker queue depth is above 1000"
runbook_url: "https://wiki.example.com/runbooks/worker-queue"
Try it yourself — pick the knobs and see the exact values.yaml snippet a team would paste into their rules: list:
rules entry per service (not one giant rule list), and let dev/values.yaml relax thresholds while prod/values.yaml keeps them strict. A CODEOWNERS file in the repo can route review of each service's rules to its owning squad.Teams only create PrometheusRule objects. The platform side does two things once to make those objects actually work.
The platform's Prometheus CRD is told to load any PrometheusRule labelled role: alert-rules, from any namespace labelled monitoring: team:
apiVersion: monitoring.coreos.com/v1
kind: Prometheus
metadata:
name: platform
namespace: platform
spec:
ruleSelector:
matchLabels:
role: alert-rules
ruleNamespaceSelector:
matchLabels:
monitoring: team
Each team namespace gets the monitoring: team label once (a platform-controlled namespace operation), which keeps the blast radius explicit:
kubectl label namespace ns-a monitoring=team
kubectl label namespace ns-b monitoring=team
# ... ns-c, ns-d, ns-e
team labelRouting stays centralized and platform-owned. Every alert already carries a team label (injected by the chart), so the routing tree is just a per-team match plus a catch-all for infra alerts that have no team label:
route:
group_by: ["alertname", "team"]
group_wait: 30s
group_interval: 5m
repeat_interval: 4h
receiver: platform-default
routes:
- match:
team: a
receiver: team-a-alerts
- match:
team: b
receiver: team-b-alerts
- match:
team: c
receiver: team-c-alerts
- match:
team: d
receiver: team-d-alerts
- match:
team: e
receiver: team-e-alerts
- match:
severity: critical
receiver: platform-oncall
receivers:
- name: platform-default
slack_configs:
- channel: "#platform-alerts"
- name: team-a-alerts
slack_configs:
- channel: "#team-a-alerts"
- name: team-b-alerts
slack_configs:
- channel: "#team-b-alerts"
- name: team-c-alerts
slack_configs:
- channel: "#team-c-alerts"
- name: team-d-alerts
slack_configs:
- channel: "#team-d-alerts"
- name: team-e-alerts
slack_configs:
- channel: "#team-e-alerts"
- name: platform-oncall
pagerduty_configs:
- service_key: "<PD_INTEGRATION_KEY>"
slack_configs:
- channel: "#platform-oncall"
Watch how a fired alert moves through that tree:
continue: true to the severity: critical route if you want critical team alerts to also page the platform on-call.From a line of YAML in a team repo to a Slack message, an alert travels this path. Walk it step by step:
role: alert-rules label or the namespace is missing the monitoring: team label, Prometheus simply ignores it — no error. That's why the chart injects the label and why verification (next section) is a hard step, not an afterthought.Self-service works only with rails. Here are the ones worth putting in place from day one.
The platform should scope each team's write access to its own namespace so a bad rule in ns-a can never rewrite Prometheus in platform or another team's rules:
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
name: alert-rule-author
namespace: ns-a
rules:
- apiGroups: ["monitoring.coreos.com"]
resources: ["prometheusrules", "servicemonitors"]
verbs: ["get", "list", "watch", "create", "update", "patch", "delete"]
On the ArgoCD side, each app's AppProject should already be restricting the destination to the team's namespace — the folder-per-app convention pairs naturally with that scoping.
# .gitlab-ci.yml (team repo)
lint-alerts:
image: dtzar/helm-kubectl
script:
- helm dependency update team-a-monitoring
- helm lint team-a-monitoring
- helm template team-a-monitoring team-a-monitoring --namespace ns-a
<Service><Condition> — e.g. ApiHighErrorRate, not Alert1.severity (critical / warning / info). The chart defaults to warning.for long enough to avoid flapping. A 5xx rate needs for: 5m; a disk-full needs less.summary, description, and runbook_url. An alert without a runbook is a notification, not an actionable alert.by/without so one incident doesn't page once per pod.team label owns the routing. That four-way split is what makes five teams on one cluster feel like five independent clusters.Here is the rollout, in order. The platform steps happen once; the team step happens per team.
# 1. Publish the chart to your OCI registry
helm package team-monitoring
helm push team-monitoring-1.5.0.tgz oci://registry.example.com/platform/charts
# 2. Configure Prometheus to discover team rules
kubectl apply -f prometheus.yaml # ruleSelector + ruleNamespaceSelector
# 3. Label team namespaces as eligible
kubectl label namespace ns-a monitoring=team
kubectl label namespace ns-b monitoring=team
kubectl label namespace ns-c monitoring=team
kubectl label namespace ns-d monitoring=team
kubectl label namespace ns-e monitoring=team
# 4. Apply the Alertmanager routing config
kubectl apply -f alertmanager-config.yaml
# 1. Add the monitoring app to the team repo
mkdir team-a-monitoring
# ... add Chart.yaml + values.yaml as shown above ...
# 2. Preview locally before committing
helm dependency update team-a-monitoring
helm template team-a-monitoring team-a-monitoring --namespace ns-a
# 3. Commit and push — central ArgoCD discovers and syncs it automatically
# The rule object exists in the team namespace
kubectl get prometheusrules -n ns-a
kubectl get prometheusrule a-api -n ns-a -o yaml
# Prometheus actually loaded it (UI: Status → Rules, or via the API)
kubectl -n platform port-forward svc/prometheus-platform 9090:9090
curl -s localhost:9090/api/v1/rules | jq '.data.groups[] | select(.name | contains("a-api"))'
# Alertmanager has the routing tree
kubectl -n platform get secret alertmanager-platform -o jsonpath='{.data.alertmanager\.yaml}' | base64 -d
# Optionally fire a synthetic alert and confirm it lands in #team-a-alerts
warning-only rule, confirm the end-to-end path (rule → Prometheus → Alertmanager → channel), then open the floodgates. The whole point is that every subsequent team and every subsequent rule is just another merge request.