DEV Community

疏影
疏影

Posted on

Argo CD Application Set for 200 Microservices: What Breaks at Scale

Argo CD Application Set for 200 Microservices: What Breaks at Scale

We scaled Argo CD from 30 apps to 230 across 12 clusters in 4 months. Here's the Application Set pattern that survived, what broke, and the metrics that told us we were hitting limits.

Why Application Set

Plain Application CRDs are fine until you have more than one cluster + one repo + one deploy per service. At 30 apps, we had 30 Application manifests per cluster × 4 clusters = 120 files in clusters/*/apps/. Drift in template values. Engineers editing the wrong file.

Application Set generates Applications from declarative generators:

apiVersion: argoproj.io/v1alpha1
kind: ApplicationSet
metadata:
  name: cluster-services
spec:
  generators:
    - list:
        elements:
          - cluster: prod-us
              url: https://k8s-prod-us.example.com
          - cluster: prod-eu
              url: https://k8s-prod-eu.example.com
  template:
    metadata:
      name: '{{cluster}}-services'
    spec:
      project: default
      source:
        repoURL: https://git.example.com/services
        targetRevision: HEAD
        path: 'charts/service'
      destination:
        server: '{{url}}'
        namespace: default
Enter fullscreen mode Exit fullscreen mode

Add a new cluster → add a line to the list. Argo CD creates the Application. No PR to merge. No drift.

The 4 patterns that scale

1. Git directory generator

generators:
  - git:
      repoURL: https://git.example.com/services
      directories:
        - path: charts/*
Enter fullscreen mode Exit fullscreen mode

Each chart in charts/* becomes an Application. Path repo must contain an app-of-apps pattern: each chart has a Chart.yaml, values.yaml, and templates/. When you add charts/new-service/, Argo CD auto-creates Application/new-service on next sync.

2. Matrix generator (cluster × service)

generators:
  - matrix:
      generators:
        - git:
            repoURL: https://git.example.com/services
            directories:
              - path: charts/*
        - list:
            elements:
              - cluster: prod-us
              - cluster: prod-eu
Enter fullscreen mode Exit fullscreen mode

This creates prod-us/payments, prod-eu/payments, prod-us/orders, etc. — every combination. But be careful: the matrix grows multiplicatively. 30 services × 4 clusters = 120 Applications. Argo CD controller starts to struggle at 300+ Applications per cluster.

3. Cluster decision resource

If you have an external cluster registry (e.g., a Cluster CRD in a management cluster), use the cluster decision resource generator:

generators:
  - clusterDecisionResource:
      configMapName: argocd-clusters
      requeueAfterSeconds: 30
Enter fullscreen mode Exit fullscreen mode

This watches a ConfigMap with cluster metadata and re-syncs Applications when clusters are added/removed. We use this for ephemeral preview environments.

4. PR generator for preview envs

generators:
  - pullRequest:
      github:
        owner: myorg
        repo: services
        labels:
          - preview-env
      requeueAfterSeconds: 60
Enter fullscreen mode Exit fullscreen mode

Every PR with label preview-env creates a new Application in the preview cluster. PR closed → app pruned. Invaluable for QA workflows.

The 3 things that break at scale

1. Application controller CPU

Argo CD has a single controller pod by default. At ~200 Applications, it consumes 2.5 CPU cores and 4GB RAM. We scaled to 3 replicas with --sharding enabled.

apiVersion: apps/v1
kind: Deployment
metadata:
  name: argocd-application-controller
spec:
  replicas: 3
  template:
    spec:
      containers:
        - name: argocd-application-controller
          args:
            - --sharding=hash
            - --shard=0
Enter fullscreen mode Exit fullscreen mode

Each shard handles a subset of Applications based on hash. We run 3 shards × 200 apps each.

2. Repo server cache

argocd-repo-server caches manifests. At 230 apps, the default 5GB PVC isn't enough. We increased to 50GB and enabled --redis for shared cache across replicas.

3. Sync windows + drift detection

Application Set syncs all 230 apps every 3 minutes b^ default. We saw throttling from our GitHub API at 250+ apps. We now use sync waves to batch:

metadata:
  annotations:
    argocd.argoproj.io/sync-wave: "5"
Enter fullscreen mode Exit fullscreen mode

Apps with wave 5 sync last. Apps with wave 0 sync first. Critical infrastructure → wave 0. Application workloads → wave 5.

The metrics we watch

  • argocd_application_info Prometheus metric — total apps by health status.
  • argocd_application_sync_total — sync rate. Target: <5 minutes for any new Application to appear.
  • argocd_repo_server_request_duration_seconds — repo server latency. Target p99 <2s.

If sync latency p99 climbs above 5s, it's time to add a repo server replica. If app controller CPU is pegged, add a shard.

On scaling dev tools

If your Argo CD syncs WebDAV-mounted Helm charts from a remote source repo, ScsDriver WebDAV mount tool for Windows handles the Windows side of the dev workflow — useful for hybrid teams where Windows engineers write Helm charts in VS Code against a remote repo, then commit to a Linux CI server.


How many Applications do you run? And what controller / sharding strategy works for you?

Top comments (0)