Skip to content
 
 

Latest commit

 

History

3,725 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Armada logo

One API. Any number of clusters. Millions of jobs.

The open-source batch job meta-scheduler that makes Kubernetes work at scale.

CircleCI Artifact Hub LFX Health Score OpenSSF Best Practices

Website · Quickstart · Documentation · Slack


What is Armada?

Kubernetes was built for services. Armada was built for batch.

When your job volume exceeds what a single cluster can handle, you need a control plane that sits above your fleet — routing jobs intelligently, fairly, and at scale. Armada is that layer.

Armada solves the problems Kubernetes wasn't designed to handle:

  • No job queue — Kubernetes has no concept of ordering. Jobs compete for resources with no fairness guarantees. Armada adds a proper queue with priority, fair-share, and rate limiting.
  • No multi-cluster coordination — Each Kubernetes cluster is an island. Armada routes jobs across as many clusters as you need from a single API.
  • Fine grained gang-scheduling — Distributed jobs that need all workers to start together (MPI, PyTorch, Spark) are either fully scheduled or held in queue. Armada's implementation is battle-tested at scale with deep fairness and preemption integration.
  • No fairness across teams — One team can starve everyone else. Armada enforces fair-share scheduling so heavy users don't permanently dominate shared infrastructure.

Armada is used in production at G-Research since 2020, processing millions of batch jobs per day across tens of thousands of nodes.


Features

Feature Description
🌐 Multi-cluster scheduling One API across unlimited Kubernetes clusters
⚖️ Fair-share queuing Dominant resource fairness across teams and queues
🔗 Gang scheduling Atomic startup for distributed workloads
⚡ Preemption Urgent jobs bump lower-priority work automatically
📊 Prometheus metrics Full observability into queue health and cluster utilisation
🔭 Lookout UI Web interface for monitoring jobs, queues, and clusters
🔒 Enterprise-ready Secure, highly available, OIDC authentication support

Getting started

The fastest way to get Armada running locally is with the Armada Operator:

git clone https://github.com/armadaproject/armada-operator.git
cd armada-operator
make kind-all

→ Full quickstart guide — get up and running in an instant!

armadactl

armadactl is installed automatically when you run make kind-all. To install it standalone or on a machine without the full Armada setup:

# download via script
scripts/get-armadactl.sh

# or grab the binary from the release page
https://github.com/armadaproject/armada/releases/latest

Local development

Armada runs locally via Goreman — dependencies (Redis, Postgres, Pulsar) run in containers, Armada components run as host processes built from source. Iteration is fast and debuggers attach directly.

mage kind                  # one-time: create local Kubernetes cluster
export KUBECONFIG=.kube/external/config

mage dev:up                # default — no auth
mage dev:up auth           # with OIDC via Keycloak
mage dev:up fake-executor  # no Kubernetes cluster needed
mage dev:down              # stop dependency containers

→ Full local development guide — profiles, procfiles, service ports, authentication, and debugging.


Use cases

Armada is used wherever batch jobs are too large, too many, or too complex for a single Kubernetes cluster:

  • Quantitative finance & HPC — millions of short-lived simulations per day with fair-share across research teams
  • ML and AI training — distributed GPU training with gang scheduling across clusters
  • Platform engineering — multi-tenant batch infrastructure with a single API surface
  • SLURM migration — familiar scheduling semantics (queues, priorities, preemption) on Kubernetes-native infrastructure
  • CI/CD at scale — priority control so critical merges always run first

In production

Armada has been running in production at G-Research since 2020.

Running Armada in production? Open a PR to add yourself to ADOPTERS.md 🙌


Community

Everyone is welcome — come and say hi! 👋


Contributing

We'd love your contributions — code, docs, bug reports, or ideas. All are welcome.


Documentation

Resource Link
Website & overview armadaproject.io
Quickstart armadaproject.io/getting-started
Architecture armadaproject.io/docs/architecture
API reference armadaproject.io/docs/api
Developer guide armadaproject.io/docs/developer-guide
Release notes github.com/armadaproject/armada/releases

Talks and videos



CNCF logo
Armada is a Cloud Native Computing Foundation Sandbox project 🚀
Apache 2.0 License

About

A multi-cluster batch queuing system for high-throughput workloads on Kubernetes.

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages