Check out TLAPS-Bench Leaderboard.
Our vision. To prove the correctness of any given critical protocols and systems using TLA+ Proof System (TLAPS)
TLAPS-Bench evaluates whether AI agents can formally prove (or disprove) the correctness of complex protocols and systems using TLAPS.
Each problem in TLAPS-Bench is a TLA+ specification (including the formal model and the invariants that specify correctness properties). AI agents are asked to prove that the formal model satisfies the invariants. We consider each task in TLAPS-Bench to formally prove one invariant of a given specification.
TLAS-Bench includes a sandboxed runtime for AI agents to faithfully prove the given specification without cheating. The runtime is equipped with extensive checks to prevent reward hacking. We have used the runtime to prove many specifications, such as 2PC, Paxos, TCP state machine, Byzantine Paxos, Byzantine broadcast, etc.
Historically, we had two types of problems: Proof-Completion and Proof-from-Scratch. We retired Proof-Completion as completing a well-structured proof is no longer a challenge for frontier AI. However, Proof-from-Scratch tasks, which AI has to invent the entire proof structure, are still nontrivial and take long-horizon efforts.
We currently focus on a few hard problems (due to token shortage).
| Problems | Type | # Spec | # Invariants |
|---|---|---|---|
| FLASH Cache coherence | Protocol | 1 | 15 |
| ZooKeeper protocol | Protocol | 1 | 9 |
| Cahill’s serializable snapshot isolation | Protocol | 1 | 1 |
| Ivy TLB | System | 1 | 2 |
| OpenAddressing | System | 1 | 5 |
| etcd Raft | System | 1 | 8 |
| HashiCorp Raft | System | 1 | 6 |
| ZooKeeper implementation | System | 1 | 9 |
| MongoDB distributed transactions | System | 1 | 1 |
| Total | 9 | 56 |
For more problems, check out the full problem set.
We have retired the following problems from TLAP-Bench, because they are well proved (or disproved) by frontier AI models and thus are no longer capable of measuring the frontier. The problems can all be found in the repository, but we no longer run them for our leaderboard.
If you need the TLA+ proofs of these problems, contact us and we can share them.
Proof checking can use substantial memory, especially for Isabelle-heavy tasks. A few of these tasks can use significantly more than 64 GB of RAM for one job.
We recommend the following hardware configurations
| Profile | vCPUs per job | RAM per job | Guidance |
|---|---|---|---|
| Recommended | 8–12 | 96 GB | Provides better memory headroom. |
| Lower-headroom | 8–12 | 64 GB | A starting point; some Isabelle-heavy tasks may require more. |
On a wimpy machine, start with --jobs 1. Increase the value after you monitor peak memory use.
git clone https://github.com/specula-org/tlaps-bench.git
cd tlaps-bench
export OPENAI_API_KEY=sk-... # This step is optional: Codex is the default backend if no OpenAI key is provided.
uv run tlaps-bench run --mode proof-from-scratch --filter Euclid/Euclid-Hyperbook/GCD.tla --jobs 1 # A small proof-from-scratch example
The above command builds a sandbox Docker image, with tlapm, SANY, and the proof checker bundled in and runs the task inside it
(a firewall allows only the LLM API hosts and the benchmarks are mounted read-only). Later runs reuse this image. Results are stored in results/<mode>/<backend>/<timestamp>/.
uv run tlaps-bench run --mode proof-from-scratch --jobs 4
How to set up an agent (--backend and --model) and its credentials, the full CLI reference, and native (--no-container) setup are described in our usage guide.
We are grateful to the generous support from
- TLA+ Foundation
- OpenAI
- Anthropic (AI for Science Program)
- Qingrong Chen
- MIT LICENSE
- Third-party benchmark sources are attributed in NOTICE