Welcome to VaxBench

AI-Powered Model Benchmark Analytics

VaxBench reveals true model performance with transparent insights on latency, cost and reliability — empowering smarter decisions for builders and teams.

Benchmarking models from

322

Hard Problems

10

Categories

---

GLM 5.2 High Overall

100%

Objective Grading

Live Demo

See a benchmark run in real time

Hit start to simulate four frontier models racing through the same prompt.

vaxbench_engine.sh
SYSTEM ONLINE
Global Leader

GLM 5.2 High

---
Reasoning0%
Coding0%

Interactive Latency Race

Gemini 1.5 Pro0%
Claude 3.5 Sonnet0%
GPT-4o0%
Llama 3.1 405B0%
Hit “Start benchmark” to run the live GPU cluster simulation.
Why VaxBench

Transparent model benchmarking

We run models through execution-graded coding and exact-match reasoning tasks, then report the results with per-task detail.

See our methodology
01

Execution-Graded Coding

Generated code runs against hidden unit tests in a sandbox — pass@1 across algorithms, software engineering, frontend and security.

02

Exact-Match Reasoning

Reasoning answers are compared to reference answers. No partial credit, no model-as-a-judge.

03

Per-Task Transparency

Every run writes a timestamped JSON file with per-task detail, aggregated scores and performance telemetry.

New · FactoryBench

Can AI run a factory?

FactoryBench scores models as factory operators across 8 axes — vision QA, predictive maintenance, scheduling, energy, safety, profit, customer satisfaction and OEE. A fundamentally different benchmark from coding & reasoning.

Explore FactoryBench

8 Axes

Vision, maintenance, scheduling, energy, safety, profit, SAT, OEE

Agent Layer

Closed-loop factory control with 10 action types

Simulated Floor

CNC, conveyor, robot & inspection stations

Real KPIs

Production, scrap rate, downtime, OEE & profit

Ready to choose the right model?

Compare model scores across coding and reasoning tasks to find the right fit for your project.