Claude leads coding!
New models weekly!
VaxBench reveals true model performance with transparent insights on latency, cost and reliability — empowering smarter decisions for builders and teams.
Benchmarking models from
322
Hard Problems
10
Categories
---
GLM 5.2 High Overall
100%
Objective Grading
Hit start to simulate four frontier models racing through the same prompt.
We run models through execution-graded coding and exact-match reasoning tasks, then report the results with per-task detail.
See our methodologyGenerated code runs against hidden unit tests in a sandbox — pass@1 across algorithms, software engineering, frontend and security.
Reasoning answers are compared to reference answers. No partial credit, no model-as-a-judge.
Every run writes a timestamped JSON file with per-task detail, aggregated scores and performance telemetry.
FactoryBench scores models as factory operators across 8 axes — vision QA, predictive maintenance, scheduling, energy, safety, profit, customer satisfaction and OEE. A fundamentally different benchmark from coding & reasoning.
Explore FactoryBenchVision, maintenance, scheduling, energy, safety, profit, SAT, OEE
Closed-loop factory control with 10 action types
CNC, conveyor, robot & inspection stations
Production, scrap rate, downtime, OEE & profit
Compare model scores across coding and reasoning tasks to find the right fit for your project.