ARXIV:2605.10141 · FORMAL THEOREM PROVING AI · SUBMITTED 12 MAY · 20:15 UTC · FRESHNESS FRESH

VerifiedSource: PDF linkedVerifiedPaperPack: citation fields availablePartialProof: unverified proof status

FormalRewardBench: A Benchmark for Formal Theorem Proving Reward Models

Zeynel A. Uluşan · Burak S. Akbudak · Can S. Erer · Gözde Gül Şahin · arXiv

A benchmark for evaluating reward models in formal theorem proving, crucial for advancing AI assistants in complex mathematical reasoning.

Ship in 2-4 weeks›Score7.0Evidence unverified

Opportunity summary

Pain A benchmark for evaluating reward models in formal theorem proving, crucial for advancing AI assistants in complex mathematical reasoning.

Evidence 0 refs | 0 sources | 0% coverage

Blocker Evidence unverified

Open Build Read PDF Signal Canvas Track

PROBLEM

A benchmark for evaluating reward models in formal theorem proving, crucial for advancing AI assistants in complex mathematical reasoning. While verifiable rewards are cheap and scalable without reward hacking issues, they suffer from sparse…

METHOD

Full abstract

Recent neural theorem provers use reinforcement learning with verifiable rewards (RLVR), where proof assistants provide binary correctness signals. While verifiable rewards are cheap and scalable without reward hacking issues, they suffer from sparse credit assignment: models receive no learning signal from difficult problems where partial progress goes unrewarded. This motivates learned reward models that can evaluate proof quality beyond binary verification. However, comparing reward models is challenging since it typically requires expensive RL training ablations. To address this, we introduce \textbf{FormalRewardBench}, the first benchmark for evaluating reward models in formal theorem proving with Lean 4. Our benchmark consists of 250 preference pairs where correct proofs are paired with incorrect variants generated through five expert curated error injection strategies: forced mistakes, minimal single-point variations, verbose incorrect proofs, natural language justification, and Python code injection. We evaluate frontier LLMs (e.g., Claude Opus 4.5), judge LLMs (e.g., CompassJudger-1-14B), general-purpose LLMs (e.g., Qwen2.5-72B-Instruct), and specialized theorem proving models (e.g., DeepSeek-Prover-V2-7B). Our results reveal that frontier LLMs achieve the highest performance (59.8\%) while specialized theorem provers perform the worst (24.4\%), suggesting that theorem proving ability does not transfer to proof evaluation. We provide further insights on various error injection mechanisms, highlighting the challenging nature of most injection mechanisms. We release \textbf{FormalRewardBench} publicly to encourage more research on developing reward models in formal mathematics.

RESULT

ScienceToStartup currently rates this 7.0/10 on the public viability pass. Our results reveal that frontier LLMs achieve the highest performance (59.8\%) while specialized theorem provers perform the worst (24.4\%), suggesting that theorem proving ability…

WHY NOW

Formal Theorem Proving AI moved forward this cycle; last verified May 2026. Public score 7.0/10. Production flags indicate code availability.

Continue into Read for claims, analysis, references, and neighboring papers.

Opportunity summary

Score7.0

PainA benchmark for evaluating reward models in formal theorem proving, crucial for advancing AI assistants in complex mathematical reasoning.

Evidence0 refs | 0 sources | 0% coverage

Blockerno shell-level blocker reported

Analysis summary

A benchmark for evaluating reward models in formal theorem proving, crucial for advancing AI assistants in complex mathematical reasoning.

VerifiedSource: PDF linkedVerifiedPaperPack: citation fields availablePartialProof: unverified proof status

Competitive landscape

A benchmark for evaluating reward models in formal theorem proving, crucial for advancing AI assistants in complex mathematical reasoning.

Segment

Formal Theorem Proving AI

Adoption evidence

No public code link in the paper record yet

Commercial read

7.0/10 public viability

Direct

not classified

Adjacent

not classified

Substitute

not classified

Unknown

not classified

{ "contract_version": "paper-r2", "paper_id": "9c503d99-4525-441b-b78a-218f7968699b", "arxiv_id": "2605.10141", "canonical_route": "/paper/formalrewardbench-a-benchmark-for-formal-theorem-proving-reward-models", "active_tab": "synced from current hash by the drawer client", "selected_artifact": "formalrewardbench-a-benchmark-for-formal-theorem-proving-reward-models", "endpoints": { "paper_pack": "/api/v1/paper/formalrewardbench-a-benchmark-for-formal-theorem-proving-reward-models/paper-pack", "build_passport": "/api/v1/paper/formalrewardbench-a-benchmark-for-formal-theorem-proving-reward-models/build-passport", "mcp_resource": "sciencetostartup://surfaces/paper-workspace" } }

{ "surface": "paper", "mode": "paper", "query": "FormalRewardBench: A Benchmark for Formal Theorem Proving Reward Models", "normalized_query": "2605.10141", "route": "/paper/formalrewardbench-a-benchmark-for-formal-theorem-proving-reward-models", "paper_ref": "formalrewardbench-a-benchmark-for-formal-theorem-proving-reward-models", "topic_slug": null, "benchmark_ref": null, "dataset_ref": null }

{ "@context": "https://schema.org", "@graph": [ { "@type": "WebPage", "@id": "https://sciencetostartup.com/paper/formalrewardbench-a-benchmark-for-formal-theorem-proving-reward-models#webpage", "url": "https://sciencetostartup.com/paper/formalrewardbench-a-benchmark-for-formal-theorem-proving-reward-models", "name": "FormalRewardBench: A Benchmark for Formal Theorem Proving Reward Models", "description": "A benchmark for evaluating reward models in formal theorem proving, crucial for advancing AI assistants in complex mathematical reasoning.", "isPartOf": { "@id": "https://sciencetostartup.com/#website" } }, { "@type": "ScholarlyArticle", "@id": "https://sciencetostartup.com/paper/formalrewardbench-a-benchmark-for-formal-theorem-proving-reward-models#scholarlyArticle", "headline": "FormalRewardBench: A Benchmark for Formal Theorem Proving Reward Models", "description": "A benchmark for evaluating reward models in formal theorem proving, crucial for advancing AI assistants in complex mathematical reasoning.", "url": "https://sciencetostartup.com/paper/formalrewardbench-a-benchmark-for-formal-theorem-proving-reward-models", "sameAs": "https://arxiv.org/abs/2605.10141", "identifier": { "@type": "PropertyValue", "propertyID": "arXiv", "value": "2605.10141" }, "isAccessibleForFree": true, "isPartOf": { "@id": "https://sciencetostartup.com/#website" }, "datePublished": "2026-05-11T07:51:15.000Z", "author": [ { "@type": "Person", "name": "Zeynel A. Uluşan" }, { "@type": "Person", "name": "Burak S. Akbudak" }, { "@type": "Person", "name": "Can S. Erer" }, { "@type": "Person", "name": "Gözde Gül Şahin" } ], "additionalProperty": [ { "@type": "PropertyValue", "propertyID": "viabilityScore", "value": 7 }, { "@type": "PropertyValue", "propertyID": "researchDomain", "value": "Formal Theorem Proving AI" }, { "@type": "PropertyValue", "propertyID": "commercialReadiness", "value": "code" } ] }, { "@type": "BreadcrumbList", "itemListElement": [ { "@type": "ListItem", "position": 1, "name": "Home", "item": "https://sciencetostartup.com" }, { "@type": "ListItem", "position": 2, "name": "Formal Theorem Proving AI", "item": "https://sciencetostartup.com/topics" }, { "@type": "ListItem", "position": 3, "name": "FormalRewardBench: A Benchmark for Formal Theorem Proving Re", "item": "https://sciencetostartup.com/paper/formalrewardbench-a-benchmark-for-formal-theorem-proving-reward-models" } ] } ] }

Competitive landscape

A benchmark for evaluating reward models in formal theorem proving, crucial for advancing AI assistants in complex mathematical reasoning.

Segment

Formal Theorem Proving AI

Adoption evidence

No public code link in the paper record yet

Commercial read

7.0/10 public viability

Direct

not classified

Adjacent

not classified

Substitute

not classified

Unknown

not classified

FormalRewardBench: A Benchmark for Formal Theorem Proving Reward Models

FormalRewardBench: A Benchmark for Formal Theorem Proving Reward Models

Claim map

Constellation map

Competitive landscape

Buzz

PDF

REFERENCES

Related Papers

Subscribe to the weekly brief

Build artifacts

Brief

Experiment plan

Validation checklist

Scientific founder

Translational engineer

Domain operator

GTM lead

Regulatory/clinical advisor

Timeline

Claim map

Constellation map

Competitive landscape

Buzz

PDF

REFERENCES

Related Papers

Subscribe to the weekly brief

Build artifacts

Brief

Experiment plan

Validation checklist

Scientific founder

Translational engineer

Domain operator

GTM lead

Regulatory/clinical advisor

Timeline