Logical Intelligence’s Aleph Leads Formal Verification Benchmarks

May 21, 2026
Logical Intelligence announced that its AI coding agent Aleph achieved top scores across four major formal reasoning benchmarks, marking a step toward verified code generation for mission critical software.

Logical Intelligence announced in a press release that its AI coding agent Aleph achieved leading results on four formal reasoning benchmarks, including PutnamBench, VeriSoftBench, LeanEval, and Verina. The company stated that Aleph’s performance demonstrates that formally verified code generation is now practical for critical infrastructure software.

Aleph solved 99.4 percent of the PutnamBench problem set, outperforming ByteDance’s Seed-Prover 1.5 at 86 percent and Hilbert at 69 percent. On VeriSoftBench, which measures real world software verification, Aleph reached 94 percent success, ahead of Harmonic’s Aristotle at 69 percent and Google Gemini-3 Pro at 65 percent. It also achieved top results on LeanEval and a perfect score on Verina, which was independently confirmed by benchmark authors.

The company said Aleph operates in environments requiring machine checkable proofs rather than probabilistic outputs. It is already being used in production verification workflows, including work involving the Ethereum Foundation’s ArkLib cryptographic libraries. Logical Intelligence plans to open a beta program for Aleph later this year.

Aleph automates formal verification and produces proofs that ensure critical logic functions correctly across all execution paths. The agent is intended for operators of infrastructure and safety sensitive systems who require verified code generation.

We hope you enjoyed this article

Consider subscribing to one of our newsletters like Daily AI Brief or AI Programming Weekly.

Also, consider following us on social media:

Free newsletter

AI Programming Weekly

Weekly news about AI tools for software engineers, AI enabled IDE's and much more.

Whitepaper

Tensordyne Napier: What If One Rack Could Do the Work of Nine?

Tensordyne

This Tensordyne whitepaper presents Napier, an inference-focused AI processor and rack-scale system based on the company’s TDN Math logarithmic number system. It examines infrastructure requirements for large mixture-of-experts and agentic models, compares major inference architecture approaches, and details the TDN AIP processor, TDN72 pod, TDN Link fabric, and Napier Ultra configuration. The paper reports simulation-based performance, cost, and accuracy-validation results, including Tensordyne’s projected comparison of one Napier rack with a nine-rack Nvidia Rubin plus Groq deployment; the chip is reported as taped out and in fabrication.

Read more
Free, six days a week

Daily AI Brief: the AI news that matters, in your inbox.