Logical Intelligence’s Aleph Leads Formal Verification Benchmarks
Logical Intelligence announced in a press release that its AI coding agent Aleph achieved leading results on four formal reasoning benchmarks, including PutnamBench, VeriSoftBench, LeanEval, and Verina. The company stated that Aleph’s performance demonstrates that formally verified code generation is now practical for critical infrastructure software.
Aleph solved 99.4 percent of the PutnamBench problem set, outperforming ByteDance’s Seed-Prover 1.5 at 86 percent and Hilbert at 69 percent. On VeriSoftBench, which measures real world software verification, Aleph reached 94 percent success, ahead of Harmonic’s Aristotle at 69 percent and Google Gemini-3 Pro at 65 percent. It also achieved top results on LeanEval and a perfect score on Verina, which was independently confirmed by benchmark authors.
The company said Aleph operates in environments requiring machine checkable proofs rather than probabilistic outputs. It is already being used in production verification workflows, including work involving the Ethereum Foundation’s ArkLib cryptographic libraries. Logical Intelligence plans to open a beta program for Aleph later this year.
Aleph automates formal verification and produces proofs that ensure critical logic functions correctly across all execution paths. The agent is intended for operators of infrastructure and safety sensitive systems who require verified code generation.
We hope you enjoyed this article
Consider subscribing to one of our newsletters like Daily AI Brief or AI Programming Weekly.
Also, consider following us on social media:
AI Programming Weekly
Weekly news about AI tools for software engineers, AI enabled IDE's and much more.
Whitepaper
Tensordyne Napier: What If One Rack Could Do the Work of Nine?
Tensordyne
This Tensordyne whitepaper presents Napier, an inference-focused AI processor and rack-scale system based on the company’s TDN Math logarithmic number system. It examines infrastructure requirements for large mixture-of-experts and agentic models, compares major inference architecture approaches, and details the TDN AIP processor, TDN72 pod, TDN Link fabric, and Napier Ultra configuration. The paper reports simulation-based performance, cost, and accuracy-validation results, including Tensordyne’s projected comparison of one Napier rack with a nine-rack Nvidia Rubin plus Groq deployment; the chip is reported as taped out and in fabrication.
Read moreYou may also like
Corbenic AI launches Galahad beta for secure AI memory
Thread AI Finds 5% of Enterprise AI Buyers Commit to Ship Dates
Google Introduces Gemini 4 Argon for Coding and Cybersecurity
Acalvio Launches ShadowPlex Security Agent for Gemini Enterprise
Cohere and Aleph Alpha Sign Definitive Combination Agreement
Daily AI Brief: the AI news that matters, in your inbox.