Apodex Launches TRACES Benchmark for Scientific Discovery AI

Aug 19, 2026
Apodex introduced TRACES, a benchmark that tests AI systems on scientific discovery tasks where the answer may not be known in advance.

Apodex announced in a press release TRACES, a benchmark for evaluating AI systems on scientific discovery tasks where answers may not be known in advance. The benchmark uses executable environments instead of static datasets with answer keys.

TRACES gives an AI system a work environment that can include scientific literature, structured datasets, code execution, specialized tools, simulators, folding engines, experimental feedback, or other interfaces. The system must choose actions, interpret results, adjust its approach, and work toward a verifiable outcome.

The benchmark evaluates both final results and the process used to reach them. Apodex said TRACES measures six capabilities: tool use, error repair, alternative hypotheses, coherence across long work sequences, evidence grounding, and scope limits for conclusions.

Apodex said TRACES is open for participation. Teams can submit solver systems for evaluation, and researchers or organizations can propose scientific problems to be turned into executable benchmark environments.

We hope you enjoyed this article

Consider subscribing to one of our newsletters like Daily AI Brief.

Also, consider following us on social media:

Free newsletter

Daily AI Brief

Daily report covering major AI developments and industry news, with both top stories and complete market updates

Whitepaper

Tensordyne Napier: What If One Rack Could Do the Work of Nine?

Tensordyne

This Tensordyne whitepaper presents Napier, an inference-focused AI processor and rack-scale system based on the company’s TDN Math logarithmic number system. It examines infrastructure requirements for large mixture-of-experts and agentic models, compares major inference architecture approaches, and details the TDN AIP processor, TDN72 pod, TDN Link fabric, and Napier Ultra configuration. The paper reports simulation-based performance, cost, and accuracy-validation results, including Tensordyne’s projected comparison of one Napier rack with a nine-rack Nvidia Rubin plus Groq deployment; the chip is reported as taped out and in fabrication.

Read more
Free, six days a week

Daily AI Brief: the AI news that matters, in your inbox.