Samaya AI Publishes FrontierFinance Results for Investment AI Agents
Samaya AI announced in a press release that it has published evaluation results for FrontierFinance, a public benchmark for AI agents used in investment workflows. The benchmark covers tasks across idea screening, research, financial modeling, portfolio tracking, and catalyst monitoring.
FrontierFinance includes 220 queries and 11,543 expert crafted rubrics. It tests workflows including screening and discovery, company research, sector analysis, earnings and events, and coverage monitoring.
Samaya evaluated frontier models including Anthropic Claude Fable 5 and Claude Opus 4.8, OpenAI GPT-5.6 Sol, and Google Gemini 3.1 Pro. It also tested open source models including GLM 5.2 and DeepSeek V4 Pro, along with its own AI system.
Samaya's system scored 50.8% accuracy in low effort mode at four times lower inference cost than Claude Fable 5. In high effort mode, it scored 56% accuracy at two times lower cost than Claude Fable 5. Claude Fable 5 was the highest scoring frontier model at 49.2%, followed by GPT-5.6 Sol at 46.8% and Claude Opus 4.8 at 45%.
We hope you enjoyed this article
Consider subscribing to one of our newsletters like Finance AI Weekly or Daily AI Brief.
Also, consider following us on social media:
More from Finance
Oct 1 bolttech and Bold Penguin Partner on Insurance Distribution Sep 30 Health In Tech launches HitRix and cuts 2026 revenue outlook Sep 30 Kin Lets Meta's Muse Shop for Home Insurance Quotes Sep 30 Millennial Shift Technologies Adds PebbleRisk Assessments to mShift Marketplace Sep 30 Flatworld Mortgage Relaunches Brand Around Agentic AI PlatformFinance AI Weekly
Weekly newsletter about AI in finance. Covers AI-driven trading, fintech innovations, and data analytics transforming markets
Whitepaper
Tensordyne Napier: What If One Rack Could Do the Work of Nine?
Tensordyne
This Tensordyne whitepaper presents Napier, an inference-focused AI processor and rack-scale system based on the company’s TDN Math logarithmic number system. It examines infrastructure requirements for large mixture-of-experts and agentic models, compares major inference architecture approaches, and details the TDN AIP processor, TDN72 pod, TDN Link fabric, and Napier Ultra configuration. The paper reports simulation-based performance, cost, and accuracy-validation results, including Tensordyne’s projected comparison of one Napier rack with a nine-rack Nvidia Rubin plus Groq deployment; the chip is reported as taped out and in fabrication.
Read moreYou may also like
Artmarket.com Tests How Five AI Systems Interpret the Same Book
Anthropic Attributes Its Largest Measured Distillation Campaign to Alibaba
Claude Leads 26% of Anthropic Model Research Work
Faye Expands AI Consulting Across Claude, Fin and ElevenLabs
62% of Finance Workers Say AI Errors Reached Clients or Decision Makers
Daily AI Brief: the AI news that matters, in your inbox.