Giskard Unveils Phare: A New Benchmark for Evaluating AI Models

Feb 19, 2025
Giskard has launched Phare, an open and independent benchmark to assess AI models on security dimensions like hallucination and bias, with Google DeepMind as a research partner.
Giskard Unveils Phare: A New Benchmark for Evaluating AI Models

Giskard has introduced Phare, a new open and independent benchmark designed to evaluate large language models (LLMs) on key security dimensions such as hallucination, factual accuracy, bias, and potential for harm. This announcement was made during the Paris AI Summit, with Google DeepMind collaborating as a research partner. The initiative aims to provide open measurements to assess the trustworthiness of generative AI models in real-world applications announced on their website.

Phare, which stands for "Potential Harm Assessment & Risk Evaluation," is designed to evaluate language models across multiple languages, initially including English, French, and Spanish. The benchmark will incorporate diverse linguistic and cultural contexts to ensure comprehensive assessments. The initial scope covers leading models from top AI labs such as OpenAI, Anthropic, Google DeepMind, Meta, Mistral, Alibaba, and DeepSeek.

The benchmark consists of modular test components focusing on four fundamental safety categories: hallucination, bias and fairness, intentional abuse by users, and harmful content generation. Giskard maintains full autonomy in determining the benchmark design, ensuring independence from model developers. The results from these assessments will be tracked on a public leaderboard, with future modules expanding to cover more languages and additional security aspects.

This collaborative effort is part of a broader initiative to improve AI security and robustness, encouraging practical developments in AI safety. Giskard plans to open-source a representative set of samples for each benchmarking module, enabling independent verification and private model testing.

We hope you enjoyed this article

Consider subscribing to one of our newsletters like AI Policy Brief or Daily AI Brief.

Also, consider following us on social media:

Free newsletter

AI Policy Brief

Weekly report on AI regulations, safety standards, government policies, and compliance requirements worldwide.

Industry analysis

2025 Global Business Services Agenda: Gen AI Takes Center Stage

The Hackett Group

This industry analysis by The Hackett Group explores the transformative impact of generative artificial intelligence (Gen AI) on global business services (GBS) in 2025. The study highlights the shift from exploration to acceleration of Gen AI initiatives, with 89% of executives advancing these projects to improve customer satisfaction, innovate products, and reduce costs. The report also discusses the challenges and strategies for successful Gen AI adoption, emphasizing the need for a technology-enabled operating model and the importance of reskilling the workforce.

Read more
Free, six days a week

Daily AI Brief: the AI news that matters, in your inbox.