Google Releases LLM-Evalkit for Structured Prompt Engineering on Vertex AI

Oct 21, 2025
Google has introduced LLM-Evalkit, an open-source framework built on Vertex AI SDKs that centralizes prompt engineering workflows. The tool enables teams to measure performance systematically using objective metrics and a no-code interface.

Google has introduced LLM-Evalkit, an open-source framework designed to organize and measure prompt engineering for large language models, according to a Google Cloud blog post. Built on Vertex AI SDKs, the tool provides a unified environment where teams can create, test, version, and benchmark prompts using consistent evaluation metrics.

LLM-Evalkit consolidates previously scattered workflows by combining prompt creation, testing, and comparison into a single interface. It allows teams to define specific tasks, assemble representative datasets, and evaluate outputs against objective benchmarks. This standardized approach replaces guesswork with measurable performance data.

The framework also includes a no-code interface, making it accessible to non-technical users such as product managers and UX writers. By enabling collaboration across disciplines, it helps teams iterate on prompt design more efficiently.

LLM-Evalkit is available as an open-source project on GitHub and integrates directly with Google Cloud tools. New users can explore it using the $300 trial credit offered through Google Cloud.

Alphabet’s new framework aims to streamline prompt engineering by providing a structured, data-driven workflow within the Vertex AI ecosystem.

We hope you enjoyed this article

Consider subscribing to one of our newsletters like Daily AI Brief.

Also, consider following us on social media:

Free newsletter

Daily AI Brief

Daily report covering major AI developments and industry news, with both top stories and complete market updates

Industry analysis

2025 Global Business Services Agenda: Gen AI Takes Center Stage

The Hackett Group

This industry analysis by The Hackett Group explores the transformative impact of generative artificial intelligence (Gen AI) on global business services (GBS) in 2025. The study highlights the shift from exploration to acceleration of Gen AI initiatives, with 89% of executives advancing these projects to improve customer satisfaction, innovate products, and reduce costs. The report also discusses the challenges and strategies for successful Gen AI adoption, emphasizing the need for a technology-enabled operating model and the importance of reskilling the workforce.

Read more
Free, six days a week

Daily AI Brief: the AI news that matters, in your inbox.