KushoAI announces APIEval-20 to benchmark AI agents with API testing

Machine Learning


  • Authentication failures cause 34% of API outages, new benchmark aims to address gap

  • APIEval-20 Built for APIs that change faster than documented

  • Recorded over 100 downloads in the first week of release

KushoAI, an AI-native platform for API testing and software reliability, introduced APIEval-20, an open benchmark designed to evaluate how effectively AI agents can identify functional bugs in APIs using only request schemas and sample payloads, without access to source code or documentation.

This change comes as concerns about API reliability continue to rise. Analysis of more than 1.4 million AI-driven test runs across 2,616 organizations shows that authentication failures alone are responsible for 34% of API outages, and 41% of APIs experience undocumented schema changes within a month. Nevertheless, most existing evaluation methods fail to capture whether AI tools are able to systematically detect such issues.

Rather than recreating an ideal testing environment, APIEval-20 intentionally introduces constraints that reflect real-world conditions, incomplete context, evolving schemas, and hidden dependencies to encourage AI agents to behave more like human QA engineers than automated validators.

Abhishek Saikia, co-founder and CEO of KushoAI, said, “The conversation around AI in testing has largely been about automation. What’s missing is accountability, a way to measure whether these systems actually work. APIEval-20 brings that accountability into the equation.”

Also read: AiThority Interview with Glenn Jocher, Ultralytics Founder and CEO

Saikia added, “The conversation we were expecting was around the benchmark itself. What we actually heard from engineers in week one was that they had been thinking about this question for months, but there was no way to answer it. For us, that validation is more important than the number of downloads.”

The benchmark includes 20 scenarios across domains such as payments, authentication, e-commerce, scheduling, user management, notifications, and search. Each environment has between 3 and 8 bugs, ranging from simple validation issues to serious logic flaws that require multi-step analysis.

measure what really matters

APIEval-20 introduces a scoring model that aligns with real-world priorities.

  • Bug detection (70%) to get practical effect
  • Coverage (20%) To assess the scope of the test
  • Efficiency (10%) Evaluate resource usage

Also read: The infrastructure war behind the AI ​​boom

[To share your insights with us, please write to psen@itechseries.com ]



Source link