What Veracode’s 45% AI Code Finding Actually Means

Yunhao Jiao
What Veracode’s 45% AI Code Finding Actually Means cover

In Veracode's 2025 GenAI Code Security Report, 45% of tested code-generation tasks introduced a known security flaw. The study used 80 tasks across four languages and selected weakness classes with more than 100 models; the result is not a measured defect rate for all AI-generated code. Source: https://www.veracode.com/wp-content/uploads/2025_GenAI_Code_Security_Report_Final.pdf.

The benchmark tests specific prompts and vulnerability classes. It does not measure the production codebases of professional developers or prove a universal security ceiling. Use the finding to prioritize independent security checks, not to predict your own defect rate.

Other studies report different rates and use different samples and definitions. Do not combine their ratios with Veracode's 45% result as though they came from one population. Test the risks relevant to your application, especially authorization, input handling, secrets, and session management.

Security failures can have real consequences, but survey results are self-reported and do not establish how often a given application will be affected. Record confirmed findings and remediation time in your own environment.

Why AI Generates Insecure Code

AI coding models can produce insecure implementations when the prompt omits application-specific controls or the output is accepted without review. The training-data explanation alone does not establish why a particular flaw occurred.

Specific patterns:

  • Input validation may be missing or inconsistent with the application's data rules
  • Authorization checks may not enforce your application's roles or object ownership
  • Secrets can be exposed in generated configuration or example code
  • Error responses can reveal internal details if not reviewed

The Fix: Automated Security Testing on Every PR

Clear requirements help, but prompts alone cannot verify security. Review the change, run relevant static and dependency checks, and test authorization and input boundaries on each high-risk change.

TestSprite can exercise configured UI and API flows, including selected negative cases. Dedicated security tests and human review are still needed for risks the suite does not cover. Inspect failures before merge.

Use the Veracode finding as a reason to verify generated code independently. Track the vulnerabilities your checks find and the ones that escape so coverage improves over time.

Try TestSprite free →