Key Takeaways
- More than half of global businesses regularly ship untested code into production.
- There's a dangerous confidence gpa: 93% of C-suite leaders feel good about their testing strategies, while 40% of QA/DevOps people aren't confident at all.
- Multiple industry reports agree AI-generated code is accelerating production but also driving up incidents, bugs and rework.
As AI-written code becomes more common in enterprises, something is becoming clear: It’s not very good.
A majority of global businesses — 60% — admit they regularly ship untested code into production, according to a 2026 Tricentis report. And the results of that poor-quality software are:
- Security breaches & compliance failures (30%)
- Technical debt & rework costs (28%)
- Loss of customer or partner trust (25%)
Additionally, 1 in 5 companies reported that poor software quality costs them up to $5 million each year.
Companies Aren't Accidentally Shipping Bad Code Anymore
“[Companies are] getting pushed to move so fast they’re almost intentionally knowing that when they’re developing agentic code, they’re skipping steps along the way..."
- Ryan Yackel
SVP of Strategy, Tricentis
The number of companies shipping untested code didn't change much compared to the previous year. What did change is why, according to Ryan Yackel, senior vice president of strategy at Tricentis.
Last year, a similar report found 40% of issues involved with shipping untested code were accidental slips. Now, enterprises are doing it on purpose.
“They’re getting pushed to move so fast they’re almost intentionally knowing that when they’re developing agentic code, they’re skipping steps along the way, assuming things aren’t as bad, and they can get away with random bugs in production,” Yackel said. “People are saying, we’re being pushed to produce. We’re going to be okay with some things being pushed into production that we can fix.”
This year, respondents said software development teams knowingly pushed untested code to production due to:
- Top-down leadership pressure to accelerate delivery (32%)
- Having too much code to test (30%)
Executives Think Testing Is Fine. QA Teams Disagree.
A bigger problem is that the C-suite and the practitioners don’t see the problem the same way.
Of C-level respondents, 93% said they felt confident their testing strategies address the most critical risk areas, and 81% of CEOs have high trust in their AI-driven systems, data and automation tools.
Of QA and DevOps leaders, on the other hand, 30% said they're uncertain or explicitly unconfident in their testing strategies' effectiveness, and only 56% have high-trust in their AI-backed systems.
Research was conducted among a sample of 2,501 senior IT decision-makers, investors and QA/DevOps professionals from organizations with more than 150 employees in the US, UK, Ireland, Japan and Singapore, according to Censuswide, which conducted the survey. Data was collected in the latter half of April 2026.
AI Code Looks Great in Review — Then Falls Apart in Production
Tricentis isn’t alone. A number of other surveys, including those from New Relic, Lightrun, Faros and Google, report similar conclusions.
Almost all of those surveyed in the New Relic report (94%) rated AI-generated code higher quality than human-authored code at review. Once it ships, however:
- 78% report more incidents
- 86% report more senior engineer firefighting
- 74% say at least a quarter of AI-generated code needs significant rework
Two-thirds of organizations say they need to manually debug 26-50% of AI-generated code changes, noted the Lightrun report, but, notably, no organization can release AI code fixes with just one redeploy cycle. More than three-quarters of engineering leaders say they require 2-3 redeploy cycles, and 11% said they need 4-6 cycles — just to verify a single change.
While the acceleration of software creation is undeniable, AI delivery has not yet positively impacted production reliability processes.
“Despite the widespread adoption and perceived benefits, some software development professionals remain cautious about using AI in their work,” the Google report noted, adding that a surprising "trust paradox" was uncovered: while 24% of respondents said they had a great deal or a lot of trust in AI, 30% said they only trust it a little or not at all.
What it means? AI outputs are perceived as useful and valuable by many, despite a lack of complete trust in them.
A Caveat: Many AI Quality Reports Come From Companies Selling the Fix
It should be pointed out, though, that while these surveys are generally performed by consulting firms, many were sponsored by vendors that make testing software for AI code. It’s a trend that bothers AI consultant and Linthicum Research founder David Linthicum.
“The enterprise AI market has a credibility problem with surveys,” Linthicum said. “Too many so-called ‘independent’ reports are really vendor-funded narratives with data attached. The pattern is obvious: Cybersecurity vendors find rising AI risk, integration vendors discover data is the blocker and cloud platforms warn that cloud complexity is spinning out of control. The conclusion often seems built in before the first question is asked. That doesn’t make every report useless, but it does make many of them compromised by design.”
That’s not to say the conclusions are incorrect, but it’s important to know where they’re coming from, Linthicum said.
For example, it’s interesting to note that, in contrast, the report from Google had a considerably sunnier outlook about AI-based development than the ones from the testing software vendors, and several of those vendor reports specifically compared their results to Google’s.
“What earns trust is straightforward: transparent methodology, repeatable analysis, independent funding and findings that don’t map perfectly to the sponsor’s sales deck,” said Linthicum. “What destroys trust is ‘commissioned by’ buried in fine print, narrow respondent pools and headlines that read like category marketing. Vendors should understand the long-term cost here: Once buyers assume the research is self-serving, even the good data gets ignored."
QA Teams Need a Seat at the Table — And a New Vocabulary
All that stipulated, what should enterprises do to prevent or solve the problem?
QA needs to work with engineering so its processes are inserted into the development process and QA isn’t seen as a bottleneck, Yackel advised. “The C-suite cares about speed, commodity and outpacing the competition,” he said. “If QA teams are still doing traditional test automation, where they wait until we get passed into the staging environment, the development team is going to say we’re going too slow. It needs to tie into agentic development.”
QA teams also need to speak in ways the C-suite and boardrooms understand, Yackel added. “Too often, software quality becomes a boardroom priority only after a major outage, customer disruption or reputational event forces the issue,” he said.
While practitioners often see quality issues first, the way those issues are communicated can determine whether leadership acts on them, Yackel said. “Bug volume, defect rates and test coverage all matter, but they often fail to resonate with executive leadership unless they are connected to business outcomes like production risk, customer disruption, revenue exposure, compliance risk or brand trust. The board doesn’t need to know every technical detail. It needs to understand what could happen if those signals are ignored.”
Editor's Note: We've been tracking AI's growing role in the development pipeline. Catch up on our latest coverage.
- Cursor Launches AI Model Router That Cuts Developer Costs by Up to 50% — Cursor launches an intelligent model routing system that automatically selects the best AI model per task.
- How to Build AI Agent Workflows in n8n: A Step-by-Step Guide — Learn how to build, run, and monitor AI agent workflows in n8n.
- Vibe Coding Goes Enterprise: Replit's $9B Moment — Shaquille O'Neal, a16z and Qatar's sovereign wealth fund all just backed the same AI coding platform. Here's why.