Before NowSecure fielded its 2026 Mobile App Risk Management (MARM) Survey, we ran an experiment most survey research skips. We asked three leading AI models, Claude Sonnet, ChatGPT and Gemini, to predict how senior mobile application security leaders would answer the survey.

Then we compared those predictions with the survey responses.

We weren’t asking AI to determine whether respondents were right or wrong, nor did we compare the models’ predictions with actual incident outcomes. We wanted to see where predictions based on the information available to these models aligned with how security leaders described their own organizations, and where they diverged.

The difference was striking. On questions about external business risk, such as the impact of a day of app downtime and the importance of mobile apps to the business, the models came relatively close to the survey responses. On questions about internal program maturity, investment and incident readiness, all three predicted substantially lower ratings than respondents reported.

Where Did AI Predictions and Survey Responses Align?

Every respondent in the survey rated mobile apps as very important or critical to their business, and 58% said a single day of downtime causes severe business damage.

The models anticipated responses to these external business-risk questions relatively well, earning a four-star AI Prediction Score, NowSecure’s measure of how closely model predictions tracked survey responses.

The business importance of mobile apps and the consequences of downtime are also widely documented, giving the models a substantial public record to draw from.

Where Were the Largest Gaps Between AI Predictions and Survey Responses?

The picture changed when respondents were asked to assess their own organizations.

Here are the five largest gaps between what the models predicted and what security leaders reported:

Across all five questions, Claude, ChatGPT and Gemini predicted lower ratings than respondents ultimately reported.

That doesn’t mean the AI models were right and respondents were wrong. The models predicted survey responses. They did not evaluate the respondents’ mobile security programs or predict which organizations would experience security incidents.

But the consistency of the gaps raises an interesting question: Why did three different AI models anticipate more conservative assessments of mobile security maturity, investment, monitoring and readiness than security leaders gave themselves?

Why the Gap Is Worth Examining

The AI experiment can’t tell us whether respondents assessed their programs accurately. Other findings in the 2026 Mobile App Risk Management Survey provide useful context.

Sixty-five percent of organizations rated their mobile app security programs as advanced or highly effective. Yet 65% of organizations reporting advanced programs also reported experiencing a mobile app security incident.

That doesn’t prove those organizations were overconfident. Mature security programs can still experience incidents, and incident rates alone don’t measure program effectiveness. But the combination of high self-reported maturity and continued incidents suggests confidence may not tell the whole story.

AI governance offers another example. The survey found widespread adoption of AI in mobile apps alongside gaps in organizations’ ability to see what that AI is doing. While 95% of respondents said their organizations deploy AI in mobile apps, 37% said they don’t monitor AI behavior in the apps they develop and deploy.

As we explored in AI Is Already in Your Mobile Apps. Most Security Programs Haven’t Caught Up, program strength and reported incident rates don’t always move together. A policy or maturity rating can tell you how an organization has structured its security program. It doesn’t, on its own, show what AI capabilities are doing inside a shipped app.

What AI Predictions Can and Can’t Tell Us

The AI Prediction Score gives us another way to examine the survey responses. It shows where expectations based on the information available to AI models aligned with how security leaders answered and where substantial gaps emerged.

It does not measure whether a security program is effective. And it doesn’t establish a relationship between AI predictions and actual security incidents.

The survey’s incident analysis is a separate finding. So is the case for evidence-based assessment. When we talk about evidence, we mean evaluating mobile apps through objective mobile app security testing rather than relying solely on perceptions of program maturity or readiness.

That distinction matters because program-level confidence can’t reveal everything happening inside an application. Runtime behavior, data flows and third-party components all require direct examination. NowSecure research analyzing approximately 105,000 mobile app assessments, for example, found that authenticated testing detected 78% more sensitive data exposure per scan than unauthenticated testing.

The gap between AI predictions and survey responses doesn’t tell us whose assessment was right. It gives security leaders another reason to ask how they validate confidence in their programs.

The Takeaway for Security and Risk Leaders

A governance policy, an approved budget and a high maturity rating provide useful information about how a security program operates. They don’t necessarily show what a mobile app does at runtime, what data it moves or which third-party services it depends on.

The gap between the AI predictions and survey responses doesn’t tell us whose assessment was “right.” It does give security leaders another reason to ask how they validate the confidence they have in their programs.

Self-assessment can describe the program. Objective mobile application security testing provides evidence about the apps themselves.

And the AI Prediction Score is only one part of what the research uncovered. The full survey examines AI adoption and governance, mobile security maturity, testing practices, third-party risk, incident experience and differences across finance, healthcare, high tech and retail.

Download the 2026 Mobile App Risk Management Survey to see the complete findings, industry benchmarks and strategic recommendations for closing the gap between policy and AI visibility.