Artificial intelligence may be accelerating business outcomes, but the cost of running AI at scale is beginning to show up in another place: resilience.
Fastly's findings show that AI-first businesses in Southeast Asia take around 80 days longer to recover from a security incident than their non-AI-first peers, representing a 135 percent gap, according to Fernando Medrano, Deputy CISO at Fastly.
Sharing his insights with iTNews Asia, Medrano said the issue is that organisations continue to view AI primarily through the lens of models and innovation while underestimating the infrastructure, security and recovery requirements.
What begins as a contained AI project can quickly spread into logins, content delivery, fraud controls and customer support. That changes the nature of an incident, he explained.
Teams may no longer be troubleshooting a single system. Instead, they have to trace the behaviour of multiple AI agents across services, regions and technology stacks, often with limited visibility. AI also creates unfamiliar traffic patterns and increases the number of privileged machine identities.
AI’s productivity dividend is being taken up by its operational load
AI is still delivering productivity gains, particularly within security and engineering teams. Medrano pointed to use cases such as alert triage, run book development, test coverage and analysing unfamiliar code as areas where AI can improve operational efficiency.
AI-assisted communications can also help organisations explain incidents more quickly and consistently to customers and regulators, helping maintain trust while systems are degraded.
But those gains can be offset by the additional operational burden created by AI itself, he added.
AI-driven endpoints can increase bandwidth and origin usage, while organisations also need to account for automated traffic and scraping. The result is a growing “AI speed tax”: productivity improvements that are increasingly being consumed by the infrastructure, security and specialist expertise required to operate AI safely, Medrano explained.
The AI bill is bigger than the model
For Medrano, one of the biggest problems is the way organisations calculate the cost of AI. Model selection, GPU efficiency, licensing and fine-tuning are relatively easy to quantify. Less visible are the costs associated with running AI reliably in production.
These include additional telemetry for AI agents, identity and access controls for machine identities, safeguards against automated traffic, observability, incident response and specialist expertise.
Medrano said the AI systems also need to be designed for failure. That means testing fallback behaviour, preparing for outages and ensuring teams can respond when AI becomes part of the failure chain.
“The ‘AI bill’ includes the cost of designing for failure, testing fall-back behaviours, and building incident response capabilities that assume AI will be involved in the next outage,” he explained.
For organisations that postpone those investments, the costs can emerge later as longer outages, higher infrastructure bills and more complex incident response.
The issues that delay recovery
Longer recovery times are not caused by AI architecture alone. Medrano sees three key issues such as technical complexity, skills shortages and governance gaps interacting.
A small code change, feature-flag error or data issue can produce unexpected traffic patterns, caching behaviour or responses across multiple regions. Incident teams also need new expertise spanning model behaviour, prompt risks, data lineage and the network paths used by AI features. Those skills are scarce, creating bottlenecks when an AI-related incident occurs. Governance can add further delays.
When an AI feature causes a problem, organisations need to know who can disable it, who owns the risk and who is responsible for explaining the impact to customers and regulators.
“If those questions aren’t answered before an incident, they get answered in real time on a crowded call,” Medrano said.
The challenge is also exposing assumptions built into traditional enterprise architectures. Many networks were designed around human-driven browsing and transactional workloads, where traffic was relatively predictable and non-human activity was the exception.
AI changes those patterns, he said.
Responses can become more personalised and dynamic, reducing cache efficiency and pushing additional traffic back to origin. AI agents can generate continuous background requests rather than conventional human sessions, while latency becomes more important when inference sits directly in the user path.
Machine identities are also multiplying as AI services gain access to applications, data and other systems.
For Medrano, the answer is not to discard traditional architectures but to revisit them through an AI-specific lens.

So the architecture question isn’t ‘is what we have fundamentally wrong?’ It’s ‘have we revisited our design with AI’s traffic and trust model in mind?’ If the answer is no, you’re effectively running new workloads on old models, and that is where fragility comes from.
- Fernando Medrano, Deputy CISO at Fastly.
Be careful in your AI design choice
Medrano highlighted several decisions that may create problems as deployments mature.The first is granting AI services broad, long-lived permissions because it makes experimentation easier.
The second is assuming every AI-generated response needs to be completely unique. Personalisation has value, but excessive uniqueness can undermine caching and edge offload. Using templates, semantic caching or shared components where appropriate can reduce origin load and limit the blast radius of failures.
Third, Medrano pointed to a distributing inference and state without investing in consistent observability and rapid rollback.
When should boards hit the brakes?
As AI moves into logins, payments, fraud checks, content safeguards and customer support, it should be tested like any other critical system. An AI failure may not appear as an “AI outage”.
Resilience testing therefore needs to cover the entire chain. Organisations should test what happens when models slow down, an edge location fails, agents generate unusual traffic or security controls throttle that traffic.
They also need to test the operational response like how quickly an AI feature can be removed from a critical path, what fallback exists and how the impact will be communicated. “If AI is in the critical delivery path, it deserves the same level of scrutiny as anything else you consider ‘too important to fail.’”
Medrano recommends boards should intervene when AI deployment begins to outpace organisational control.
● Lack of visibility: Leaders cannot clearly see how much AI is contributing to traffic, costs and risk.
● Recurring incidents: AI-related problems keep appearing, but organisations are only making small fixes instead of addressing deeper architectural issues.
● Limited human expertise: The same small group of engineers and security specialists is repeatedly handling AI deployments and major incidents.
When these warning signs emerge, Medrano said boards should consider slowing AI expansion and strengthening resilience before scaling further.
AI’s infrastructure crunch
Medrano predicts the next AI challenge could resemble the cloud cost shock experienced by enterprises a decade ago, but this time the pressure may extend beyond compute into network and edge economics.
AI workloads can create persistent connections, dynamic responses, heavier security inspection and more automated traffic. That can push bandwidth and origin costs higher than organisations anticipated during the pilot stage.
Capacity teams may also need to revisit infrastructure sizing because AI-driven endpoints can behave very differently from the workloads they replaced.
If adoption continues without visibility into these dynamics, Medrano expects the resulting crunch to emerge through unplanned spending, performance degradation and more frequent high-severity incidents.
The next phase will be about sustainable AI
According to Medrano, some AI projects could ultimately be abandoned not because they fail technically, but because their infrastructure economics do not work. A model may perform exactly as intended while the end-to-end system proves too expensive to operate across regions and users once bandwidth, caching, security, observability and recovery are included.
Projects that constantly bypass caches, require manual tuning from security and SRE teams or make incidents slower and more expensive may need to be redesigned or retired.
“Technical success doesn’t guarantee operational or financial sustainability,” he added.
For enterprises, the next phase of AI adoption will therefore be less about proving that AI works, and more about proving that it can operate, secure and recover at scale.
Medrano said, organisations that build AI-aware architectures now, with appropriate traffic policies, caching strategies, machine-identity controls, observability and rapid rollback, will be better positioned to capture AI's productivity gains without allowing its operational costs to overwhelm them.





