Why AI Agents Stumbled in 2025 and A Potential Path Forward

A
Andrew D'AmbrosioCo-founder
·
Cover Image for Why AI Agents Stumbled in 2025 and A Potential Path Forward

As 2025 draws to a close, it’s pertinent to reflect on what was hyped to be “the year of AI agents”. Instead, it seems 2025 could be more accurately characterised as the year many teams realised the limitations of AI agents operating in the real world.

Companies spent millions discovering that getting agents reliably working in production takes a lot more effort than initially expected or perhaps budgeted for. Then in July researchers at MIT went viral for reporting that only ~5% of custom enterprise GenAI tools make it through to successful implementation.

There were obvious exceptions to this rule, none more evident than Anthropic’s Claude Opus 4.5 model and its use in their agentic coding tool Claude Code, which is broadly redefining the job description of software development. Claude has proven that AI agents can be supremely useful for long chain, complex tasks to produce work of meaningful value. At Voxworks we use Claude to build out product features at a pace that was unimaginable just 12 months ago. It's a glimpse of where the rest of the world is headed.

In the meantime, the fundamental question is why haven’t AI agents had much broader impact on other jobs and industries? Many are wondering what will it take to get a Claude for insert my job description.

The Reliability Cliff

The fundamental problem with AI agents is the stochastic nature of AI models (as opposed to being deterministic) and the way in which AI agents chain together multiple responses to complete a complex task.

Hallucinations and model errors are largely acceptable in a parallel process such as a maths test, where a 96% score is considered a pass and in some cases better than most humans could manage. But in a sequence, where each answer is dependent on a correct prior answer, even a small error rate results in massive deterioration in reliability.

The reason for this can be stated in simple mathematical terms: the probability of success of the entire process is equal to the success rate of a single step raised to the power of the number of sequential steps. From the prior example, if the LLM hallucinates only 4% of the time, after 10 iterations the probability of failure for the entire sequence is more than one in three. If you reduce the error rate to 1%, that same probability drops to one in ten.

We see this problem as particularly acute in Voice AI. Voxworks is building an AI calling platform that can theoretically automate the majority of any business’s low complexity or routine call volumes. The challenge is in modelling conversational turns that are dependent on prior turns. If the AI agent makes a mistake on any one turn, the entire conversation is often compromised.

On the flipside, if we could make minor improvements in the success rate of each individual step, the overall performance of the agentic system is exponentially improved. So then we should just use smarter models, right?

Well, the AI industry as a whole has been narrowly focused on improving intelligence as measured by “evals”, that is improving the response quality of a single step in a maths test type environment, by scaling these neural nets to unimaginable size. Then came the “thinking models” like GPT o3, which entered deep thought patterns with recursive LLM calls. Running these models to tap intelligence gains is impossible in voice AI because we can’t run them with low enough latency to sound natural in conversation.

But what if you could simply change the structure of the agentic system to improve the single step accuracy without requiring a super intelligent LLM?

The Error-Proof Agent

In a recent study Solving a Million-Step LLM Task with Zero Errors, researchers at Cognizant AI Lab in collaboration with UT Austin introduced an agentic system called MAKER that successfully completed the Towers of Hanoi experiment with over 1,000,000 sequential steps without a single mistake, a feat that the top frontier AI models typically fail at after less than a hundred steps.

The Towers of Hanoi experiment is a recursive logic puzzle where one wrong move invalidates the entire solution. The puzzle is also significant for another reason. Earlier this year researchers at Apple published a paper titled The Illusion of Thinking which, amidst Apple’s widely publicised failure to meaningfully adopt AI in their products, was surprisingly critical of the notion that AI exhibited true thinking behaviour.

Apple argued that while LLMs appear to reason, they are mostly just matching patterns. Part of the evidence was what they coined the Reliability Cliff. Using the Towers of Hanoi puzzle, Apple showed that as the number of disks increased, the model's probability of success drops to zero. Therefore you cannot trust an LLM to plan more than ~50 steps into the future because the probability of a hallucination compounds until failure is guaranteed.

In rebuttal, the Cognizant team proved the bug Apple had discovered was a limitation of single models, not of AI systems. By taking the exact same puzzle Apple used, Cognizant proved they could bypass the intelligence limit by redesigning the architecture, extending the reliability horizon from Apple's ~100 steps to 1,000,000+ steps.

The Secret Sauce: MAKER

The researchers at Cognizant developed a framework called MAKER (Maximal Agentic decomposition, K-threshold Error mitigation, and Red-flagging), a specific architecture based on three fundamentals:

  • Extreme Decomposition: Instead of asking one model to "solve the puzzle," they broke the task down into the smallest possible atomic units (micro-steps).
  • Multi-Agent Voting ("First-to-ahead-by-k"): For every single step, multiple small agents would propose a move. The system wouldn't proceed until a consensus was reached.
  • Red-Flagging: Agents could self-assess and red-flag outputs that looked structurally wrong or confused, discarding them before they could pollute the vote.

The MAKER framework and its resultant performance proves that it is possible to build highly reliable autonomous agents by wrapping AI models in a rigorous voting and error-checking architecture. By demonstrating a simple system with perfect accuracy on multi-chain tasks, the researchers have effectively converted a non-deterministic technology (LLMs) into a somewhat deterministic system.

The question then becomes how does this theory work in practice for more complex and messy real world environments, such as a customer conversation, where instead of 3 potential options as in the Towers of Hanoi there are infinite combinations of responses that could occur.

Bridging the Reliability Cliff in Voice

At Voxworks, we spent most of 2025 dealing with the exact same problems with reliability that Apple identified. Under the existing agentic frameworks available to us, we could never guarantee the voice agent wouldn’t go down an incorrect path or say something completely rogue. We realised it would never be a reliable, compliant and controllable system in large scale industrial or commercial applications.

So we decided to build our own agentic framework architecture specific to low-latency and high-reliability voice agents. In doing so, we independently reached many of the same conclusions from the MAKER paper.

Our voice agents follow similar paradigms highlighted by the researchers. For example for every turn of a conversation we might run the same LLM input prompt in parallel and take the most consistent answer to feed into the conversation. This eliminates outlier hallucinations to an acceptably small probability and also yields a small latency benefit.

The research also highlights you don’t need super intelligent models to do useful work. We know that AI model performance degrades as the context window is used up, and we had issues trying to get the LLM to make too many decisions in a single call. By breaking down tasks into as many sub-tasks, small AI models working together in parallel can often outperform the larger models, and do so with far lower cost and latency.

We implement decomposition in our agentic framework by breaking down every decision point into a separate LLM microservice, such as whether a user has decided to hang up, or whether the voice on the other end is a voicemail message.

None of this is necessarily groundbreaking in voice AI but we’ve taken it a step further to attain a higher reliability standard, and putting it all together into one unified system has shown material improvement in the robustness of our product.

Where to in 2026?

If there’s one lesson from applied-AI in 2025 it’s that AI Agent reliability at this point is mostly an architecture problem.

At Voxworks we’ve attacked it by running complex voting mechanisms and LLM decomposition under the hood in hopes that we can provide Australian businesses with a reliable tool to automate routine phone calls. We’re still working on this challenge and the system is by no means perfect, but the improvement to date has been compelling.

Other teams will start experimenting with their own agentic frameworks to deconstruct LLM workloads for their own specific job application.

The good news is that none of this requires more intelligent LLMs. The models are already smart enough. The differentiator is how you chain them together, provide or remove context, carry memory forward, combine tool-calling at the appropriate time, and manage multiple models in parallel to maximise the chance of the overall system achieving its goal.