The Underlying Intent
← Blog

Leadership · Business · For Leaders & Professionals

Build AI Around Its Limits to Create Reliable Business Value

Ram Konduru · July 22, 2026

Build AI Around Its Limits to Create Reliable Business Value

Enterprise AI creates lasting value when leaders turn model limits into clear requirements for data, workflow design, verification, governance, human decision-making, and the full operating system around each model.

Two Days in July Settled the Enterprise AI Question

Two days in July settled a question enterprise buyers had been circling for a year: what are you actually paying for when you pay for AI?

On July 1, 2026, Palantir CEO Alex Karp went on CNBC and said out loud what many buyers had only muttered. Companies were pouring money into AI tokens, he argued, with little to show for it in business terms, and worse, they were handing their proprietary data, workflows, and internal knowledge to outside model providers in the process. He called the industry’s fixation on usage “tokenmaxxing.”

Take his framing with some distance. Palantir sells a competing platform, so the critique is also a sales pitch. But the question underneath it is real: beyond access to a powerful model, what is the buyer getting?

Microsoft answered the next day. On July 2 it launched Microsoft Frontier Company, backed by a $2.5 billion investment and roughly 6,000 industry and engineering experts who would sit inside customer organizations to design, deploy, govern, and improve AI systems tied to measurable outcomes. Model access, in other words, now comes bundled with engineering, trusted data, governance, and process redesign.

Put the two days together and the real product comes into focus. A model supplies capability. The operating system built around it is what turns that capability into dependable work, and whether you pay off on AI depends far more on the second thing than the first.

Capabilities vs. Boundaries: Start With the Failure Conditions

Most executive conversations about AI start with capability: what can the model do? That question is useful for comparing models and scoping pilots. It tells you almost nothing about what happens when the model hits an incomplete record, a conflicting instruction, an edge case, or a decision with legal and financial weight.

The more useful question is where the model stops being dependable.

Take loan underwriting. In a demo, the model reads an application, summarizes the supporting documents, and flags risks in seconds. Impressive, and beside the point. Real files are messier. One document reports a different income figure than another. A borrower has a legitimate situation that falls outside the usual pattern. A record is stale, incomplete, or simply entered wrong.

Something has to decide which record wins, what evidence the model is allowed to use, when the case goes to a trained human, and how the reasoning gets logged for later review. That “something” is the operating system around the model, and its requirements come straight from the model’s limits. They shape the data structure, workflow, approval rules, and audit trail before launch, not after.

Engineers already work this way. They study how a component fails under load, then build the surrounding structure to carry that load when the component reaches its limit. Enterprise AI deserves the same treatment. Reliability starts when the rollout plan treats failure conditions as a design input, not a surprise.

The Closed-Loop Requirement: Build Verification Into the Workflow

Language models generate likely responses from patterns in their training and the context they’re given. They can critique their own draft, but that critique is just another model output. It is not independent proof that the answer matches your records or your rules. A model cannot verify itself. The workflow around it has to.

That means an external loop with four jobs. Grounding: the model draws from an approved, current source, not an unspecified mix of files and memory, which in underwriting is the verified application data, policy rules, and current risk criteria. Verification: the system checks the output against that source, so a generated summary has to match the figures in the application and a recommended action has to satisfy a policy rule. Control: known conditions get handled by rule, where a missing document blocks approval and a request above a set value forces a second review. Escalation: low confidence, conflicting data, or an unusual case goes to a person, and that person gets enough context to see why it escalated and what the model already did.

Microsoft’s Frontier Company announcement describes the same shape, calling for platforms that let customers observe, govern, secure, and manage AI across the stack, plus a continuous loop that improves the process over time.

The loop is where the value is. It connects model output to trusted evidence and a controlled action. Skip it, and your staff spend their days checking, correcting, and explaining results, and the model has quietly created more work than it removed.

Constraint-First Strategy: Redesign the Process Before Scaling It

A constraint-first strategy maps the failure modes before deciding how widely to use AI. It feels slower than shipping a quick pilot. It buys you a real view of cost, risk, and the process changes the rollout will actually require.

Erik Brynjolfsson, Daniel Rock, and Chad Syverson named this pattern in their NBER paper on the modern productivity paradox. General-purpose technologies tend to need new skills, processes, and organizational change before their larger gains show up. The technology arrives first; the measurable productivity arrives later, once the supporting structure exists.

The enterprise data reads the same way. A 2026 study of S&P 500 firms found 11 percent had deeply integrated AI into business processes in 2025, with another 10 percent using it in producing goods or delivering services, up from 5 percent in 2022. The researchers found a J-curve relationship with profitability but no significant relationship with productivity or capital spending, and they were explicit that this is association, not causation.

The practical lesson: buying AI does not finish the job. Roles, data flows, approval steps, quality checks, and performance measures usually have to change before results show up across the business.

And local wins can hide system-level strain. A coding team ships more pull requests with an AI assistant while an unchanged review process buckles under them. An underwriting model scores applications faster while every unclear exception lands on compliance and customer support. The first step improves; the whole workflow carries more weight. Constraint mapping surfaces that early, by asking where errors travel, who absorbs the extra review, and whether the final business outcome improves or just one isolated step.

Three Objections Leaders Should Take Seriously

Closed-Loop Systems Can Cost More Than the Task Is Worth

Some AI tasks carry little risk and come with a built-in reviewer. A developer using an AI coding tool inspects the output, runs tests, and rejects bad suggestions before anything reaches production. A 2026 study of tens of thousands of Microsoft engineers found that Claude Code and GitHub Copilot CLI users merged about 24 percent more pull requests over four months. The researchers used merged PRs as a measure of output and said plainly that output is not the same as business value.

That is the right instinct: match the controls to the task. A private coding assistant in the hands of an experienced developer can run on a lighter process, because the user already supplies the review and the accountability. Automated underwriting, supply-chain routing, and payment approval involve more people and bigger consequences, so they need stronger data checks, decision rules, audit records, and escalation paths. Spend on control where an error can move fast, cross teams, or reach a customer. That is proportional design, not bureaucracy.

Large AI Providers Can Survive a Slow Return Cycle

A second objection points to the balance sheets of the major cloud providers, which can keep investing even when enterprise returns take years to arrive. True, and not the real risk. David Cahn’s 2024 Sequoia analysis put a roughly $600 billion gap between the revenue implied by AI infrastructure spending and the revenue actually visible across the market. Cahn also argued AI could create enormous long-term value; his worry was timing, investment levels, and who eats the losses while adoption catches up.

The nearer-term pressure comes from the buyer, not the hyperscaler. CFOs will keep asking whether rising token spend produces lower costs, more revenue, faster service, or better decisions. When teams cannot draw that line, spending slows even while the providers stay financially secure. A customer-led slowdown would not signal a cash crisis at the top. It would signal that businesses had hit the limit of paying for activity with no operating result attached.

Vendors Are Already Building the Missing System

Microsoft Frontier Company and Palantir’s forward-deployed model suggest the big vendors already grasp the integration problem and can place engineers inside customer organizations to build the systems directly. For some buyers, that closes the gap. It also proves the point about how much work sits beyond the model. Microsoft’s own announcement lists industry knowledge, change management, engineering, governance, model choice, data protection, cost control, and continuous improvement, every item a piece of the surrounding operating system.

The decisions still belong to the buyer, even when a vendor does most of the building. You set the acceptable risk, approve the data access, assign decision rights, choose the success measures, and own accountability when the system fails. A vendor can build the machinery. Only the company can decide how it fits the business.

Treat AI Limits as Design Inputs

Start with a short, blunt set of questions. Which source should the system trust? What output needs verification? Which actions can AI take without approval? What conditions should stop the workflow? Who reviews the unusual cases? How do we log decisions and learn from errors?

The answers are the operating system around the model. They tie together data, policy, software, roles, review steps, and business goals, and they set a higher bar than parking a human at the end of every AI task. Oversight matters, especially where a decision needs judgment or accountability. But a person cannot patch weak data, unclear rules, missing logs, and bad workflow design at scale. Reliability has to be distributed across the whole process.

AI pays off when the model’s limits shape the design from the start. Do that and teams pick better use cases, put controls where they matter, set honest expectations, and measure the result that actually reaches the customer or the balance sheet. The model is one important component. The system around it decides whether its output becomes reliable business value.

Source List