Why General AI Tools Like ChatGPT Are Not Enough for Enterprise AP Automation
MIT found that 95% of enterprise AI pilots deliver zero ROI and named memory, grounding, and workflow integration as the structural gaps. Here's why general AI tools fail at AP, and what Blackbee AI does instead.
There's a story that plays out in almost every finance team that discovers AI.
Someone on the AP team starts using ChatGPT to extract invoice data. It works. It's fast. They build a few prompts. Then someone else uses it to analyse spend, and someone else drafts vendor emails with it. Within a few months, the team has a genuinely useful set of AI-assisted workflows and a growing sense that this might be the answer to the AP automation question nobody wanted to spend six figures on.
Then the volume climbs. An auditor asks a question. A duplicate payment slips through. And the team discovers something uncomfortable: the thing that made ChatGPT so useful at the start, its flexibility, its willingness to reason about anything you paste into it, is exactly what makes it insufficient as the backbone of an enterprise AP operation.
This isn't a criticism of ChatGPT, Claude, or Gemini. These are extraordinary tools. But there's a hard structural difference between a general-purpose AI assistant and a system built to govern financial decisions at enterprise scale, and the data on what happens when organisations confuse the two is now genuinely sobering.
This piece explains why. Not with hand-waving, but with what the research actually shows about where general AI tools break down in enterprise workflows, why finance is the worst possible domain for those particular failure modes, and what a purpose-built alternative like Blackbee AI does differently.
The Data: Enterprise AI Is Failing, and We Know Why
Start with the number that should be pinned above every finance leader's desk.
MIT's Project NANDA published The GenAI Divide: State of AI in Business 2025, based on 52 executive interviews, surveys of 153 leaders, and an analysis of 300 public AI deployments. Its central finding: 95% of enterprise generative AI pilots delivered no measurable P&L impact. Only 5% of integrated systems created significant value.
Ninety-five percent. Not a low return. Zero.
And the abandonment rate is accelerating. Analysis of the same research notes that 42% of companies scrapped most AI initiatives in 2025, up from 17% in 2024, a rate that more than doubled in twelve months.
Now, here's the part that matters for this conversation. The MIT researchers were explicit about the cause, and it wasn't model quality. As Fortune's reporting on the study put it, the core issue was not the quality of the AI models, but the "learning gap" for both tools and organizations. While executives often blame regulation or model performance, MIT's research points to flawed enterprise integration. Generic tools like ChatGPT excel for individuals because of their flexibility, but they stall in enterprise use since they don't learn from or adapt to workflows.
The report itself is even more direct. Coverage of the study quotes it, stating: ChatGPT's very limitations reveal the core issue behind the GenAI Divide: it forgets context, doesn't learn, and can't evolve. For mission-critical work, 90% of users prefer humans. The gap is structural; GenAI lacks memory and adaptability.
Read that once more. The gap is structural. Not a version problem. Not something that gets fixed when the next model ships. A structural property of what a general-purpose assistant is.
There's one more finding in the MIT data that finance leaders should sit with. The report found a mismatch of use case and value: many organizations allocate more than half of their GenAI budgets to sales and marketing tools, while the highest ROI lies in back-office automation, document processing, compliance, and internal workflows.
The highest ROI is precisely where accounts payable lives. Which makes it all the more important to get the how right, because the opportunity is real, and the general-purpose approach to capturing it is demonstrably failing.
Why Finance Is the Worst Domain for General AI's Weaknesses
Every general-purpose LLM shares a set of characteristics. In most domains, those characteristics are acceptable trade-offs. In accounts payable, they're liabilities, and the academic research on this is unambiguous.
The arithmetic problem. Large language models are, at their core, prediction engines. As one analysis puts it bluntly, LLMs are prediction engines, not knowledge bases. They generate the most statistically plausible next word, not the most factually accurate one.
That property has a specific, well-documented consequence for numbers. Peer-reviewed research published as the FAITH benchmark, built from S&P 500 annual reports, found that proprietary frontier models achieve high overall accuracy, yet even these systems exhibit 10–20% error rates on multi-step numerical reasoning. The same body of work found that top models collapse from 95.6% accuracy on simple lookups to near 0% on multivariate calculations.
An invoice is a multivariate calculation. Quantity times unit price, summed across line items, plus tax, minus a volume discount, against a contracted rate cap. This is precisely the class of chained numerical operations where the research shows accuracy degrades fastest.
The FAITH authors are direct about the implication: even modest hallucination rates, such as 10–20% in complex financial reasoning, could translate into substantial financial misjudgments when scaled to production systems. Related research goes further, with one group arguing that LLMs must be "surgically" relieved of arithmetic responsibility altogether, in favour of neuro-symbolic approaches with deterministic fact ledgers.
Think about what a 10-20% error rate means at AP volumes. Another analysis puts the scale problem starkly: a 15% hallucination rate on 10,000 daily queries is 1,500 wrong answers per day entering business decisions. For an AP team processing a few hundred invoices a month, the arithmetic is less dramatic but no less consequential, because in AP, a wrong answer is a wrong payment.
The grounding problem. The second structural issue is that general AI tools reason from whatever you happen to hand them, not from a governed source of truth.
The evidence here is striking. Gartner research cited in the same analysis compared identical models running identical architectures against different data conditions: on ungoverned data, 52% of responses contained fabricated information; on governed data, using the same model, that rate collapsed. IBM's research reinforces the point, 72% of AI failures in enterprise settings are attributable to inadequate context, not model capability. The failure is upstream, not in the model weights.
This is the whole ballgame for AP. When you paste an invoice into ChatGPT and ask whether the pricing is right, you are asking a model to reason without access to the contract that defines what "right" means. It has no vendor master. No payment history. No purchase order. No contract terms. It has an invoice and whatever context you thought to include.
It will give you a confident answer anyway. That's what a prediction engine does.
The auditability problem. The third issue is the one that ends the conversation for anyone with financial control responsibility.
As one financial analysis firm observes, LLMs for financial analysis function as opaque systems, making it hard to retrace how a specific output was formed. In regulated industries, this lack of clarity can conflict with documentation and audit trail requirements. If a model identifies a transaction as suspicious, regulators expect the reasoning behind it.
When ChatGPT helps your AP team approve an invoice, where does that decision live? In a chat window. Not timestamped in your system of record. Not linked to the invoice. Not part of your financial controls documentation. Not available to an auditor asking why invoice #4471 was approved and on what basis.
And this is no longer a matter of best practice. Academic work on financial AI notes that the EU AI Act mandates compliance for high-risk financial AI systems by August 2026, requiring human oversight with interpretable outputs (Article 14) and accuracy guarantees (Article 15). The regulatory floor is rising to meet exactly the gap that general-purpose tools leave open.
The shadow AI problem. There's a fourth issue that most finance leaders haven't fully registered, and the MIT research surfaced it plainly. The report describes a shadow AI economy, where employees in over 90% of firms use personal AI tools even when official pilots fail.
Translate that into an AP context. Your team is almost certainly already pasting invoice data, vendor terms, and payment details into consumer AI tools, outside any data governance framework, with no audit trail, and no visibility for you. The question was never whether AI enters your AP process. It already has. The question is whether it enters through a governed system or through a browser tab.
The Core Distinction: An Assistant Versus a System
Everything above points at a single distinction, and once you see it, the whole category makes sense.
MIT's researchers put their finger on it when they described why enterprise deployments stalled: most tools cannot retain feedback, adapt to context, or improve over time. And the report found that even purpose-built enterprise systems were being rejected; 60% of organisations evaluated such tools, but only 20% reached pilot stage and just 5% reached production, most failing due to brittle workflows, lack of contextual learning, and misalignment with day-to-day operations.
The lesson isn't "AI doesn't work in the enterprise." It's that a conversation doesn't work as an infrastructure.
A general AI assistant is stateless. It has no memory between sessions, no connection to your systems of record, no persistent view of your vendors, no access to your contracts, and no ability to act on anything it concludes. It produces outputs. A human carries every output into another system. That human is the integration layer, and at enterprise volume, that human is the bottleneck that defeats the entire purpose.
A system is different in kind. It has memory. It holds state. It connects to your ERP. It reads your contracts and holds them as active rules. It monitors continuously rather than when someone remembers to run a prompt. It takes action rather than producing suggestions. And it logs every decision as a byproduct of doing the work, not as a separate documentation exercise.
That's not a better assistant. That's a different category of thing.
What Blackbee AI Does Differently
Blackbee AI was built for exactly this gap, as an agentic Intake-to-Pay platform, not an AI assistant with finance features bolted on. The distinction shows up in every dimension where the research says general tools fail.
Blackbee AI has memory and persistent context. MIT identified the absence of memory as the structural gap: tools that forget context, don't learn, and can't evolve. Blackbee AI holds a continuous, persistent picture of your AP operation: what you paid each vendor, which exceptions each supplier has generated, what normal billing looks like for this relationship. When an invoice arrives, Blackbee AI isn't reasoning from a blank slate. It's reasoning from everything the organisation already knows. The context that a ChatGPT conversation loses when you close the tab is the context Blackbee AI treats as the foundation of every decision.
Blackbee AI is grounded, not free-floating. The Gartner finding that ungoverned data produces fabricated answers at dramatically higher rates than governed data is the single most important design constraint for a finance AI system. Blackbee AI doesn't ask a model to guess whether a price is correct. Its Clause Agent reads your vendor contracts, extracts the rates, caps, discount triggers, and payment terms, and holds them as live, checkable guardrails. When an invoice arrives, Blackbee AI validates it against the actual contractual terms, not against a plausible-sounding inference. The grounding isn't a retrieval trick layered on top. It's the architecture.
Blackbee AI relieves the model of arithmetic responsibility. The academic research is explicit that LLMs should be surgically relieved of numerical computation in favour of deterministic verification. Blackbee AI's Parse Agent extracts and confidence-scores every field on every invoice, then subjects the arithmetic to deterministic validation, line totals against quantity times price, subtotals against line sums, totals against subtotal plus tax. The reasoning is agentic. The arithmetic is verified. Blackbee AI doesn't ask an LLM to be a calculator because the research says LLMs shouldn't be calculators.
Blackbee AI acts rather than suggests. A general AI tool produces an output that a human must carry into another system. Blackbee AI's Sync Agent connects directly to NetSuite, Sage Intacct, Dynamics 365, Workday, or SAP and posts validated decisions back as clean transactions. Its Route Agent doesn't draft an approval email; it routes the approval through your policy, enforces the threshold, escalates on SLA breach, and refuses to advance an invoice that hasn't cleared the required sign-off. Where ChatGPT ends its involvement at the suggestion, Blackbee AI carries the decision through to the system of record.
Blackbee AI produces an audit trail as a byproduct of the work. The auditability problem, opaque reasoning that can't be retraced, is the one that disqualifies general tools from financial control entirely. Every decision Blackbee AI makes is logged automatically: what data was evaluated, what policy was applied, what the confidence level was, what the reasoning was, and who confirmed it. When an auditor asks why a specific invoice was approved, the answer is in the system, timestamped and linked to the transaction. Not in a chat window, someone may or may not have saved.
Blackbee AI monitors continuously. ChatGPT analyses what you give it, when you give it. Blackbee AI's Spend Intelligence Agent watches continuously, surfacing a vendor whose pricing has drifted the moment the invoice arrives, flagging an anomaly before payment rather than in next month's report. And its Trust Agent scores vendor risk continuously rather than at onboarding, so today's decision reflects today's risk profile, automatically.
Blackbee AI governs upstream, where general tools never look. Every general AI tool begins when you hand it a document. Blackbee AI begins at spend intent, the moment someone proposes a purchase, before any commitment exists. Its Intake Agent captures the request; the Clause Agent checks it against the governing contract; the Route Agent enforces the approval policy. By the time an invoice arrives, Blackbee AI already knows it was coming, who approved it, and what it should cost. The exception is prevented rather than detected.
And critically, Blackbee AI operates above your ERP rather than inside it. Your ERP remains the authoritative system of record, unmodified. Blackbee AI governs the decisions that produce ERP transactions and hands the ERP the validated outcome, which means the integration doesn't carry the brittleness and upgrade fragility that MIT identified as a leading cause of enterprise AI failure.
What This Doesn't Mean
Let's be fair to the general tools, because overstating the case would be its own kind of error.
ChatGPT, Claude, and Gemini are excellent at what they're built for. Reading an unusual invoice format. Drafting a vendor dispute email. Interrogating a messy dataset. Answering a question nobody anticipated. These are genuinely valuable capabilities, and any AP team that isn't using them for that class of work is leaving productivity on the table.
Blackbee AI doesn't ask you to give them up. The two coexist naturally because they serve different functions. Use a general AI assistant for the unstructured, one-off, in-the-moment work. Use Blackbee AI for the structured, high-volume, governance-critical work that constitutes the actual AP process.
The mistake isn't using ChatGPT. The mistake is mistaking ChatGPT for infrastructure.
The Bottom Line
MIT found that 95% of enterprise AI pilots produce no measurable return and identified the cause as structural: general tools forget context, don't learn, and don't integrate with real workflows. The academic literature shows frontier models carrying 10-20% error rates on exactly the kind of multi-step numerical reasoning that an invoice requires. Gartner's data shows that ungoverned data drives fabrication rates above 50%. And the regulatory floor is rising to demand interpretable, auditable financial AI by August 2026.
Against that evidence, the case for running enterprise accounts payable on a general-purpose AI assistant doesn't survive contact with the details. Not because the models are weak, they're remarkable, but because a conversation is not a controlled environment.
What enterprise AP requires is memory, grounding, deterministic verification, the ability to act, continuous monitoring, and an audit trail that exists whether or not anyone thought to save it. Blackbee AI was designed around each of those requirements because each one is a place where the research says general tools break.
The opportunity MIT identified is real: back-office automation is where the highest AI ROI actually lives. Capturing it just requires the right kind of system. Blackbee AI is that system, the agentic Intake-to-Pay platform that governs every spend decision from intent through to payment, above your ERP, with the contract intelligence, continuous monitoring, and complete auditability that a chat window can never provide.