Lessons from managing AI transformation in a $40 billion contract portfolio
Artificial Intelligence (AI) & Machine Learning (ML)

Lessons from managing AI transformation in a $40 billion contract portfolio

October 1, 2026/6 min read

At Amazon, I was the Senior Technical Product Manager for a cross-functional team of procurement managers, engineers, data scientists, and legal and operations partners in the last-mile supply chain organization. The team procured contracted transportation and logistics capacity to move packages from warehouses to customers, handled cross-border import and export shipments, and contracted third-party warehouse storage. I managed the Salesforce platform that mapped this end-to-end procurement lifecycle; from bid to award to payment into a single view, and I co-built a machine learning system that pulled legal clauses out of contracts across that same $40 billion portfolio. My role was to own the product and data strategy that got the system to 95% extraction accuracy. The lessons that stuck with me had almost nothing to do with ML model architecture. They were about data that nobody clearly owned.

Contracts at that scale show up in every format you can think of. Years of negotiations, different templates, separate legal teams, several countries. A single liability clause might appear a dozen ways across thousands of documents and mean the same thing every time. Before a model can read any of it, a person has to decide what “the same thing” means across the whole portfolio. That work is slow. It looks boring on a project plan. Most teams skip it.


Since Amazon, I have watched a lot of “AI transformation” projects, and the same thing happens almost every time. A team builds a demo on clean, hand-picked data. Everyone loves the demo. Then production data shows up and looks nothing like the sample.

By then the project has a name and a sponsor. It is “the AI initiative.” Try telling a room full of excited executives that you now need six months on data quality before the real thing works. It rarely goes well. So teams push ahead, ship something that handles a slice of the real cases and park the rest on a future roadmap. The clever AI layer ends up sitting on contract text, product catalogs, and clause libraries that were never built for a machine to read.

That is the main reason AI programs stall after the pilot. The data underneath them was never ready. The unglamorous work nobody wants to fund is the work that decides the outcome.

What getting to 95% accuracy levels took

The accuracy number came from work that happened before we trained anything.

First, a taxonomy. We had to write down every clause type and what counted as an example of each one. That sounds simple until you put legal, procurement, and operations in the same room and realize they have never agreed on a shared vocabulary. A term one team treated as routine, another team read as a liability. Those conversations took weeks, and they were the real foundation of the project.

Then, labeling. My team hand-labeled thousands of real contract sections, including the messy, ambiguous ones that never make it into a demo. We pulled in the people who read these contracts for a living: legal reviewers who knew why a clause was worded a certain way, procurement and operations staff who knew how it played out in practice, and the engineers who had to turn all of it into training data. Nobody could have done this alone. The legal reviewers caught meaning the engineers would have missed, and the engineers caught inconsistencies the legal team had lived with for years.

Last, error checking. We built review steps so that when the model got something wrong, we caught it fast and fixed it before it fed into anything downstream.

None of that is a modeling problem. It is product and data governance. It needs a product owner who will sit with legal, operations, and engineering and get them to agree on what the data means before a single training pipeline runs.

I see the same pattern now in my current work. I lead Salesforce and Vlocity CPQ transformation for T-Mobile through Mphasis, and CPQ and contract lifecycle management (CLM) hit the exact same wall. A pricing engine or an approval workflow can appear fully automated while quietly relying on product catalogs full of duplicate SKUs, inconsistent attributes, and clause libraries maintained separately by three different teams. Any AI you build on top of that, a pricing model or a contract-drafting agent, inherits every one of those problems.

Clean, structured data in the quote-to-cash pipeline is what makes an AI layer possible at all. I think the industry undervalues this badly. The lasting advantage goes to whoever keeps their revenue and contract data clean. The newest model matters far less than most people assume.

What I’d tell a product leader starting this work

Treat data readiness as its own workstream. Give it an owner, a timeline, and a budget. It is not a subtask of model development, and the moment you treat it like one, it loses every priority fight to work that demos well.

Build the taxonomy with your domain experts before anyone labels a thing. Legal, RevOps, whoever owns the data. If they do not agree on definitions up front, you will relabel everything later, and relabeling thousands of documents twice is how timelines slip by a quarter.

Set your accuracy target against a real business decision, like whether extraction is reliable enough to drop manual legal review on a category of contracts. Tie the number to something the business actually does, so “95%” means “we can stop paying people to check these” rather than a score on a slide.

Plan your correction process before you go live. The first thousand production documents will hand you edge cases the training set never had. Decide now who reviews the errors, how fast they turn around, and how a fix gets back into the system without breaking what already works.

And expect a fight over the timeline. The data phase always competes with feature work that produces better demos. The programs that survive have a product leader willing to defend that phase out loud, in the room, when the pressure is to skip ahead to something more visible. I have had to be that person more than once. It is uncomfortable, and it is the difference between a pilot that ships and one that quietly dies.

Wider implications of Data-Driven AI Transformation

The clause extraction system took a large amount of manual legal review out of a $40 billion portfolio. That is AI changing how a real team works every day, well past anything a demo could show. It came from data work that ran longer than we planned and stayed almost invisible to everyone outside the team.

The invisible parts of AI transformation, the contract data, the catalog quality, the clause taxonomies, carry most of the work and most of the risk. Models get the excitement in a steering committee. The data underneath decides whether the thing still works in production a year later.