Procurement

Manufacturing procurement data: the SAP problem hard FM already solved

95% of manufacturing spend auto-classified, without touching the SAP structure underneath it.

Procurement27 August 20265 min read

Quick answer: A manufacturer with years of history in SAP usually has all the data it needs and no way to use it. Item descriptions vary by whoever entered the purchase order. Supplier names drift across plants. Categories were set up locally by whoever configured the ERP years ago. This is not a missing data problem. It is a structure problem, and it is the same one hard FM has with invoices, just with parts and supplier codes instead. The fix is the same three steps in both cases: resolve the supplier list, apply a taxonomy that cannot invent its own categories, then classify every line.

On this page: Why manufacturers assume this is a data problem · What does unclassified spend actually cost? · Who is responsible for fixing it? · What this looks like on a real SAP export · Why not just point a general AI tool at the export? · What the first step actually looks like · FAQ

Why manufacturers assume this is a data problem

If we invest, we check how the company has performed before we commit. If we plan a route to an important meeting, we account for the traffic. Most decisions we trust are decisions we have already structured, even if we do not think of it that way. Procurement spend is no different. A manufacturer with a decade of SAP history has more transactions than it will ever need. What it does not have is a consistent category on each one, a single resolved supplier per entity, or a taxonomy that survives contact with a second plant.

That is why "we have all the data" and "we cannot answer a basic question about it" are true at the same time. The transactions exist. Nobody built the structure that would let anyone query them.

What does unclassified spend actually cost?

Legacy classification tools generally land at 75 to 85 percent accuracy on structured spend. The remaining 15 to 25 percent is usually tail spend, services, and P-card transactions, which is exactly where supplier consolidation and price benchmarking would do the most good. Anything above roughly 10 percent unclassified is worth treating as a live gap, not background noise.

Poor master data costs organisations an average of 12.9 million dollars a year. For a manufacturer running material master records across several plants and ERP instances, that number climbs, because the same part often exists under three different material numbers, one per site that set it up independently. That produces duplicate purchase orders, redundant inventory sitting in two warehouses under two part numbers, and a maintenance team searching for a part that technically exists in the system, just not under the code they are searching.

Who is responsible for fixing it?

Nobody, usually. Finance owns the ledger. Procurement owns the negotiation. Neither owns the taxonomy, so it degrades quietly until someone needs a category-level view and discovers there isn't one.

The fix is not a new system. It is the same three-step method that works on any procurement dataset with no structure.

The first step is: resolve the supplier list. Supplier identity is one of the strongest signals for classifying a line correctly, and deduplicating the list is usually how a company first discovers that two plants have been buying from the same distributor under different names, at different prices.

The second step is: apply a constrained taxonomy. A part cannot be given a category that does not exist in the standard. This is the step that stops the system inventing its own codes.

The third step is: classify every line and route the uncertain ones to a person. Confidence scoring is what makes this repeatable. A description that is only a part number, or an item that could sit in two families, goes to review instead of getting guessed at.

What this looks like on a real SAP export

We ran this on a mid-sized manufacturer's SAP data: years of procurement history with no consistent categorisation, no visibility into whether the same part was bought at different prices across sites, no way to see which suppliers could be consolidated. We resolved the supplier list, applied a constrained taxonomy, encoded the categorisation rules the procurement team actually used, and classified the backlog. 95 percent of lines were auto-classified directly against the existing SAP structure. That is what turned "we have the data somewhere" into a category-level view the team could negotiate from.

The same unstructured-data problem shows up outside spend, too. One air filtration manufacturer we work with had site visit reports arriving in every format imaginable: handwritten notes, inconsistent layouts, misspelled part names, that had to be retyped by hand into a proposal document the business could trust. Different data. Same failure: information that existed but was never organised into something usable.

Why not just point a general AI tool at the export?

It will produce an answer. It will not produce a repeatable one. A general-purpose model invents category codes that do not exist in your taxonomy, classifies the same description differently depending on the run, and gives no signal about which of its answers to trust. On a one-off sample of a few hundred lines, that is tolerable. On a live SAP feed that has to stay accurate as new purchase orders land every day, it is not, because the whole point is that the answer still has to be correct next quarter.

A pipeline is different from a classification exercise in one respect: it holds the taxonomy constant and resolves supplier identity before it guesses at a category, instead of starting from zero each time. New purchase orders get classified the same way the backlog was. That consistency is the only reason a 95 percent figure still holds six months later instead of drifting back into an unclassified pile the moment nobody is watching it.

We have not tested this on every ERP configuration a manufacturer might run. What we can say is what the method requires: a supplier list that gets resolved before classification starts, a taxonomy the system cannot deviate from, and a review step for the lines it is genuinely unsure about.

What the first step actually looks like

Start with the supplier list, before touching a single category. If two sites turn out to be buying from the same distributor under different names, that is usually the first finding. Then decide how deep the taxonomy needs to go. Most manufacturers do not need commodity-level detail everywhere. They need it where they are actually negotiating, and family or class level everywhere else. Write the edge cases down with the people who actually buy: what counts as MRO versus capital spend, how subcontracted labour gets coded, what happens to freight and surcharge lines that are not really a product at all. Then classify, score confidence, and route what the system is unsure about to a person, because that is where the rules that stop the same ambiguity recurring get written down.

Frequently asked questions

Why does a manufacturer with years of SAP history still have unclassified spend data?

Because item descriptions vary by whoever entered the purchase order, supplier names drift across plants, and categories were set up locally when the ERP was configured. It is a structure problem, not a missing data problem: the transactions exist, but nobody built the taxonomy that lets anyone query them consistently.

How much does poor master data cost manufacturers?

Poor master data costs organisations an average of 12.9 million dollars a year. For manufacturers running material master records across multiple plants and ERP instances, the same part often exists under several different material numbers, producing duplicate purchase orders and redundant inventory.

What accuracy do legacy classification tools typically achieve on spend data?

Legacy classification tools generally land at 75 to 85 percent accuracy on structured spend. The remaining 15 to 25 percent is usually tail spend, services, and P-card transactions, which is exactly where supplier consolidation and price benchmarking would help most.

What are the steps to classify unstructured SAP procurement data?

Three steps: resolve the supplier list to find duplicate suppliers bought under different names, apply a constrained taxonomy so a part cannot be given a category outside the standard, then classify every line and route uncertain cases to a person through confidence scoring.

Why not use a general AI tool to classify SAP spend data instead?

A general-purpose model invents category codes that do not exist in your taxonomy and classifies the same description differently across runs, so it produces an answer but not a repeatable one. On a live SAP feed that needs to stay accurate as new purchase orders land daily, that inconsistency is not tolerable.

How does Pearstop classify manufacturing procurement data in SAP?

Pearstop resolves the supplier list, applies a constrained taxonomy, encodes the categorisation rules the procurement team already uses, and classifies the backlog. On one mid-sized manufacturer's SAP export, this auto-classified 95 percent of lines directly against the existing SAP structure.


Free resources

Stephanie Wiechers

Stephanie Wiechers

CEO & Co-founder, Pearstop

Stephanie leads Pearstop's go-to-market and strategic direction. She works directly with procurement and FM leaders across Europe to understand how data quality affects margins, contracts, and AI readiness.

LinkedIn →

Further reading

Latest Insights

Procurement

Manufacturing procurement data: the SAP problem hard FM already solved

95% of manufacturing spend auto-classified, without touching the SAP structure underneath it.

Read more
AI & Digital

How to build a spend cube with AI

The five steps to building a spend cube, what data you need, what usually goes wrong, and when a con…

Read more
Procurement

UNSPSC classification in construction: a practical guide

How construction and engineering teams apply UNSPSC to procurement data, who owns the decisions, how…

Read more