Everybody Built The Pipeline, Nobody Built The Refinery: Managing Data In An AI World
Shawn Rosemarin is the Global Vice President, R&D, Customer Engineering at Pure Storage.
gettyYou’ve likely been in this meeting: Two leaders walk in with different numbers for the same metric. The next 20 minutes are wasted arguing over whose data is right. By the time it’s settled, the meeting is over, and a decision never gets made.
I’ve witnessed this more times than I care to admit. It’s frustrating, expensive and a complete waste of the leadership team’s bandwidth.
The truth is that, in many cases, both numbers are technically correct. They came from the same expensive platform you spent three years consolidating, pulling from the same underlying data. The problem was that different teams had different ideas about what the data meant. One team might count a customer as “active” based on a recent transaction, while another might use an open account as the threshold. The platform didn’t fail you. The definitions were never aligned.
That isn’t a reporting failure. It’s a data refinement failure, and it’s probably why your AI pilot is stalling.
We’ve all heard “data is the new oil.” It’s a tired cliché, but most leaders miss the critical engineering reality behind it.
Oil’s value isn’t in the ground. It’s created in the refinery—the machinery that converts contaminated crude into standardized fuel that engines consume. You don’t pour crude into a jet engine. The refined specification is the product.
When a peer brags about their “data lake,” I still hear more about reserves: volumes, aggregation, scale-out analytics—it’s about quantity. I hear less about value.
• Upstream (Drilling): Capture data and establish meaning at the source. You may not control what’s in the ground, but you can control how it’s labeled.
• Midstream (Pipes And Tanks): Move and store it. This is where many companies stop, optimizing for volume and throughput rather than usability.
• Downstream (Refinement And Distribution): Turn raw data into something teams and AI can actually consume: consistent definitions, trusted identities and usable context.
Most organizations think they’ve reached the finish line when they’re still stuck in the middle. I see this as the “death zone” for AI projects.
Gartner indicates that 63% of data leaders lack or are unsure they have the foundational practices required for AI. They predict that through 2026, 60% of AI projects will be abandoned. The reason isn’t model failure, but unrefined data. Raw data was never the product.
I’ve found it helpful to think about data as having “grades.” Structured data, such as database tables and CRM records, is like light crude: relatively easy to refine, but limited in volume and variety. Semi-structured data, such as productivity documents, logs and nested files, is more like heavy crude. It requires more mapping and enrichment before it becomes useful. Unstructured data, including photos, videos, call recordings and emails, is closer to oil sands: abundant, but far more expensive to recover and refine.
For oil sands, an “upgrader” is an industrial facility that converts heavy bitumen into lighter synthetic crude oil that can be more readily transported and refined. Without upgrading, raw bitumen is too heavy and viscous for many conventional refineries to process efficiently.
Unstructured data is no different. Before a model can utilize a contract, something must read it, link it to a customer identity and attach permissions. Too many companies pipe raw bitumen into a model and wonder why the results are garbage.
Shammy Narayanan, the senior vice president for data, AI and architecture at Welldoc, told me that over half of clinical data in healthcare is stranded behind technical debt. This isn’t a storage problem—we already pay for the bytes. It’s an actionable data problem.
This is where I might lose some of you. You should not process every barrel. If recovery and refinement costs exceed expected value, you leave it in the ground.
Raw storage is affordable. Refinement is the expensive part, and it scales with how much you decide to refine. In the AI era, a strategy to extract the optimal utility of data includes deciding which data to exclude; the cost of refinement is not worth the incremental value it will provide.
Most organizations have never made that decision explicitly. They’ve made it by accident by refining whatever the loudest project asked for. To turn the idea of a data refinery into a build plan, think of it in terms of six core processes:
1. Distillation: Separate inputs and tag them with lineage and permissions.
2. Desalting: Remove noise, duplicates and unusable personal data.
3. Cracking: Break messy logs and other complex inputs into small, usable facts.
4. Reforming: Consolidate records into a trusted source of truth, such as a single customer identity.
5. Blending: Establish and enforce a consistent definition for every metric.
6. The Lab: Give every data product a specification, and document what happens when a bad batch reaches production.
In AI innovation, “yield” is the share of raw info that becomes safely usable for a business decision. Gartner shows that 59% of companies don’t measure data quality, even though poor quality costs millions.
In my experience, maybe five out of 400 use cases have the data ready to support them. That’s barely a 1% yield. Everyone is investing in producing code faster, but few measure the return.
To increase that yield, stop thinking of governance as a gate. In a “data intelligence” world, governance is a product specification. It’s what makes reuse and scale possible.
Measurable quality, traceable lineage and permissions enforced at the point of use—that’s the goal. Done correctly, governance stops being the department that says “no” and becomes the engine that says “go.”
On another podcast, Armel Roméo Kouassi of Northern Trust put it well: “I cannot tell the regulator that...these earnings increases are due to AI.” You need the receipts. “The model decided” is not a legal defense. Refining gives you the lineage and permissions to defend your results.
Building pipelines was the heavy lift of the last decade, but a lake is not a product. The leaders who build the refinery can better see compounding returns, with their next project starting at week three, not week one.
The models are ready. The question is, what are you feeding them?
Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?
