Share this article:

Beyond the PoC: Designing AI Initiatives That Reach Production on Palantir

Share This Article

The proof of concept is rarely the hard part. Getting past it is.

Almost every organisation we speak with has AI pilots running. Far fewer have AI in production. That gap is the real story. The numbers around it are not encouraging: CEOs reporting no financial return, organisations reporting no enterprise impact, PoCs that never leave the sandbox. Out of every thirty-three AI PoCs, only four reach production. At Unit8, we have worked closely with Palantir Foundry & AIP for years, and those figures align with what we see on the ground when we first walk into a project.

The reason is almost never the system. BCG framed it well: AI success is 10% algorithms, 20% data and technology, and 70% people and process. The algorithm still matters. Without it, nothing works. But it is the cherry on top. The 70% is where PoCs go to die.

The six patterns that keep a PoC a PoC

Across our projects, the same six failure modes come back again and again.

The first is misunderstanding the problem. You solve the wrong thing, and you solve it well. The second is a technology-first mentality, where the team thinks harder about the model than about what the model is for. The third is the last-mile problem, and as engineers, we know it from the inside: we build a clean pipeline and a smart model, we feel good about the dataset, and then we never connect any of it to the workflow people actually use. The fourth is data quality discovered too late. The fifth is missing business ownership, where a project is run by IT alone and the business never owns the problem it is meant to solve. The sixth is measuring the wrong thing, celebrating model accuracy while nobody asks whether decisions got better.

None of these are exotic. Most of them look like Excel sheets, copy-pasted scores, and emails. That is the shape the problem usually takes in real life.

A supplier risk model that never left the notebook

Picture a traditional manufacturer. Call them Orkova. They build automotive and aerospace components, and a few months ago, they built a supplier risk scoring model.

It runs in a Python notebook. Every Monday, it produces a CSV, and the CSV is emailed to the procurement team. The team copy-pastes the scores into another spreadsheet, makes decisions off those numbers, and records those decisions nowhere. The accuracy was 84%. Everyone celebrated.

Now play it forward. The data scientist who wrote the notebook leaves. Nobody can reproduce the pipeline. The model starts to drift quietly, and no one notices for eight weeks. The scores are wrong, but the email still lands every Monday morning, and the procurement team still acts on it. This is what we call the Excel problem, and it kills good projects at the finish line.

The alternative uses the same use case, supplier risk scoring. What changes is the design: it is built for production from the first day.

Eight questions to answer before the first sprint

Production thinking starts before a single line of the PoC is written. We work through eight questions with the business and the engineering team together. The platform gives you the answers by default. The discipline is asking the questions early.

What decision is changing? 

“Can we score suppliers?” is a capability, and plenty of teams can deliver it. The real question is sharper: should a procurement manager escalate this supplier relationship right now? In Foundry, this is where the Ontology earns its place. The Ontology is a digital twin of the process, and every object in it exists because someone has a decision to make. A disruption alert exists so a person can decide whether to escalate. An action type is the production contract: who can act, on what, with what parameters, with what audit trail. The AI  can change, the pipeline can change, the interface can change. The contract holds.

Who owns it in production? 

In the notebook version, ownership is a conversation for month five, and by then the system is an orphan. Encode it on day one instead. Name the point of contact, the business owner who defines what “at risk” means, the engineers who build, and the platform owner who keeps it running. Governance is part of this, not a legal review bolted on at the end. Access markings mean every AI capability sees exactly what the user sees, and nothing more.

Are the metrics tied to business outcomes? 

84% accuracy told us nothing about whether procurement made better decisions. So the KPI became the “average time” from “disruption alert” to “escalation decision”, measured inside the platform rather than reconstructed from a notebook eight months later.

Is the data audited? 

We define data expectations in business terms, not schema checks. A risk score has to sit between zero and one. When a value of 1.45 shows up, the pipeline stops and the Ontology never receives the bad data. That test lives in the pipeline because we asked the data quality question on day one. When compliance asks where a score came from, the lineage is one click away.

Is governance designed in? 

Six months of work can still die when InfoSec/data protection officers  finally asks who can see supplier financials, who is accountable when the AI recommends the wrong thing, and whether every decision can be audited. When everything is recorded from day one, those answers already exist. Edit history, timestamps, and every call to the AIP logic function, including the chain of thought behind each recommendation, are all there to review.

Is the AI output in the workflow people use? 

This is the line between a dashboard and a system. Instead of the Monday CSV, the procurement team opens one operational command centre every morning. Active alerts, values at risk, resolution times. Click an alert and you see the supplier, the severity, the score, the recommended action, and the linked purchase orders, all pulled from the Ontology so nobody logs into a second system. Two decisions are available: approve the recommended action, or override and escalate. Approve it, and the alert resolves, the action logs, the status updates, and the manager is notified. No email bouncing around. No spreadsheet. This is the last mile, closed.

How is it monitored? 

Health checks show job status at a glance. A monitoring view sends notifications, emails, or a pager when something fails, so a broken data set surfaces to the person on call rather than to the procurement team.

How is it handed over? 

Developers rotate and leave. Version history records what was built and when, and evaluation suites, with their code and metrics, catch drift before it reaches the business. The system is designed to survive the people who built it. AI FDE makes documentation for free; no production system can live without proper handover docs, system design, escalation paths, etc., especially with LLMs’ capabilities embedded in the platform.

Start with the decision, not the AI

There is a simple rule underneath all of this: if everyone is accountable, nobody is. 

That is why ownership comes first, and why the first question is about the decision being changed rather than the model being built.

This is not one person changing how they work. It is IT, data, and business teams agreeing on a way of working, and it needs senior leaders to model it. The prompts are simple. Have you asked the AI tool? Have you looked at the data? If you are in the business, have you talked to IT, and if you are in IT, have you talked to the business?

Three ideas hold this together. The platform is rarely the bottleneck: on Foundry & AIP, the Ontology is already there to use. The design of the workflow around the model matters more than the sophistication of the model itself, because a system only creates value once people rely on it. And the work begins from the business decision, before anyone opens an IDE.

What we have seen in practice

The Orkova example is notional, but the pattern is not. We are a certified Palantir partner with more than 80 Foundry and AIP specialists, working with Fortune 500 companies since 2017. Across those programs, the organizations that get the most from Foundry & AIP are the ones that use it to rethink which processes need to exist at all, rather than to automate the spreadsheet they already have.

Take one multinational pharmaceutical manufacturer. Before our implementation of Foundry and AIP, the team investigated around 20 disruption alerts by hand in a whole year. Now they screen roughly 10 a day. That shift freed more than 3,000 hours of sourcing-manager time annually, with an expected 10x ROI.

We also stay involved past go-live. That is where most initiatives quietly stall, and it is the part where a partner earns its place on.

Out of every thirty-three AI PoCs, only four reach production. If you are scoping an initiative on Foundry & AIP, we can run this eight-question checklist against your own use case in a 30-minute working session and tell you, plainly, where the risk sits. That is a more useful first step than a demo.

Want to receive updates from us?

agree_checkbox 

By subscribing, you consent to Unit8 storing and processing the data provided above in order to provide you with the requested content. For more information, please review our Privacy Policy.

Our newsletter features industry news, the latest case studies, and future Unit8 events.

close

This page is only available in english