Fundamentals 25 min read

8 Palantir Ontology Anti-Patterns That Break Your Data Model

This article details eight common Palantir Ontology anti-patterns grouped into four root causes—system silos, department silos, overloaded objects, wrong execution layers, and poor naming—and provides a prioritized repair checklist to keep ontology models aligned with business reality.

Linyb Geek Road
Linyb Geek Road
Linyb Geek Road
8 Palantir Ontology Anti-Patterns That Break Your Data Model

Previous articles in this series explained how to build an ontology correctly by modeling the business world. However, many model problems do not appear in the first version. When the first department onboards, Customer is clean; the second department adds Sales Customer; the third, unwilling to change others' work, creates Billing Customer. Six months later, the same customer has four identities, dozens of statuses, and scores of Actions in the ontology.

Each decision seems reasonable in isolation, but together the model spirals out of control. Palantir's official documentation summarizes 8 common anti-patterns . They are not obvious technical errors; rather, they are convenient shortcuts that accelerate initial delivery but defer rework until the model needs reuse.

The author reorganizes these 8 anti-patterns into four root-cause categories and provides a repair sequence:

Root cause: Treating source systems as reality → Anti-patterns: System Silos , Department Silos → Symptom: Same person or customer appears as multiple object types.

Root cause: Unwilling to make trade-offs → Anti-patterns: Kitchen Sink , God Object → Symptom: Excessive fields, or one object representing too many entities.

Root cause: Choosing the wrong execution layer → Anti-patterns: Golden Hammer , Action Sprawl → Symptom: All logic crammed into one tool type, actions becoming increasingly fragmented.

Root cause: Distortion of time and language → Anti-patterns: Time Machine , Misnomer → Symptom: Historical versions become object copies; names rely on verbal explanation.

These four problem categories reinforce each other. For example, departments each building their own Customer often retain technical fields; to accommodate multiple customer views, a God Object full of nulls emerges; finally, the same modification action gets copied multiple times. Therefore, fixes must target root causes first, not individual tables or Actions.

System Silos: Number of Source Systems ≠ Number of Real-World Objects

System Silos is the most common issue. A company has HR, badge, and project management systems, so the ontology contains:

HR System Employee
Badge System Employee
Project Management Employee

All three objects query normally and have their own primary keys. The trouble: they describe the same real-world Employee. When the business asks "What projects is this employee responsible for? Does they have campus access? Who is their manager?" users must jump across three objects. Applications, Links, and Actions must be duplicated.

The correct direction is usually not to keep three "system employees" in the ontology, but to perform identity resolution, field merging, and conflict prioritization in the data pipeline, then expose a single Employee upward.

Unifying the object does not mean deleting source information. Source system, extraction time, and matching criteria can remain in the underlying data for debugging and audit; they just shouldn't dictate business object boundaries.

Three questions to check for System Silos:

Can these objects point to the same person or thing in reality?

Do users often need to join them to complete a task?

Are multiple Actions and applications needed only because data comes from different systems?

If all three answers are "yes", the problem is likely not in Ontology Manager but in incomplete upstream identity resolution and data governance.

Department Silos: Org Chart Changes Should Not Force Ontology Redesign

Department Silos mirrors System Silos, but the split is by department instead of system. Sales builds Sales Customer, Support builds Support Customer, Finance builds Billing Customer, Marketing builds Marketing Contact.

Each team says "Our customer is different." Half true: different departments care about different attributes, relationships, and workflows, but the real-world customer is usually the same.

A more robust model shares one Customer and uses properties and relationships to carry departmental perspectives:

Sales cares about salesStatus and Opportunity.

Support cares about supportTier and Support Ticket.

Finance cares about Billing Account and Invoice.

Different applications present different views; sensitive fields remain permission-controlled.

Don't replicate business entities to give teams "control." Autonomy should live in extensions, applications, and permission boundaries—not in manufacturing multiple copies of reality.

Is the ontology describing reality, or is it diagramming the company's system inventory and org chart?

Kitchen Sink: Including Every Field Seems Safe but Pushes Judgment to Consumers

Kitchen Sink" literally means "throw everything in." When building Customer from CRM, besides ID, name, email, all these fields become Properties:

_crm_extracted_at
_crm_received_at
_crm_sequence
_crm_table_version
last_etl_update_timestamp

The rationale: "Might be useful later." The cost is direct: business fields drown in technical metadata, search and indexing burden grows, new users don't know which fields are trustworthy, and agents guess which timestamp means "customer updated time."

This doesn't mean underlying data must drop technical metadata. Palantir docs suggest keeping technical metadata in the backing dataset for data engineering and troubleshooting, but not necessarily exposing it as an Ontology Property.

A Property is worth exposing only if:

Someone needs to view, search, filter, or act on it.

Its business meaning can be stated in one sentence.

Deleting it would break a specific workflow.

If you can't answer, don't expose it. "Put everything in now, clean up later" rarely happens. Once fields are used by applications and APIs, deletion cost only rises.

God Object: Trying to Reduce Object Count Creates an Object Nobody Can Explain

Kitchen Sink stuffs too many fields into one object; God Object goes further: one object represents multiple distinct entities. The official docs give a classic example: trucks, machines, software licenses, real estate, financial instruments, even employees all shoved into Asset.

Signals appear quickly:

Property count grows, but most fields on most objects are null. value, location, status change meaning depending on assetType.

Validations and Actions are full of if assetType = ... branches.

Searching Asset returns a pile of incomparable things.

At this point, "Asset is a unified abstraction" exists only in name. Different real-world entities should be split into Equipment, Vehicle, Software License, Property, etc. If they truly share attributes or behaviors, use an Interface (e.g., Depreciable Asset) to express that.

As discussed in the Interface article: shared fields ≠ shared capability. God Object makes the opposite mistake: it treats "sounds like they belong together" as "they should share a lifecycle, rules, and actions."

Golden Hammer: The Team's Favorite Tool Becomes the Universal Solution

Golden Hammer means: "When you only have a hammer, everything looks like a nail." In ontology projects, the hammer might be Action, Pipeline, Automation, or Function.

Common scenes:

Daily regional sales rollup—better precomputed by Pipeline—becomes a user-clicked Action.

New alert needs on-call assignment—suited for Automation reacting to events—stays in a polling Pipeline.

Simple derivation like fullName = firstName + lastName gets written as a real-time Function.

Complex cross-object real-time validation gets stuffed into batch processing, forcing users to wait for the next build.

The tools aren't wrong; their placement is. First classify by "who triggers, when results are needed, compute scale":

Batch / Streaming Pipeline → Cleansing, aggregation, bulk or continuous data processing.

Automation → Automatic reaction to Ontology events.

Function → Complex real-time logic depending on current ontology state.

Action → User-initiated, business-meaningful modifications.

This table isn't a universal router, but it catches the most common misuses. The real danger is when tool choice becomes team identity: "We do Pipelines" "We only use Functions." Once discussion stays at tool preference, business latency, scale, and responsibility boundaries get ignored.

Action Sprawl: Actions Are Not Field-Update Interfaces

Action Sprawl often imports traditional CRUD thinking into the ontology. An Employee object spawns dozens of Actions:

Update Employee First Name
Update Employee Last Name
Update Employee Email
Update Employee Department
Update Employee Manager

Official docs flag "single object type has >10 Actions", "multiple Actions always executed sequentially", "Action names heavily prefixed with Set or Update" as typical signals. The number isn't a hard limit; the real problem is that the user's business task gets fragmented.

Employee transfer is a single business operation. It may simultaneously change department, manager, office location, effective date, and trigger permission adjustments. A proper Action is Transfer Employee, not four sequential field updates.

An Action should leave an auditable record:

Who, when, for what business reason, transferred the employee to which department.

If logs only show "manager field changed from A to B", context is lost. This also hurts Agents: exposing 30 single-field tools forces the Agent to guess call order and handle intermediate failures; a well-bounded business Action makes input, validation, and audit far easier to control.

Time Machine: History Is Not Infinite Copies of the Same Object

To preserve contract history, some create:

Contract v1
Contract v2
Contract v3

Even more extreme: Contract 2024 and Contract 2025 as separate object types. Short term, each version retains its state. Long term, relationships become ambiguous:

Which version should Vendor link to?

Which object represents the current contract?

How many deduplications for cross-year stats?

Does a contract amendment require copying all Links?

Palantir docs recommend keeping one Contract representing the real contract, with Properties reflecting current state; historical changes go into linked Contract Amendment / Contract History, or use time series, edit history, and underlying data for audit trails.

Choice still depends on business. If a contract amendment has its own ID, effective date, approval process, and attachments, it likely deserves its own object. Merely saving every field change doesn't justify copying the entire Contract. The criterion remains: how many things exist in reality, not how many historical snapshot rows sit in the database.

Misnomer: One Bad Name Forces Every Caller to Guess Again

The last anti-pattern looks lightest but spreads farthest. Object named Item, Properties named value, type, date, Link named related to.

The original developers know what these mean. A new team only sees questions:

Is Item a product, order line, or inventory record?

Is value an amount, quantity, or risk score?

Is date creation date, due date, or effective date?

Is related to parent-child, substitute, or procurement relationship?

Documentation can explain, but docs shouldn't rescue bad naming. Names are the first semantic layer read by every application, API, and Agent.

Renaming value to monetaryValue or riskScore, date to orderPlacedDate or contractEffectiveDate, related to to Purchased By or Supervised By directly reduces ambiguity.

This is especially critical for Agents. Humans ask colleagues when fields are ambiguous; Agents often pick the "most plausible" interpretation from context. The more capable the model, the smoother the guess—and the easier to mistake that fluency for true business understanding.

How to Fix: Don't Start with the Easiest Changes

After spotting anti-patterns, teams tend to fix field descriptions, add docs, clean a few Actions—because those deliver quickly. A more effective sequence tackles the issues that distort "object identity" first.

Priority 1: How Many Objects Exist in Reality?

Address System Silos , Department Silos , and God Object first. If the same real entity is duplicated, or multiple entities are crushed into one object, subsequent Properties, Links, and Actions are built on a faulty foundation.

Priority 2: What Does Each Name and Action Actually Mean?

Address Misnomer and Action Sprawl . Once object boundaries stabilize, unify business language and consolidate fragmented field updates into complete business operations.

Priority 3: Which Execution Layer Should Run the Logic?

Address Golden Hammer . Review each existing Pipeline, Automation, Function, and Action: who triggers, what latency is required, does it need real-time ontology state, does it involve human judgment.

Priority 4: Clean Up Exposure Surface and Historical Structures

Address Kitchen Sink and Time Machine . Hide or remove Properties with no business use; migrate historical copies into clearer version, amendment, or time-series models.

A Checklist for Ontology Design Reviews

When reviewing an existing ontology, ask these 8 questions directly:

Is the same real-world entity modeled as multiple Object Types because of different source systems?

Is the same customer, employee, or supplier modeled separately by different departments?

Are there Properties that only serve ETL debugging and are never viewed or filtered by the business?

Does an object have many nulls, with Property meanings dependent on a type field?

Does the team habitually use the same tool for almost all computation and workflow problems?

Does a business task require the user or Agent to chain multiple single-field Actions?

Are historical versions copied into multiple objects, making the "current object" hard to identify?

Can Object, Property, and Link names be understood without the original author?

Don't treat this as a pass/fail scorecard. Some systems genuinely need multiple source objects; some historical versions have independent business identity. Anti-pattern judgment is about form—it's about whether identity, semantics, and action boundaries become increasingly inexplicable.

Conclusion

The hardest part of ontology design is that bad design usually still runs. It queries, renders UIs, and lets Agents call it. Until the third department, second application, and first large-scale change arrive together, the early shortcuts start collecting interest.

These 8 anti-patterns compress into three checks:

Does the model describe reality, or the systems and org chart?

Can every object, field, and action be accurately understood without the original author?

Are logic and actions placed in the execution layer that fits them?

When answers get fuzzy, it's time to stop and pay down the debt.

References

[1] Palantir Docs: Ontology design — Anti-patterns<br> https://www.palantir.com/docs/foundry/ontology/ontology-anti-patterns/

[2] Palantir Docs: Ontology design — Best practices<br> https://www.palantir.com/docs/foundry/ontology/ontology-best-practices/

[3] Palantir Docs: Ontology design — Structural guidance<br> https://www.palantir.com/docs/foundry/ontology/ontology-structural-guidance/

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

data modelingNaming ConventionsAnti-patternsAction SprawlGod ObjectOntology DesignPalantir OntologySystem Silos
Linyb Geek Road
Written by

Linyb Geek Road

Tech notes

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.