What is Product Data Quality?

Executive Summary

Most organisations discover the true condition of their product data at the moment they try to publish it. Until then, the data works well enough: a missing weight is filled in by someone who knows the product, a contradictory material figure is quietly reconciled in a spreadsheet, and an expired certificate remains attached to a record because nobody was accountable for noticing. A Digital Product Passport removes those informal repairs. Information must be assembled from several systems, presented coherently, and stood behind in public.

Product data quality is the discipline that makes that possible. It is not a general aspiration to have good data. It is a measurable statement about whether a specific attribute, for a specific product, is fit for a specific use. The same material composition record can be perfectly adequate for internal engineering and inadequate for a regulated public declaration, because the requirement changed, not the data.

This article sets out The Trusted Product Data Quality Model, six dimensions assessed together: Completeness, Accuracy, Consistency, Validity, Timeliness and Traceability. The central point of the model is that these dimensions are independent. An attribute can be present, correctly formatted, internally consistent and recently updated, and still be wrong. That is why passing validation is not the same as being trustworthy, and why quality has to be judged at attribute level rather than declared at system level.

The article also separates three ideas that are routinely conflated: quality, which asks whether information meets the requirements of its intended use; validation, which asks whether information passes defined structural, format, range and business rule checks; and truth and evidence, which asks whether the underlying claim is factually correct and adequately supported. Programmes that treat these as one thing tend to build strong validation, report high quality scores, and remain exposed on exactly the claims that matter most.

Everything here is vendor neutral. The six dimensions are a tieback educational model, not a regulatory or standards requirement, and they are intended to be used alongside, not instead of, the formal data quality vocabulary published by standards bodies.

FrameworkTBF-025
The Trusted Product Data Quality Model

Assesses product information across completeness, accuracy, consistency, validity, timeliness and traceability to judge fitness for intended use.

Table of Contents

Definition

Definition
Product Data Quality

Product data quality is the degree to which product information satisfies the requirements of the use it is put to, assessed per attribute rather than per record or per system. It is expressed through measurable dimensions, typically completeness, accuracy, consistency, validity, timeliness and traceability, and it is meaningful only in relation to a stated purpose: an attribute is not high quality in the abstract, it is high quality for internal planning, for a customer specification, or for a regulated public declaration.

Three clarifications follow directly from that definition.

Quality is relative to purpose. A nominal product weight rounded to the nearest hundred grams is excellent data for a catalogue and inadequate for a shipping declaration. There is no absolute quality level, which is why quality programmes that begin by cleaning everything usually run out of sponsorship before they reach the attributes that carry regulatory consequence.

Quality is a property of attributes, not systems. Statements such as “our product data is 92 percent clean” describe an average across attributes with entirely different consequences. Within that average, a marketing description and a substance-of-concern declaration are weighted identically, which is precisely the wrong result.

Quality is measurable, or it is opinion. The purpose of naming dimensions is to convert vague dissatisfaction into specific, assignable defects: this attribute is missing on 340 products, this identifier fails its check digit, this certificate expired eleven months ago.

The one sentence version

Quality asks whether the information is good enough for what you are about to do with it, and the answer changes when the use changes even if the data does not.

Why Product Data Quality Matters

Product data has always mattered operationally. What changes under passport regimes is exposure, durability and accountability.

Errors become public. An incorrect internal value causes friction; the same value published against a persistent identifier becomes a claim visible to customers, retailers, recyclers and market surveillance authorities, who can compare it against packaging, laboratory results and competitor declarations.

Errors become durable. Published information persists and is versioned. An organisation may be asked years later why it stated a figure, which turns provenance and change history from technical detail into the substance of the answer.

Errors become attributable. Regulation attaches responsibility to economic operators. Data quality defects therefore stop being an internal cost and become a compliance exposure carried by a named party.

Errors become expensive to correct at scale. Correcting one product record is trivial. Correcting an attribute that was systematically mis-defined across a portfolio, and re-publishing every affected passport, is a programme in its own right.

Errors compound downstream. Poor data propagates into recycling instructions, repair information, customs declarations and sustainability reporting, where it is re-used by parties who have no means of detecting that it was wrong.

Example
The cost of a definition, not a typo

A manufacturer records recycled content as a percentage. Engineering measures it by mass of the finished article; a supplier reports it by mass of a single input material; a sustainability team reports it by mass excluding packaging. All three values are entered correctly, pass every format check, and disagree. No individual made an error. The defect is definitional, and it becomes visible only when the three figures must be reconciled into one published claim.

The Trusted Product Data Quality Model

The model assesses product information across six dimensions that operate together. It is deliberately not a staircase or a maturity ladder: there is no order in which the dimensions are satisfied, and no dimension is a prerequisite for another. All six are evaluated for the same attribute, and the attribute is fit for its intended use only when each dimension meets the threshold that use requires.

OutcomeTrusted Product Data

Information an organisation is willing to publish, defend and be held to

01
CompletenessPresence

Is the required information present?

Every attribute the intended use requires exists, at the required granularity, for every product in scope. Absence is recorded explicitly rather than left indistinguishable from an unanswered question.

02
AccuracyCorrespondence

Does the information correctly represent the product or underlying fact?

The recorded value matches the physical product or the verified fact it describes. Accuracy is established by measurement, test, verification or authoritative evidence, never by inspection of the record alone.

03
ConsistencyAgreement

Does the same information agree across systems, suppliers and representations?

One product means one answer, whether read from a design system, a planning system, a catalogue, a label or a passport. Disagreement between sources is a defect even when one of the values is correct.

04
ValidityConformance

Does the information conform to required formats, schemas, vocabularies and rules?

Values match the declared type, syntax, unit, controlled list and business rules for the attribute. Validity is machine checkable, which makes it the cheapest dimension to enforce and the easiest to mistake for quality overall.

05
TimelinessCurrency

Is the information sufficiently current for its use?

The value reflects the product as it is now, and any supporting evidence remains within its validity period. Timeliness requires a last verified date, which is a different fact from a last modified date.

06
TraceabilityProvenance

Can the information be connected to its origin, owner, evidence and change history?

Each value can be followed back to where it came from, who is accountable for it, what evidence supports it and how it has changed. Without traceability the other five dimensions can be asserted but not demonstrated.

Fit for intended use

Framework note: thresholds are set per attribute, per use case, per regulatory requirement and per lifecycle stage. A descriptive attribute and a regulated substance declaration do not warrant the same controls, and satisfying one dimension says nothing about the other five.

The model sits alongside four existing tieback frameworks rather than replacing any of them. The Product Data Foundation Model classifies what kind of data an attribute is. The Trusted Product Data Governance Framework establishes who is accountable and what the rules are. The Enterprise Integration Model moves the data from source systems to the point of publication. The Product Data Maturity Model describes how consistently the organisation can do all of this.

Stated compactly: governance establishes responsibility and control, quality measures whether the data meets requirements, integration moves data between systems, and maturity describes the organisation’s ability to sustain those capabilities. Quality is the measurement layer. It tells governance where to intervene and tells a passport programme what it can safely publish.

The Six Dimensions of Product Data Quality

Completeness

Completeness asks whether the information a use case requires is present, at the granularity it requires. Three failure modes recur.

Absent attributes. The field exists in the model but holds no value for part of the portfolio. This is the easiest defect to count and the one most often reported as the whole of data quality.

Insufficient granularity. A value exists at product model level when the use requires it at batch or item level, or a material is recorded as a family when the requirement is a specific substance. The record looks complete and is not.

Unrecorded unknowns. An empty field can mean the value does not apply, has not been collected, or was refused by a supplier. Collapsing three different situations into one blank makes the backlog impossible to prioritise. Mature programmes record a reason code alongside the absence.

Completeness must be defined against a requirement set, not against the schema. A system with two hundred available fields is not incomplete because a hundred are empty; it is incomplete when one of the fifteen attributes a passport requires is missing.

Accuracy

Accuracy asks whether the value corresponds to the product or fact it describes. It is the only dimension that cannot be assessed by examining the record, which is why it is systematically under-measured: every other dimension can be checked by software against the data itself, and accuracy requires reference to something outside the data.

Accuracy is established by measurement, laboratory testing, physical inspection, an authoritative third party declaration, or documented calculation from verified inputs. It is maintained by re-verification on a defined cycle and on material change, because a value that was accurate at design freeze can be inaccurate after a component substitution nobody propagated.

An important nuance: accuracy is bounded by definition. A value cannot be accurate if the attribute it populates has no agreed meaning, because there is no fact for it to correspond to. Definitional disputes therefore have to be resolved before accuracy can be measured at all.

Consistency

Consistency asks whether the same information agrees wherever it appears. It has three distinct forms.

Cross-system consistency. A design system and a catalogue hold different weights for the same product. One may be correct, but the organisation cannot demonstrate which without further work, so both are unusable for publication until reconciled.

Cross-representation consistency. The published passport, the printed label, the product specification sheet and the customs declaration state different things. This is the form most visible to external parties and the most damaging to trust.

Internal logical consistency. Values within a single record contradict each other: component masses that do not sum to the declared product mass, a recycled content figure that exceeds the mass of the material it applies to, or a country of manufacture inconsistent with the declared manufacturing site.

Consistency defects are frequently symptoms of an ownership gap rather than a data entry gap. Where no single system is authoritative for an attribute, divergence is not an accident; it is the predictable result of the architecture.

Validity

Validity asks whether the value conforms to its declared rules: data type, syntax, unit, permitted range, controlled vocabulary, schema and business rules. It is the dimension that automated checks handle well, and it should be enforced at the point of entry rather than at the point of publication.

Validity is necessary and never sufficient. A GTIN can be fourteen digits with a correct check digit and still identify a different product. A country code can exist in the reference list and still be the wrong country. Confusing validity with quality is the single most common analytical error in product data programmes, and it is the reason the next two sections exist.

Timeliness

Timeliness asks whether the value is current enough for its use. Three facts are needed to answer that, and most systems record only one of them.

Last modified tells you when the record changed. It says nothing about whether anybody has confirmed the value is still right.

Last verified tells you when the value was last checked against reality. This is the fact timeliness actually depends on, and it is the field most often missing.

Valid until tells you when supporting evidence expires. Certificates, test reports and supplier declarations have validity periods, and an expired certificate turns a previously supported claim into an unsupported one without anything in the record changing.

Timeliness thresholds vary enormously by attribute. A dimension that has not changed in six years may be entirely current; a recycled content figure from a supplier who has since changed input streams may be stale after six months.

Traceability

Traceability asks whether a value can be connected to its origin, its accountable owner, its supporting evidence and its change history. It is the dimension that converts assertion into demonstration.

Four links are needed. Origin: which system, supplier or process produced the value. Ownership: who is accountable for it now. Evidence: which document, test report or declaration supports it, and whether that evidence is retrievable. History: what the value was previously, when it changed, and on whose authority.

Traceability is what allows an organisation to answer a challenge. Without it, the response to a regulator or customer is an assurance rather than a record, and the distinction between those two is the distinction between a defensible position and a hopeful one. It is closely related to product traceability in the physical sense but is not the same thing: this is traceability of the information, not of the goods.

Data Quality vs Data Validation

These are different questions, applied at different points, answering to different authorities.

Data quality asks whether the information meets the requirements of its intended use. It is a judgement made against a purpose, and it spans all six dimensions.

Data validation asks whether the information passes defined structural, format, range or business rule checks. It is a mechanical test made against a rule set, and it principally addresses validity, with partial coverage of completeness and internal consistency.

Validation is a component of quality assessment, not a synonym for it. Seven distinct classes of control appear in practice, and they serve different purposes.

Checks that the value conforms to its declared type and schema position: a number where a number is expected, a required object present, a repeating group correctly structured. Catches integration and mapping faults. Says nothing about whether the value is right.

Checks that the value follows the expected syntax: an identifier of the correct length with a valid check digit, a date in the required form, a percentage expressed as a number rather than a string with a symbol. Catches transcription and entry faults.

Checks that the value exists in an approved list: a country code present in the reference register, a material term drawn from the agreed taxonomy, a unit from the permitted set. Eliminates free text drift, and is the highest value control per unit of effort in most portfolios.

Checks that values which should logically agree actually agree: component masses summing to the total, a recycled content percentage not exceeding one hundred, a manufacturing date preceding a shipment date, a hazard classification consistent with a declared substance.

Compares the same attribute across two or more systems and raises a defect where they disagree. This is the only control that detects consistency failures, and it is the one most often absent, because it requires an integration path that exists solely to compare rather than to move data.

Checks that a claim has supporting documentation of the required type, that the document is retrievable, and that it refers to the product or batch in question. Addresses the gap between a value being recorded and a value being supported.

Checks currency: whether a certificate or test report has expired, whether a verification cycle has lapsed, whether a supplier declaration predates a known change in the product. Detects decay, which no other control detects because nothing in the record changed.

The first four classes can be run inside a single system against a single record. The last three require context beyond the record, and they are the classes that most closely correspond to whether a published claim can be defended.

Common Mistake
Reporting validation pass rates as data quality

A validation suite reports the proportion of records that satisfy the rules that were written. Attributes with no rules, and dimensions no rule addresses, are invisible to it. A ninety eight percent pass rate frequently means the rule set is narrow rather than the data is good.

Data Quality vs Data Truth and Evidence

The third idea, and the one most often missing entirely, is whether the underlying claim is factually correct and sufficiently supported by authoritative evidence.

Consider a recycled content field on a single product.

Example
Four questions about one value

It exists. The field is populated. Completeness is satisfied.

It contains 42%. The value is a number, within the permitted range, expressed in the declared unit, and drawn from the permitted precision. Validity is satisfied and every format check passes.

It is attached to the correct product. The value is linked to the right item, at the right granularity, and does not contradict other values in the record or in other systems. Consistency and context are satisfied.

It may still be wrong. The supplier declaration behind the figure covers a different input batch, was issued before a change of recycled feedstock supplier, or reports recycled content on a basis the organisation has not adopted. Accuracy fails, and no structural, format, reference or cross-field check would ever have detected it.

The general principle: technically valid data is not automatically trustworthy data. Validity is a property of the record. Truth is a property of the world. Evidence is the bridge between them, and it is the only thing that allows an organisation to make the second claim rather than only the first.

This has three practical consequences.

Evidence has to be a first class object. A claim without a retrievable document supporting it is an assertion. The document needs an owner, a validity period, a scope, and a link to the specific attribute and product it supports, rather than sitting in a repository that a person once looked at.

Evidence has its own quality. A supplier declaration, a self-assessment, an accredited test report and a third party certification are not interchangeable. The required evidence standard should be set per attribute according to the consequence of being wrong, and recorded alongside the requirement.

Some values cannot be verified, and that must be visible. Where accuracy cannot be established, the honest position is to record the basis of the value, its uncertainty and its evidence gap. That is a defensible disclosure. Presenting an unverified figure with the same confidence as a tested one is not.

Best Practice
Separate the value, the basis and the evidence

Store what the value is, how it was arrived at (measured, tested, calculated, declared by a supplier, estimated) and what supports it, as three distinct facts. Programmes that store only the value can report completeness and validity, and can never report whether a claim is supported.

Product Data Quality and Digital Product Passports

Passport programmes are routinely described as having created a data quality problem. They have not. They have exposed one that already existed.

Product data historically lived in the systems that used it, in the form those systems needed, and was reconciled by people at the point of use. That arrangement tolerates fragmentation extremely well. A planning system and a design system can disagree indefinitely, because no process ever asks them the same question at the same time.

A passport asks exactly that question. Information from several systems must be assembled into one coherent, externally visible representation of a single product, at a single point in time, under a named accountability. Four weaknesses surface immediately.

Fragmentation. The same attribute exists in several systems with different values, and nobody had previously needed to decide which one was authoritative.

Missing ownership. Attributes that sit between functions, typically sustainability and regulatory data, turn out to have no accountable owner, only contributors.

Inconsistent definitions. Terms that everyone assumed were shared turn out to mean different things in different functions, so the values were never comparable.

Weak evidence chains. Claims that were adequate internally have no retrievable, in-date documentation behind them, which becomes apparent only when someone is asked to produce it.

None of these are caused by the passport. They are pre-existing conditions of a distributed product data estate, and the passport is simply the first process that requires them all to be resolved simultaneously. Recognising this matters politically as well as technically: programmes that treat the exposure as a passport problem tend to try to fix it in the publication layer, which is the one place it cannot be fixed.

Common Mistake
Fixing data in the publication layer

Correcting values on the way out produces a passport that disagrees with every internal system, cannot be reconciled, and has to be corrected again at the next publication. The passport should publish what the authoritative sources hold; defects belong upstream, where the value is owned.

Where Product Data Quality Problems Originate

Defects are rarely created where they are discovered. Six origins account for most of what a passport programme finds.

Creation without definition. An attribute is introduced for a project, populated by whoever needed it, and never given a written definition, unit or permitted values. Every later inconsistency traces back to this.

Migration and consolidation. Data moved between systems, or merged after an acquisition, carries mappings that were approximate at the time and are undocumented now. Values look native and are not.

Manual entry and spreadsheets. Information collected outside governed systems, then loaded in bulk, arrives without validation, without provenance and often without an accountable sender.

Integration mapping faults. A field mapped to the wrong target, a unit conversion applied twice or not at all, a truncation applied silently. These produce values that are valid and wrong, and they persist because nothing rejects them.

Change propagation gaps. A component substitution, a supplier switch or an engineering change is made in one system and not propagated to the derived values elsewhere. The original data was correct; the estate simply drifted apart.

Passive decay. Nothing changed in the record, but the world moved: a certificate expired, a regulation changed the required basis of a calculation, or a supplier altered a process. This is the category conventional data cleansing never finds.

Classify defects by origin, not by symptom

A backlog organised by attribute tells you what to fix once. A backlog organised by origin tells you what to stop causing. The second is what reduces the volume of the first.

Supplier Data Quality

A substantial share of the information a passport carries did not originate inside the organisation. Material composition, recycled content, substance declarations, certifications, origin and component specifications commonly arrive from suppliers, and frequently from suppliers several tiers away who have no direct relationship with the party that must publish.

This creates a specific problem: the organisation is accountable for information it did not create, cannot directly measure and often cannot independently verify. Six issues recur.

Supplier provided attributes arrive in the supplier’s structure, vocabulary and units, and must be mapped into the receiving organisation’s definitions. The mapping is where most silent defects are introduced.

Declarations are statements of fact made by the supplier. They carry the supplier’s accountability, not the receiving organisation’s assurance, and they should be recorded as declarations rather than promoted to verified values.

Certifications have a scope, an issuer, a validity period and a subject. A certificate covering a production site is not a certificate covering a specific batch, and treating the two as equivalent is a common and consequential error.

Supporting evidence is frequently the weakest link: a document referenced in an email, a report covering a superseded specification, or a claim with nothing behind it at all.

Missing information must be tracked as an open request with an owner and a due date, not as an empty field. Where a supplier cannot or will not provide an attribute, that fact is itself information the programme needs.

Conflicting information arises when a supplier’s declaration disagrees with the organisation’s own records or with another tier of the chain. Conflict resolution needs a defined rule, an owner and an audit trail, because whichever value wins becomes a published claim.

Two further mechanisms are needed to make supplier data usable. Supplier corrections must have a route: a defined way for a supplier to supersede a previous submission, with the previous version retained and any published passport re-evaluated. And approval workflows must distinguish receiving data from accepting it. Data can be received, validated, reviewed against evidence requirements, and only then approved for use in a published claim.

Common Mistake
Treating receipt as acceptance

Loading a supplier submission into a system does not make it trustworthy. Accepting supplier data is a decision, made by an accountable person against a defined evidence standard, and it should be recorded as such. Without that step the organisation has inherited the supplier’s uncertainty and published it under its own name.

The mechanics of turning supplier submissions into governed, evidenced, publishable values, including onboarding, data specifications, escalation and tiering, are a subject in their own right and are covered separately in a forthcoming article on how supplier data becomes trusted product data. The broader movement of information along the chain is covered in How Product Data Moves Through the Supply Chain.

Product Data Quality Across ERP, PLM, PIM and MDM

Quality behaves differently in each system because each holds a different kind of data for a different purpose. The system boundaries are set out in ERP vs PIM vs PLM; what follows is the quality view of the same landscape.

PLM holds engineering truth: specifications, bills of materials, materials and design revisions. Its characteristic quality risk is version drift, where the published value reflects a superseded revision, and granularity mismatch, where materials are recorded for engineering purposes at a level of detail that regulatory disclosure does not accept.

ERP holds operational and commercial reality: suppliers, sites, costs, logistics attributes and transactions. Its characteristic risk is that operationally sufficient values, nominal weights, approximate dimensions, generic origin, are treated as declarations of fact.

PIM holds customer facing content: descriptions, marketing attributes, images, channel specific variants. Its characteristic risk is that content optimised for persuasion is reused as a compliance statement, and that per channel variants create multiple conflicting versions of the same attribute.

MDM, where present, holds the reconciled golden record. Its characteristic risk is false confidence: a survivorship rule silently selects one of several conflicting values, and the result is presented as authoritative without any evidence that the surviving value is the correct one.

Two conclusions follow. First, quality rules must be attached to attributes and applied wherever the attribute lives, rather than being implemented once inside whichever system the programme happens to start in. Second, cross-system reconciliation is not an optional refinement; in an estate of four systems it is the only control capable of detecting the most consequential class of defect.

Example
One product, five sources

A manufacturer preparing a passport for a single product finds material composition in PLM, supplier and manufacturing site in ERP, customer facing attributes in PIM, recycled content declarations in a supplier portal, and certification documents in a document repository.

Applying the six dimensions to that assembly finds different problems in each source. Completeness shows the substance breakdown in PLM is held at material family level, not at the substance level disclosure requires. Accuracy shows the recycled content declaration in the portal was issued against a superseded input specification. Consistency shows the manufacturing site in ERP and the site named on the certificate in the repository are different facilities. Validity shows PIM holds the product weight as free text with a unit inside the string. Timeliness shows the certificate expired four months ago. Traceability shows no record of who accepted the supplier declaration or on what basis.

Each of those is a different defect, with a different owner and a different remedy. A single quality score for the product would have reported one number and directed the team nowhere. And none of them is fixed by the passport: the passport is where they became visible.

Note what the example does not imply. The passport is not the place to store the corrected values. Each attribute remains owned by its system of record, and the passport publishes from those sources. A passport that becomes the master database for values it did not originate is a new silo with an external audience.

Measuring Product Data Quality

Measurement should be attribute level and domain level, never a single organisational number. The practical instrument is a scorecard that records, for each attribute in scope, whether it is required for the intended use and how it performs against each dimension.

An illustrative scorecard for a small set of passport relevant attributes:

AttributeRequired?Complete?Valid?Current?Source known?Evidence available?
GTINYesYesYesYesYesNot applicable
Material compositionYesPartialYesYesYesPartial
Recycled contentYesYesYesNoYesNo
Manufacturing locationYesYesYesYesYesYes
Repair informationYesNon/an/aNoNo

The scorecard is illustrative, not a template to be adopted verbatim. Its value lies in four properties.

It is per attribute. Each row is separately actionable and separately owned. Recycled content here is complete and valid and simultaneously the highest risk row on the sheet, because it is neither current nor evidenced. An aggregate score would have concealed that.

It distinguishes not applicable from failing. Evidence is not applicable to a GTIN in the way it is to a sustainability claim. Marking that explicitly prevents the scorecard from manufacturing false defects.

It separates the dimensions. Reading across a row shows which control failed, which determines who fixes it. Completeness gaps go to the data owner, evidence gaps go to procurement or sustainability, validity gaps go to the point of entry.

It carries a required flag. Attributes that are not required for the intended use do not generate work. Scope discipline is what keeps a quality programme finishable.

On scoring, two cautions. There is no universal tieback scoring formula, deliberately. Averaging six dimensions into one percentage implies that a validity failure and an evidence failure are interchangeable, and they are not. More importantly, critical regulatory attributes warrant hard pass or fail thresholds rather than averages: an attribute required by a delegated act either meets its requirement or it does not, and a portfolio at ninety four percent on such an attribute is not ninety four percent compliant, it is non-compliant for six percent of products. Averages are appropriate for prioritising remediation effort; they are inappropriate as a statement of readiness.

Useful operational measures, reported per attribute and per domain, include population rate against the required set, validation failure rate by rule class, cross-system disagreement rate, proportion of values with a recorded last verified date within the required interval, proportion of claims with in-date retrievable evidence, and open defect age. Each of those points at a specific control.

Best Practice
Measure the attributes you are about to publish first

Scoping measurement to the attribute set required by the nearest obligation produces a scorecard that can be completed, acted on and finished. Portfolio wide measurement across all attributes produces a dashboard that is never green and never prioritised.

Improving Product Data Quality

Improvement is a sequence, and the order matters more than the tooling.

Define before you measure. An attribute without an agreed definition, unit and permitted values cannot be assessed, because there is no standard to assess against. Definition work is usually the cheapest and highest yield activity in the programme.

Establish ownership per attribute. Every defect needs somewhere to go. Unowned attributes generate reports that nobody actions, which trains the organisation to ignore quality reporting generally.

Prioritise by consequence. Rank attributes by the cost of being wrong: regulated disclosures and safety relevant values first, commercial attributes next, descriptive content last. This is the opposite of prioritising by defect count, which pulls effort towards whatever is most numerous.

Move controls upstream. A rule applied at the point of entry prevents a defect; the same rule applied at publication rejects work already done. Every control that can be moved earlier should be.

Remediate in campaigns, not continuously. Fixing a defined attribute across a defined product set, with an owner and an end date, completes. Open ended cleansing does not.

Close the loop with the origin. Each remediation should end with a change to the process, integration or supplier specification that produced the defect. Without that step the same backlog regenerates.

Re-verify on a cycle. Set a verification interval per attribute according to volatility and consequence, record last verified dates, and treat a lapsed interval as a defect in its own right.

Report defects, not dashboards. Quality reporting that produces assignable items with owners and due dates changes behaviour. Quality reporting that produces a trend line does not.

A realistic expectation is worth stating plainly: perfect product data is not achievable, and pursuing it is a poor use of a programme’s credibility. The objective is data that is demonstrably fit for its intended use, with known gaps that are recorded, owned and disclosed where disclosure is appropriate.

Common Misconceptions

Common Mistake
Complete data is high-quality data

Completeness is one dimension of six. A fully populated record can be inaccurate, internally inconsistent, years out of date and entirely unevidenced. Population rate is the easiest metric to produce, which is why it is so often mistaken for the whole picture.

Common Mistake
If data passes validation, it must be correct

Validation tests the record against rules. Correctness is a relationship between the record and the world, and no rule written inside the system can inspect the world. A valid value that corresponds to nothing real will pass every check that exists.

Common Mistake
The Digital Product Passport should clean the data

A passport is a publication mechanism. Correcting values at the point of publication creates a representation that no internal system agrees with, cannot be reconciled, and must be corrected again at every subsequent publication. Defects are fixed where the attribute is owned.

Common Mistake
One enterprise system should become the source for everything

Different attributes have legitimately different authoritative sources: engineering specifications in PLM, supplier and site data in ERP, customer facing content in PIM. Consolidating everything into one system replaces a distribution problem with a migration problem and does not by itself improve any of the six dimensions.

Common Mistake
Data quality is an IT responsibility

IT operates the systems, the integrations and the validation engine. It cannot decide what an attribute means, whether a value is accurate, what evidence standard applies, or whether a supplier declaration should be accepted. Those are business decisions, and quality stalls wherever they are delegated to a technical function.

Common Mistake
A single data-quality score tells us whether the product is compliant

A composite score averages across dimensions and attributes with different consequences, and can be high while a legally required attribute is missing or unevidenced. Compliance is assessed per requirement, on a pass or fail basis, not against an aggregate.

Frequently Asked Questions

Is product data quality the same as data governance?
No. Governance establishes accountability, definitions, policies and decision rights. Quality measures whether the resulting data meets requirements. Governance decides what should be true; quality reports how far reality differs from that, and hands the difference back to governance as assignable work.

Can data quality be fully automated?
Validity, and much of completeness and internal consistency, can be automated. Accuracy and evidence sufficiency cannot, because they require reference to something outside the data. Automation should be used to remove all the mechanical failures so that human attention is available for the two dimensions that need it.

What quality level is good enough?
It depends on the attribute and the use. A defensible answer sets a threshold per attribute derived from the consequence of being wrong, and treats regulated attributes as pass or fail. There is no credible universal target.

Should we fix everything before publishing a passport?
No, and the attempt usually prevents publication indefinitely. Fix the attributes the obligation requires, to the standard it requires, and record known gaps for the rest with owners and dates.

Who should own product data quality?
Accountability sits with the business owner of each attribute domain, operating through stewards, with a governance function maintaining the standard and reporting. A central quality team without attribute owners produces measurement without remediation.

How does this relate to ISO data quality standards?
Formal standards, notably the ISO 8000 series and ISO/IEC 25012, define data quality vocabulary and characteristics and are the appropriate reference for a formal quality management approach. The six dimension model in this article is a tieback educational framework intended to make the subject tractable for passport programmes; it is not a standard and does not substitute for one.

Does higher data quality guarantee regulatory compliance?
No. Compliance requires the right attributes, to the right definition, with the required evidence, published in the required form. High quality data makes compliance achievable; it does not establish it.

How often should product data quality be measured?
Continuously for automated dimensions, since the cost is near zero once the rules exist. Periodically and by exception for accuracy and evidence, on a cycle set per attribute by volatility and consequence.

Key Takeaways

Key Takeaways

Definitions of record for the terms used above live in the glossary.

References

About This Article

tieback Knowledge is a continuously maintained reference library covering Digital Product Passports, product traceability, product compliance and related regulations. Articles are reviewed regularly as legislation, standards and implementation guidance evolve.