It can take years to build and only a moment to lose. AI can be a powerful growth engine for the enterprise, but when handled poorly, it becomes a brand risk. Data quality management is no longer just a technical task. It is enterprise risk management.
ISO/IEC 5259-2: The Data Quality Standard for Reliable AI
Enterprise trust takes years to build and only a moment to lose.
💡
ISO/IEC 5259-2, the focus of this article, establishes that benchmark as an international standard. Pebblous assesses and improves data quality based on ISO/IEC 5259-2 for enterprise organizations across telecommunications, manufacturing, mobility, and other industries, including LG Uplus, one of Korea's three major telecommunications companies.
In this article, we will walk through what ISO/IEC 5259-2 is, how it applies to real-world data, and how to interpret measurement results.
What Makes ISO/IEC 5259-2 Different?
At Pebblous, we place particular emphasis on ISO/IEC 5259-2 among the many data quality standards. Why does this standard matter so much?
Put simply, it's the data quality standard best suited to where AI is headed.
To explain why, let's compare ISO/IEC 25012, ISO/IEC 5259, and ISO/IEC 5259-2.
ISO/IEC 25012
Introduced around 2015, this standard was designed to manage the quality of structured data stored in databases. To date, many government agencies and enterprises have used this standard as a basis for data quality management.
🤖
However, in the AI era, ISO/IEC 25012 alone is not enough. AI and databases have different requirements. AI does not just need data that is correctly formatted. It needs data that is suitable for learning.
Consider a dataset. The file names are valid, the image resolution is uniform across files, and there are no obvious missing records. The labeling is also well executed. Under ISO/IEC 25012, this dataset would likely be considered high quality.
But what if 90% of the samples come from only one specific condition, or if all images were captured from the same angle and environment? The data format may be flawless, but AI performance will suffer.
ISO/IEC 5259
This is why ISO/IEC 5259 was developed. ISO/IEC 5259 builds on ISO/IEC 25012 and extends it for AI. It also defines additional data quality characteristics essential for AI and machine learning, such as diversity, representativeness, and similarity.
👏
The ISO/IEC 5259 series evaluates content rather than format, distribution rather than raw counts, and diversity rather than mere existence. Only with this standard can organizations accurately assess data density and AI-specific quality indicators.
Pebblous leverages cutting-edge research on the ISO/IEC 5259 series to drive real-world impact. To dig deeper into ISO/IEC 5259 measurement criteria, check out our detailed guide to ISO/IEC 5259-2 and our ISO/IEC 5259 series information hub.
ISO/IEC 5259-2
As described above, ISO/IEC 5259 is an international standards series for managing the quality of AI training data. Within this series, ISO/IEC 5259-2 focuses on how data quality should be measured and evaluated in practice.
ISO/IEC 5259-2 Measurement Criteria
From this point on, we'll focus specifically on ISO/IEC 5259-2. Its measurement items are broadly divided into four groups.
1. Intrinsic Data Quality Characteristics
These characteristics assess the fundamental properties of the data itself. They evaluate whether the content of the data is valid based on five criteria: accuracy, completeness, consistency, credibility, and currentness.
2. Intrinsic and System-Dependent Data Quality Characteristics
Data quality can change depending on the system that stores and processes the data. In other words, this category does not evaluate the properties of data in isolation. It evaluates the quality created by the combination of the data and the system that stores and processes it. It examines how well data is managed within a system, including accessibility, compliance, and traceability.
3. System-Dependent Data Quality Characteristics
These items evaluate the performance of the IT systems and infrastructure that hold the data. Even high-quality data is difficult to use in practice if the system is unstable. This category assesses whether data is available when needed, portable to other environments, and recoverable in the event of failure.
4. Additional Data Quality Characteristics for Analytics and ML
These ML-specific data quality characteristics consist of 26 items across 10 subcategories, including auditability, balance, diversity, validity, identifiability, suitability, representativeness, similarity, and timeliness.
💡
This is what differentiates ISO/IEC 5259-2 from conventional standards that focus mainly on defects in the data itself. The core value of ISO/IEC 5259-2 is that it's designed from the data user's perspective, asking how useful the data is for a specific AI project's purpose. It is the only AI-specific quality standard that enables quantitative measurement of balance, representativeness, diversity, and similarity, all of which directly influence unbiased model training.
For detailed measurement criteria for each item, including QM codes, see the Pebblous blog.
How Does Pebblous Apply ISO/IEC 5259-2 to Real Assessments?
Pebblous applies ISO/IEC 5259-2 to Data Clinic, our data quality management solution, to assess real datasets. We organize the quality characteristics of ISO/IEC 5259-2 by assessment level and analyze data quality accordingly.
Level I : Foundational Quality Assessment
This level checks the basic quality condition of a dataset. Think of it as a baseline health check for your dataset.
Pebblous Assessment Item | ISO/IEC 5259-2 Measurement Item | Details |
|---|---|---|
Missing Values | Completeness | Checks whether data values, records, and labels are missing. |
Consistency Check | Consistency | Checks whether data formats are valid and whether labeling errors are present. |
Class Balance | Balance | Identifies whether data is excessively concentrated in specific classes and diagnoses the potential for training bias. |
Statistical Analysis | Accuracy | Analyzes the distribution and attributes of the data to evaluate how accurately it reflects real-world phenomena. |
Level II·III : AI-Centric Quality Assessment
However, AI data quality cannot be determined by missing values or formatting errors alone. To build data that AI can learn from effectively, Data Clinic uses DataLens. It projects data into an embedding space and analyzes quality in the way AI perceives data.
Think of an 'embedding space' as a massive digital map where similar data points are clustered together like neighboring cities. By analyzing this map, AI doesn't just read data; it visualizes the entire landscape, allowing us to spot 'deserts' where data is missing, or 'congested zones' where redundant data is piled up.
Pebblous Assessment Item | ISO/IEC 5259-2 Measurement Item | Details |
|---|---|---|
Density Analysis | Similarity | Analyzes data density in the embedding space to diagnose excessive concentration of duplicate or similar data. |
Manifold Shape Analysis | Diversity, Representativeness, Balance | Identifies underrepresented regions and edge cases to analyze whether the data distribution is biased. |
Intrinsic Dimension Analysis | Similarity, independence-related characteristics | Measures how much diverse information the data actually contains and evaluates information complexity. |
Why ISO/IEC 5259-2 Is Difficult to Apply in Practice, Even When You Understand It
Even when organizations understand ISO/IEC 5259-2, improving data quality to the desired level is often difficult in practice. Pebblous has already addressed many of these challenges. Based on that experience, we will explain why many teams struggle and how those challenges can be solved.
1. Data Quality Cannot Be Solved by Standards Alone.
Standards are important, but standards are not everything. (quote)
In real-world operations, a single standard can't answer every question that comes up. Conditions vary widely by industry, and so does the complexity of each problem. Standards alone can't keep pace with these realities.
Organizations must also comply with AI regulations and governance requirements. Yet many ethical criteria remain insufficiently defined.
In this context, Pebblous’ data quality management solution, Data Clinic, does not rely solely on ISO/IEC 5259-2.
💡
Pebblous’ Data Clinic 2.0 uses Agentic AI technology to measure and improve data quality more comprehensively from a neuro-symbolic perspective. An AI agent trained on 285 expert documents related to data quality understands complex standards and diverse domains, then recommends and performs the most appropriate quality assessment.
The agent's training scope consists of the following five areas. These bodies of knowledge work together to improve data quality.
Data quality standards, including ISO 5259 (28.84%)
Regulations and governance frameworks, such as the EU AI Act, GDPR, NIST AI RMF (U.S. AI Risk Management Framework), and the Executive Order on AI (45.09%)
Domain applications in robotics, manufacturing, and public safety (14.06%)
Foundations of data science based on recent academic conference papers, including ICML (9.89%)
Pebblous internal technical documents (2.11%)
2. Data Is Viewed Only from a Human Perspective, Not an AI Perspective.
Artificial intelligence may appear highly capable, but its operating principles are fundamentally different from those of the human brain.
Humans understand data intuitively, but AI does not understand data in the same way. AI learns patterns in data. As a result, data that appears normal to humans can become biased training data for AI. (call out)
That is why organizations need to shift perspective. Data should be examined not from the human point of view, but from the AI point of view.
Two Solutions Developed by Pebblous
After extensive research, Pebblous identified two solutions.
1. Neural Network-Based DataLens
Rather than treating AI data as a simple table, this technology analyzes how AI actually perceives the data, how data points relate to one another, and what the statistical distribution looks like.
Data Clinic uses two types of DataLens: general-purpose and customized. Among them, customized DataLens is developed to reflect the characteristics of each domain.
2. Data Imaging
Data Imaging is a technology that makes the data seen by AI visible to humans. Data is high-dimensional, often containing thousands of features. Data Imaging visualizes it in 2D or 3D space, allowing teams to see at a glance where data is concentrated, which areas are empty, and where outliers exist. Pebblous holds a U.S. patent for this approach, which combines ISO 5259 quality criteria with Data Imaging to assess data quality, and the technology is proprietary to Pebblous.
A Real Data Quality Assessment Process Using Data Clinic
Let's look at how Data Clinic, which combines all the technologies described above, assesses data quality. We'll use a diagnostic report on a recyclable waste dataset as an example.
Below are two issues Data Clinic identified by projecting images into an embedding space and analyzing them.
Class Imbalance
According to ISO/IEC 5259-2, this dataset showed a clear class imbalance issue. When comparing the video_tape class with the filament class, the counts were 646 images and 33 images, respectively. In class analysis, a ratio exceeding 10:1 between the largest and smallest classes is generally considered a severe imbalance. In this case, the ratio was 19.6:1.
We also analyzed the degree of imbalance using L2 (general DataLens) and L3 (domain-specific DataLens). L2 density values ranged from 0.06 to 0.18, while L3 values ranged from 0.5 to 1.5. The units differ because L2 is based on general-purpose embeddings with 1,280 dimensions, whereas L3 uses optimized embeddings with 32 dimensions through fine-tuning for the recycling domain. As a result, L3 captures recycling-related features more sensitively.
Looking first at the L2 chart, we can see that video_tape samples are clustered with similar samples in the embedding space.
Through the L3 domain-specific lens, this pattern becomes even clearer. When analyzed from the recycling-domain perspective, video_tape shows a much larger density gap compared with other classes. This clearly reveals that video_tape accounts for 37.8% of the dataset and that similar patterns repeat within that class.
You can find the full explanation of the diagnostic report at the link below.
Where Pebblous Is Headed Next
Pebblous is not satisfied with today's data quality standards. We anticipate the quality standards of the future.
When Pebblous was founded in 2021 and first began developing Data Clinic, its data quality management solution, ISO 5259 did not yet exist. Because there was no adequate quality standard tailored to AI training data, Pebblous developed its own criteria.
🙂
Today, a significant portion of what Pebblous anticipated is reflected in ISO 5259. Pebblous aims to do more than help your AI data meet standards. We help organizations understand why those standards exist and fundamentally improve their data. We also aim to provide solutions that remain safe and reliable not only today, but also in the future.
1. Under review for designation as a KOLAS-accredited body
KOLAS is a Korean government accreditation scheme that evaluates and officially accredits the technical competence of testing, inspection, and calibration bodies under Korea's Framework Act on National Standards. It can be understood as Korea's counterpart to ANAB in the United States.
Pebblous is currently undergoing review to be recognized by KOLAS as a trusted organization authorized to issue certificates for the ISO/IEC 5259-2 standard. We have successfully completed the on-site assessment and expect to obtain the accreditation body designation in August.
2. Implementing automated assessment for all items in the AADS project
In the AADS (Agentic AI Data Scientist) project currently being carried out by the Ministry of Science and ICT, we are developing technology that automatically assesses data quality based on the 5259 standard. We successfully automated part of the process in Phase 1, and in Phase 2, all 5259 items will be assessed automatically. Neuro-symbolic technology and AI Agent technology are at the core of this work.
Are you curious whether your AI data currently meets the criteria above?
Data Clinic assesses your dataset against ISO/IEC 5259-2 and delivers a report that pinpoints exactly where the issues are. Using Pebblous' proprietary technology, we help you build a data quality foundation that stays resilient as AI regulations and standards continue to evolve.