Can We Trust the Data? Understanding Data Quality Before Gaining Insights

Introduction: Before We Use Data, Should We Trust It?
A dashboard reports that customer satisfaction is 92%.
A machine-learning model achieves 95% accuracy.
A university reports that 87% of students are satisfied with a programme.
The numbers look precise. The charts may look professional. The analysis may appear convincing.
But before asking:
“What does the data tell us?”
Perhaps we should first ask:
“Can we trust the data?”
This is an important question because data is not automatically reliable simply because it is presented as a number, displayed in a dashboard or processed by sophisticated technology.
Data may contain errors, missing information, inconsistencies, duplication or bias. It may also be collected for one purpose and later used for another.
In the previous article, Why Data Literacy Is Becoming the New Digital Literacy, we explored why data literacy matters. In this article, we take the next step: understanding data quality and trustworthiness.
For professionals who want to strengthen their ability to interpret, analyse and use data more effectively, the Data Analytics Powered by AI course from Digital Regenesys provides a practical pathway for developing data analytics capabilities that can support informed decision-making.
This does not mean that data must be perfect before it can be useful. Rather, we need to understand whether the data is good enough and appropriate for the purpose for which it is being used.
This article outlines how to assess data quality before using data to generate insights, train artificial intelligence or machine-learning models, or support organisational decisions. It explains six practical checks for determining whether data can be trusted: understanding its source and provenance, checking accuracy, assessing completeness, testing consistency, considering potential bias and determining whether the data is fit for purpose. The article also explains how unreliable or unrepresentative data can affect analytics and AI model performance, why context matters when interpreting statistics and dashboards, and how organisations and data professionals can apply stronger data-quality practices to support more responsible, evidence-based decision-making in South Africa and other data-driven environments.
What Makes Data Trustworthy?
Consider two datasets containing customer information.
The first contains:
- Complete customer records
- Consistent definitions
- Accurate transaction information
- A clear collection method
- Regularly updated records
The second contains:
- Missing customer details
- Duplicated records
- Outdated information
- Inconsistent formats
- Unclear information about how the data was collected
Both may contain thousands of records.
But more data does not necessarily mean better data.
The quality of data depends on characteristics such as accuracy, completeness, consistency, timeliness and relevance.
A useful question is therefore not:
“How much data do we have?”
But:
“Is this data fit for the purpose?”
The 6 Checks Before You Trust Data
Whenever you encounter a dataset, statistic, report, dashboard or AI-generated insight, consider these six checks.
1. Where Did the Data Come From?
The origin of data influences how confidently it can be interpreted.
Data may come from surveys, transaction systems, sensors, experiments, social media, administrative records, websites or research studies.
Consider the statement:
“90% of customers are satisfied with our service.”
Before accepting it, ask:
- Who collected the data?
- How was it collected?
- How many customers participated?
- When was it collected?
- Were all customers given an opportunity to respond?
A company survey and an independent research study may produce different results because they may use different methods, populations and questions.
Understanding data provenance — where data originated and how it was collected and transformed — is therefore an important part of responsible data use.
2. Is the Data Accurate?
Accuracy refers to whether the recorded information correctly represents what it is intended to measure.
A simple example is a customer database containing an incorrect age, address or transaction value.
Such errors may appear minor, but they can influence analysis and decisions when they occur repeatedly.

Accuracy can also become more complicated when data is collected automatically.
A sensor may record an incorrect reading. A form may allow invalid entries. A system may convert measurements incorrectly.
The presence of data in a database does not guarantee that the data accurately represents reality.
3. Is the Data Complete?
Missing information can change the story told by a dataset.
Imagine a university analyses student engagement using online-platform activity. Some students show very little activity.
Does this necessarily mean they are disengaged?
Not necessarily.
They may access learning materials offline, experience connectivity difficulties, use alternative resources or simply interact with the learning environment in ways that are not captured by the system.
This illustrates an important principle:
What is not recorded can sometimes be as important as what is recorded.
Completeness therefore requires us to consider both missing values within the dataset and missing aspects of the real-world situation.
4. Is the Data Consistent?
Data can come from different systems, people and time periods. As a result, the same information may be recorded in different ways.
For example, one system may record a South African province as:
- Gauteng
- GP
- Gauteng Province
A person may recognise these as the same location, but an analytical system may treat them as different categories unless the data is standardised.
Consistency becomes particularly important when datasets from multiple sources are combined.
Without appropriate data cleaning and standardisation, apparently simple analysis can produce misleading results.
5. Could the Data Be Biased?
A dataset can be technically accurate and still provide a distorted picture if certain groups or situations are under-represented.
Consider a survey about students’ experiences that receives responses mainly from students who are highly engaged with the institution.
Can the results automatically represent all students?
Not necessarily.
This is why sampling and representation matter.
Bias can enter through the way data is collected, who participates, which variables are selected, how questions are designed and how information is subsequently interpreted.
The key question is:
“Whose experience is represented — and whose might be missing?”
This question is particularly important when data is used to gain insights and support decisions affecting people.
6. Is the Data Fit for Purpose?
This may be the most important question of all.
A dataset can be accurate, complete and consistent and still be unsuitable for a particular purpose.
For example, historical sales data may be useful for understanding previous purchasing patterns. It may not, by itself, be sufficient for predicting future demand if market conditions have changed significantly.
Similarly, social-media data may provide useful insights into online opinions, but it may not represent the views of an entire population.
Therefore, data quality cannot be judged independently of the question being asked.
Good data is not simply data that is accurate. It is data that is appropriate for the purpose.
A Closer Look: Data Quality in AI and Machine Learning
For readers working with data science, data analytics, AI or machine learning, data quality becomes especially important because models learn from the data provided to them.
Missing values, outliers, sampling bias, poorly defined features or data leakage can affect analysis and model performance.
Therefore, a technically sophisticated model cannot compensate for unsuitable or unreliable data.
Consider a machine-learning model reported to have 95% accuracy.
Before trusting this result, ask:
- What data was used for testing?
- Was the test data representative?
- How was accuracy measured?
- Were some classes under- or over-represented?
- Is accuracy the right metric for this problem?
A 95% accuracy score may sound impressive, but a performance measure needs context.
Reliable AI begins with reliable data.
The same principle applies outside machine learning.
A percentage, average, prediction or visualisation should not be interpreted independently of how it was produced and what it is intended to represent.
Turn Data Skills Into Career Growth
Learn Python, machine learning & Power BI with hands-on, AI-enabled projects.
Explore ProgrammeTry This: Can You Trust This Statistic?
Consider the statement:
“90% of students prefer online learning.”
Before accepting it, apply the six checks:
- Source: Who collected the data and how?
- Accuracy: Were the questions clearly designed?
- Completeness: How many students responded?
- Consistency: Was “online learning” defined consistently?
- Bias: Were particular groups under-represented?
- Fitness for purpose: Can the findings reasonably represent all students?
The aim is not to reject the statistic, but to understand the evidence before using it to make decisions.
Good data thinking begins with good questions.
Data Quality Is a Shared Responsibility
Data quality is not only a technical concern.

It can be influenced at every stage of the data lifecycle — from collection and entry to cleaning, analysis, modelling and interpretation.
Data professionals play an important role, but so do the people who collect, manage and use data.
Good data is everyone’s responsibility.
For professionals who want to build stronger practical capabilities in analysing, interpreting and communicating data, the Data Analytics Powered by AI course from Digital Regenesys offers a pathway for developing skills that can support data-informed business decisions.
Conclusion: Trust, but Verify
Data can be powerful, but more data does not necessarily mean better evidence.
Before trusting a number, dataset, dashboard, prediction or AI-generated insight, consider whether the data is reliable and fit for purpose.
Ask six simple questions:
- Where did the data come from?
- Is it accurate?
- Is it complete?
- Is it consistent?
- Could it be biased?
- Is it fit for purpose?
These questions complement technical skills in data analysis, data science and machine learning and support more informed decisions.
Before asking what the data tells us, ask whether the data deserves our trust.
Practical Recommendations
Before using or acting on data:
- Check its source and quality.
- Look for missing information, errors and inconsistencies.
- Consider bias and representation.
- Check whether the data is fit for purpose.
- For AI and machine-learning applications, evaluate the data as carefully as the model.
Good decisions begin with data that is fit for purpose.
Next Data Science Batch Starts Soon
Join 500,000+ alumni. Real projects, AI workflows, IITPSA-accredited certificate.
Enrol NowReferences
- Gartner (2024). Data Literacy. Gartner Research.
- IIEP-UNESCO (2026). Data Literacy in the Age of Generative Artificial Intelligence. International Institute for Educational Planning, UNESCO.
- OECD (2024). Do Adults Have the Skills They Need to Thrive in a Changing World? The Relevance of Information-Processing Skills in Rapidly Changing Societies. Paris: OECD Publishing.
- OECD and European Commission (2026). Empowering Learners for the Age of AI: An AI Literacy Framework for Primary and Secondary Education. Paris: OECD Publishing. DOI: 10.1787/65cd27d4-en.
- Ridsdale, C., Rothwell, J., Smit, M., Ali-Hassan, H., Bliemel, M., Irvine, D., Kelley, D., Matwin, S. and Wuetherick, B. (2015). Strategies and Best Practices for Data Literacy Education: Knowledge Synthesis Report. Halifax: Dalhousie University.
- World Economic Forum (2025). The Future of Jobs Report 2025. Geneva: World Economic Forum.
Last Updated: 25 September 2026