Data is powerful, but it needs to be used carefully. One of the big challenges with using data is ensuring that it represents everyone fairly and doesn’t contain biases that can lead to wrong or unfair conclusions.
You might be wondering what does bias mean. So, let’s start with this.
Bias means having a preference or unfair opinion about something or someone without knowing all the facts. It’s like when you think something is true just because of one thing you heard or saw, rather than understanding the whole story. For example, if you think all cats are mean because one cat scratched you, that’s a bias.
Bias can make us treat people or things unfairly because we’re not seeing the full picture.
On the other hand, we have the concept of representation, which means making sure different kinds of people, are included and shown in a fair way. Going back to the example we used previously, about cats, if you only hear stories, for example, where cats are always shown as mean and scary, you might think all cats are like that. But if there are also stories where cats are shown as playful and loving, it helps you see a more complete picture.
For this specific case, good representation helps reduce bias because you get to understand and appreciate the different ways cats can behave.
Let’s explore what representation and bias mean in the context of data and how they can affect data interpretation.
Representation in data means making sure that the data includes all the different kinds of people or things it’s supposed to represent. For example, if you’re collecting data about favourite school subjects, you want to ask students from different grades, schools, and backgrounds. If you only ask students from one school, your data won’t represent all students accurately. This is important because decisions based on this data, like which subjects to focus on more in school, might not be the best for everyone.
Bias in data happens when the data collected favours one group over another unfairly. This can happen for many reasons. Sometimes it’s because the data collectors have their own opinions and assumptions that affect their work. Other times, it’s because the data doesn’t cover everyone equally. For example, if you’re collecting data on the best sports, but you only ask boys and not girls, your data will be biased. This kind of bias can lead to conclusions that don’t truly reflect what everyone thinks or feels.
Now, let’s look into how bias can affect our lives. Let’s start with AI systems used for facial recognition. These systems learn to recognise faces by studying lots of images. If most of the images they study are of people from one ethnicity only, the AI system might not be very good at recognising faces from other ethnicities. This is a problem because it means the AI is biased. It might make mistakes, like not recognising someone or confusing one person with another, which can have serious consequences, especially if this technology is used in security.
Another example is in healthcare. Imagine an AI system designed to help doctors diagnose illnesses. If the data used to train this system mostly comes from people of a certain age, the AI might not perform well for people outside that group. This can lead to misdiagnoses and unfair treatment. For instance, if an AI system has mostly been trained with data from adults, it might not be as effective in diagnosing children’s illnesses.
Representation and bias can also affect other simpler aspects of our lives, like recommendations on streaming services (like Netflix). If the data used to suggest movies or series is mostly based on what older adults watch, younger audiences might not get good recommendations that match their interests. This can make the service less enjoyable for them.
So, now you might be guessing, how do we ensure that data is well-represented and free from bias?
Well, it starts with being mindful about how data is collected. It’s important to include a diverse range of participants and to check for any patterns that might show bias.
When building AI systems, for example, it’s crucial to use diverse and comprehensive data sets. This helps AI learn accurately about all the different groups of people it needs to serve.
One way to tackle bias is by continuously testing and updating the data and systems. Just like we update our apps to fix bugs, we need to update our data and AI models to correct biases.
It’s also essential to have diverse teams working on data collection and AI development. When people from different backgrounds work together, they can spot biases that others might miss.
In conclusion, representation and bias are critical issues in data interpretation. Ensuring that data represents everyone fairly and is free from bias helps in making accurate and fair decisions. By understanding and working against bias, we can create a more inclusive and fairer digital world.
The next chapter is Data Visualisation.
I’m on YouTube now!
Check out Accessible BI for practical Power BI tutorials and tips on making data accessible to everyone. Subscribe here: Accessible BI YouTube Channel





[…] The next chapter is Bias and representation: why should we worry about them when working with data? […]