← Back to articles

Data & Research

I Stopped Cleaning Data Before Understanding It

I used to open a messy dataset and immediately start fixing things.

Remove duplicates. Correct spellings. Fill blanks. Standardise formats.

It felt productive.

Then I realised I was sometimes fixing problems I hadn’t even properly understood.

Now, before touching a dataset, I do a quick audit.

How many duplicates are there? How many missing values? Which formats are inconsistent? Are different categories actually describing the same thing?

That small step changes everything.

A few inconsistent entries might be a five-minute cleanup. But if hundreds of records are inconsistent because different people have been collecting information differently, the problem isn’t really the spreadsheet.

It’s the process behind it.

I also stopped assuming that manual checking is automatically more accurate. Going through hundreds of rows by hand at 2am doesn’t make me careful. It makes me tired.

If a formula, validation rule, or script can reliably identify the problem, I’d rather automate that part and use my attention where it’s actually needed.

The goal of data cleaning isn’t to make a spreadsheet look perfect.

It’s to make the information trustworthy enough to use.

Enjoyed the article?

Read more of my notes and experiences.

More articles →