Why original data stays untouched
Cleaning is an argument about the evidence. The source file remains the evidence.
The original dataset is the record of what was collected or received. It is evidence. Cleaning, recoding, and variable construction are arguments about that evidence. Arguments get written in working files. They do not overwrite the source.
Why the source has to survive
If the original file is edited in place, nobody can say what arrived. A later dispute about an outlier, a duplicate, or a skipped item cannot be resolved. Audit, supervision, and reanalysis all depend on a file that still looks like the field or the sender.
The practice is simple. Store the file as received. Give it a name that says it is original. Restrict who can write to it. Do every change in a new file or a script that reads the original and writes a working copy.
Changes are methods statements
Document the changes. A note that age 220 was set to missing because it is outside a plausible range is a methods statement. A silent overwrite of 220 to 22 is a fabrication risk, even if the intent was honest.
Scripts are better than pointing and clicking for this reason. A script can be rerun. A manual edit in a spreadsheet cannot. If a spreadsheet must be used, keep the original sheet frozen and work on a copy.
Version the working files. Analysis ready is a destination, not a licence to forget the steps. A codebook and a short decision log should travel with the analytical file.
A documented correction is a sign of control. An original that has already been tidied is a sign that control was lost.
Custody starts at receipt
This rule is the same whether the data came from AcadStat field teams, a partner organisation, a ministry extract, or a public archive. Custody starts at receipt.
Researchers sometimes fear that leaving errors visible will embarrass the study. The opposite is true. The original data stay untouched so that every later claim can be traced. That is not bureaucracy. It is how evidence remains evidence.