The Census Is a Privacy Nightmare. I Also Can’t Wait to Get the Data.
Australia wants to digitise everything. It might be worth asking whether everything actually needs to be digital.
So the Census happened last night, and as I somewhat expected, my experience with the website was an absolute cesspit of frustration. Things were slow, pages were struggling, and I spent a decent portion of the process wondering whether clicking a button again would help or whether I was about to accidentally submit myself to the Australian Bureau of Statistics seventeen times. By the end of it, I was a very frustrated boy.
It did get me thinking about something much bigger than the Census itself though: why are we so obsessed with digitising every single piece of government infrastructure? There seems to be this automatic assumption now that digital means better. Digital means faster, cheaper, more efficient and more modern. Except sometimes it doesn’t. Sometimes we’ve taken something incredibly simple, built an enormous technological system around it, spent an obscene amount of money doing so, introduced cybersecurity risks that didn’t previously exist and somehow made the experience more annoying.
The Census is probably one of the best examples of this because it operates at a completely ridiculous scale. Australia has roughly 27 million people, and the entire purpose of the Census is to create a statistical snapshot of basically all of us. Obviously that doesn’t mean 27 million people simultaneously opening a browser at 7:03pm and attacking the ABS servers. People complete it at different times, households submit information together, some people use paper forms, some people do it early and others presumably forget until someone reminds them.
But even after accounting for all of that, you’re still potentially dealing with millions of Australians interacting with the same government service over a relatively concentrated period. That’s not a normal website anymore. That’s a major infrastructure problem.
The Census Is Basically One Giant Traffic Spike
From a technical perspective, I actually find the whole thing fascinating. You need infrastructure capable of scaling rapidly, load balancing, databases that can deal with enormous spikes in activity, redundancy, monitoring, backups, disaster recovery and enough cybersecurity protection to deal with whatever nonsense people inevitably decide to throw at a major government system while the entire country is using it.
The particularly strange thing about the Census is that you’re designing all of that around an event that happens once every five years. Most websites can provision infrastructure around relatively predictable traffic patterns and then leave some breathing room for unexpected spikes. The Census is different because the event itself is essentially one enormous national traffic spike. You know it’s coming, you know roughly when it’s coming and you know millions of people are going to pile into the same system around the same time.
Modern cloud infrastructure obviously makes this considerably more manageable than it would have been twenty years ago. Elastic computing exists specifically because throwing more capacity at temporary demand is a solvable problem. But that doesn’t make the engineering, security, testing and procurement around the system trivial. If anything, the fact that this is government infrastructure collecting incredibly sensitive information makes the requirements considerably more complicated.
And because we’re talking about government procurement, I can only imagine the number of zeroes involved. I would have loved to be the contractor that got that job.
This is where I start committing technological heresy, because despite building software for a living and generally believing that technology can make almost everything better, I genuinely think there are situations where the government should just use paper.
Not exclusively. I’m not suggesting we burn down AWS, purchase 27 million HB pencils and return Australia to 1973. But somewhere along the way we’ve started treating analogue systems as inherently outdated rather than recognising that analogue and digital systems have completely different strengths and weaknesses.
Paper has one particularly wonderful built-in feature: it doesn’t go offline because six million people are using paper at the same time. My Census form doesn’t suddenly display a 503 because someone in Perth is also holding a pen. There’s no database connection to time out, no JavaScript bundle refusing to load and no national authentication service deciding that tonight would be a fantastic evening for an existential crisis.
Of course, paper creates an entirely different set of problems. Someone has to distribute it, collect it, scan it, process it, store it and eventually transform millions of handwritten answers into useful structured data. That isn’t cheap or easy either. But that’s why I don’t think the answer should necessarily be paper or digital. It should be both, and both should be treated as genuine ways of completing the Census rather than one being the modern system and the other being the backup for people who haven’t discovered computers yet.
A hybrid system also distributes risk. Which brings me to the part of the Census that is considerably more uncomfortable.
This Is an Absolutely Delicious Database to Hack
Think about the information contained in an individual Census response. Not the nice aggregated statistics that eventually appear on the ABS website. I mean the actual information Australians submit.
Your name, age, address, household, relationships, employment, education, ancestry, religion, income range, health circumstances and family structure can all form part of an extraordinarily detailed picture of a person and the people they live with. From a statistical perspective, that information is incredible. From a cybersecurity perspective, it is the sort of database that makes me slightly uncomfortable knowing it exists at all.
Before someone interprets that as me saying the Census has been hacked, I’m not. I have seen no evidence that Census responses have been leaked, compromised or stolen. But the risk inherent in collecting enormous quantities of sensitive information is still something worth talking about. Any sufficiently valuable centralised collection of information becomes an attractive target, and Australia doesn’t exactly have a spotless recent history when it comes to organisations holding gigantic amounts of personal information.
Imagine, purely hypothetically, the catastrophe if raw Census responses were ever compromised. It would be an absolutely monumental mess. Horrifying, obviously, but also almost comically Australian. “Yeah, sorry everyone. Remember that legally required national survey where we asked you to describe basically your entire household? Funny story...”
That is one of the fundamental trade-offs of centralising information. Centralisation makes data incredibly useful because you can analyse everything together, but it also makes whatever is holding that information incredibly valuable. The bigger and more comprehensive the dataset becomes, the more catastrophic a serious breach potentially becomes.
There is actually something funny about comparing this with snail mail. Physical information absolutely isn’t secure by default. Someone can steal your mail, intercept a form, break into an office or access records they shouldn’t. But physical theft has a scaling problem. If you want to steal ten million paper Census forms, you’re going to need a fairly impressive van.
Digital information doesn’t have that limitation. A sufficiently serious compromise can potentially expose enormous quantities of information at once. That’s one of the strange things about our transition towards digital infrastructure: we’ve massively improved accessibility, processing speed and analytical capability while simultaneously inventing entirely new categories of catastrophic failure.
Unfortunately, I Absolutely Love Census Data
This is where my moral high ground completely collapses, because despite everything I’ve just written, I freakin’ love Census data.
I use it constantly.
Once all of those individual responses have been processed, protected and turned into aggregated statistical datasets, the Census becomes an extraordinary resource for understanding Australia. You can analyse demographics, population movements, household structures, employment, education, age distributions, housing, income, cultural backgrounds and the differences between completely different parts of the country.
This isn’t just interesting trivia either. It can genuinely improve decision-making. If an area suddenly has significantly more young families living in it, infrastructure requirements are going to change. If another region is ageing rapidly, healthcare requirements are going to change. If Australia’s population is gradually migrating from one part of the country towards another, housing, transport, schools, hospitals and public services eventually need to follow them.
Good government requires good information. Otherwise we’re effectively allocating billions of dollars based on vibes, and while that occasionally appears to be our existing economic policy, I’d rather we didn’t formalise it.
The slightly awkward part is that I also use this information for marketing.
Obviously I’m not accessing individual Census responses. I’m using the aggregated public datasets released afterwards. But Census data is enormously useful for demographic modelling because it allows you to understand the characteristics of different geographical areas and build much better population profiles.
What I’m particularly interested in isn’t simply asking where people live now. I want to understand where they’re going.
If you combine multiple Census periods, you can start looking at how demographic characteristics move geographically over time. People age, families form, people move, suburbs gentrify, housing affordability pushes populations further outward, employment centres change and infrastructure creates entirely new areas of growth. Entire demographic groups can gradually migrate across a city or region.
With enough historical information, you can start modelling those movements and eventually ask what the next movement might look like. That’s incredibly useful for governments planning infrastructure, but it’s also incredibly useful for businesses. If you’re deciding where to open a physiotherapy clinic, restaurant, retail store or practically any location-dependent business, knowing where your customers might live five years from now is considerably more useful than simply knowing where they lived five years ago.
So yes, I am deeply suspicious of enormous government datasets.
Please release the enormous government dataset as soon as possible.
Thank you.
Doesn’t the Government Already Know Most of This?
There was another thought that kept popping into my head while I was completing the Census: why am I telling the government some of this information in the first place?
The Australian Government and state governments already maintain an extraordinary number of administrative datasets. Births and deaths are registered. The tax system contains financial and employment information. Governments maintain business records. Medicare exists. Centrelink exists. Immigration records exist. Education systems exist. Property records exist. ASIC exists. The ATO exists. There are already enormous quantities of information about Australians scattered throughout government.
Now, there are extremely good reasons those databases aren’t all joined together into one gigantic government super-database where every public servant can search your name and see your entire existence. Different agencies collect information for different purposes and operate under different legislation, privacy requirements and access controls. Combining all of it indiscriminately would arguably create an even bigger privacy nightmare than the Census itself.
But from the citizen’s perspective, it still creates this bizarre contradiction where the government simultaneously knows an enormous amount about us while behaving like none of its departments have ever met each other.
We submit information to one department. Another department asks us for the same information. Then another government service asks again. Then every five years we fill out a Census form containing information that may overlap with administrative records already sitting somewhere else inside government.
That’s partly because there are legitimate privacy protections preventing government agencies from simply sharing everything with each other, which is probably a good thing. But it also exposes one of the biggest problems with government digital transformation: government itself is incredibly fragmented.
So rather than digital systems removing bureaucracy, sometimes we just digitise the bureaucracy.
And Who Has to Clean This Data?
The other thing that caught my attention while filling it out was the number of questions that allow relatively open text responses.
This is admittedly the data nerd part of my frustration because the moment I see a free-text field in a dataset I immediately think about the poor bugger who eventually has to clean it.
If you allow humans to freely type things into a database, humans will very quickly demonstrate why humans should not be allowed to freely type things into databases. One person writes “Marketing Manager”, another writes “Marketing Mgr”, another writes “Digital Marketing”, another just writes “marketing”, and someone inevitably describes themselves as a “professional Facebook wizard”.
There are legitimate reasons to allow free-text responses. Predefined categories can miss unusual occupations, cultural identities, emerging industries and answers that whoever designed the form simply didn’t anticipate. You don’t want the structure of your questionnaire determining what kinds of Australians are allowed to statistically exist.
But from a data-processing perspective, it must be an absolute nightmare. You’re dealing with spelling mistakes, abbreviations, slightly different descriptions of identical things, jokes, nonsense and completely legitimate answers that don’t fit neatly into existing classifications.
Enjoy your fuzzy matching.
Godspeed.
So... Should We Actually Be Worried About the Census?
This is where I reach the incredibly satisfying conclusion that I don’t actually know.
I started this entire thought process because I was annoyed at a website, but the Census sits directly in the middle of two things I strongly believe. I believe personal information should be protected aggressively and that governments shouldn’t collect information merely because collecting it might someday be useful. At the same time, I believe good decisions require good data, and you cannot meaningfully understand a country of roughly 27 million people without collecting information about the people who actually live in it.
Collect too little information and government planning becomes worse. Collect enormous amounts of information and you’ve created an enormous privacy and cybersecurity responsibility. Digitise the process and the information becomes faster to process, easier to analyse and potentially cheaper to collect at scale, but you’ve also created technological dependencies and concentrated cyber risks that paper never had. Keep everything analogue and you reduce some of those digital risks while making collection and processing slower, more cumbersome and potentially more expensive.
It’s a Catch-22. Damned if you do, damned if you don’t.
Maybe that’s actually what bothers me about the constant push towards government digitisation. Digital transformation isn’t inherently progress. Data collection isn’t inherently surveillance. Paper isn’t inherently obsolete. Technology isn’t automatically the correct solution simply because it’s newer. Every one of these systems involves trade-offs, and we’ve become extraordinarily good at talking about the benefits of digitisation without spending nearly enough time talking about the new risks and dependencies we’re creating along the way.
Maybe the question governments should be asking isn’t, “How do we digitise this?”
Maybe it should be, “Does digitising this actually make it better?”
Because if the answer is no, I’m perfectly happy with a piece of paper.
Just make sure you digitise the results afterwards.
I have models to build.


