If your company uses the Meta pixel, Customer Match, or LiveRamp, your privacy policy probably says something that a Yale paper just proved is false. And that wasn’t even new.
If your company uses the Meta Pixel, Customer Match or LiveRamp, your privacy policy probably says something that a Yale paper has just proven false. And it wasn’t even news.
Many of us have come across advertising leaflets or — much worse — contracts from advertising companies that say things like, “If you give us your customer or user lists — their identifiers, such as email addresses or phone numbers — we’ll cross-reference them with our own databases to optimise the audience, find lookalikes and maximise the perfect match rate for your campaign.”
And here’s the kicker: “Don’t worry, we’ll never be able to read those identifiers: we anonymise them using a ‘hash function’ or convert them into hashes before uploading them to our platform, so we can’t read them.”
1.- Heads or tails?
If I toss a coin in the air and hide the result, you will never be able to know whether it came up heads or tails.
You won’t be able to know it either if I tell you that I’ve hashed the result and show it to you:
The result was: “f61f5c5e22749aeba2bd920352dbd4c9f63e0ebaed6df623b09c836ffcbec2eb”
Or can I know it?
Of course I can know it: I need to know which hash function you used and hash the two possible results. At a glance I will know if it was heads or tails.
Now it is much easier to understand a slightly less silly example.
Meta, in explaining the “advanced matching” of its Meta pixel, says it collects the gender data of its users (masculine “m” or “f” feminine) and then “hashes” it and boasts that it protects user privacy because who could know that, for example, “62c66a7a5dd70c3146618063c344e531e6d4b59e379808443ce962b3abd63c5a” is “masculine”?
Again: extracting the original value “m” back from “
62c66a7a5dd70c3146618063c344e531e6d4b59e379808443ce962b3abd63c5a” is impossible.
The problem is that you don’t need to:
There are only two options:
“m”: 62c66a7a5dd70c3146618063c344e531e6d4b59e379808443ce962b3abd63c5a
“f”: 252f10c83610ebca1a059c0bae8255eba2f95be4d1d7bcfa89d7248a82d9f111
The first is “masculine” and the second is “feminine”.
Ergo, the protection of that attribute is, in reality, null.
Like that of a coin toss.
Cryptography experts speak in Sanskrit but the concept is simple, right?.
Let’s go with another example, a bit less silly.
You are reading ZERO PARTY DATA. The newsletter on current affairs and technology law by Jorge García Herrero and Darío López Rincón.
In the free time that this newsletter leaves us, we solve complicated issues related to personal data protection regulations and artificial intelligence. If you have any of these, make a sign with your hand. Or contact us by email at jgh(arroba)jorgegarciaherrero.com
2.- How to know who will win the Ballon d’Or
Imagine I tell you that I’ve installed Pegasus on the person counting the votes to choose the Ballon d’Or and that I’ve intercepted the communication in which they reveal the winner. But this person was careful and had “hashed” the name before sending it.
The winner is:
“8cf7dd0d99b272e32d34a6a18dd87b1ddde9b297aa6b0980c1d6a75ef7a886d6”
Yes, unfortunately the winner’s name was hashed with SHA-256. Impossible to crack.
Since there are only 30 candidates, we can simply “hash” them one by one or all at once and compare, and the match will be the winner.
You can try doing it in this artifact (link).
All these transparent examples come from the same paper: “Anonymity, Consent, And Other Noble Lies: An Empirical Study of The Data Economy” by Reardon, Egelmann, and others.
You should all read this paper: it is as devastating, precise, and transparent in its explanations as any of Daniel Solove’s.
There are other delicious moments like the Meta lawyers’ quotes claiming contrary arguments in different trials, depending on the interest defended at each moment.
In Meta Pixel Healthcare it maintains that it sends de-identified information, and in Meta v. BrandTotal their expert declares that hashing user IDs does not anonymize.
Ah, damn lawyers.
It is no coincidence that it is cited in our own paper “Prompt like a Butterfly, Sting like a Tracker” with IMDEA Networks when we highlight that when AI agents share your hashed username and/or email address with third parties (expert data brokers), they have not provided you with any protection thereby.
It is no coincidence that the FTC has already warned about this a few times as well. The latest in a post with an eloquent title: “No, Hashing Still Doesn’t Make Your Data Anonymous”.
3 Do Bradley Cooper or Jessica Alba leave good tips?
My favorite example is that of a smart-aleck journalist who tried to extract information from an “anonymized” data set that would be interesting to many: Which famous New Yorkers are big spenders and which ones are stingy when tipping taxi drivers?
Someone filed a Freedom of Information request against the NYC government, which published a data set of taxi trips that included, among other columns, the license number, the “Medallion”—that is, the little number each taxi carries on its roof—the customer’s pickup point, the drop-off point, the type of fare applied, the fare amount, and the passenger’s tip.
To “preserve the privacy” of the taxi drivers, the license number and medallion columns were—can you guess?—hashed with MD5, converted into an alphanumeric string like the old ones.
Hashes are irreversible. The problem is that there is a lot of additional information within our reach directly related to the data set.
The savvy journalists looked for photographs in which one could identify: (i) a minor celebrity getting into a taxi, (ii) the “medallion” of that taxi, and (iii) the street where it happened (“customer pickup point”).
They found many.
By now you can probably imagine the rest:
They hashed the taxi medallion from the photo and looked up, for that medallion, the ride that started that day on that street. And BAM, all that’s left is to go to the tip column.
Spoiler: Bradley Cooper and Jessica Alba had no tip recorded in a country where the custom is to leave substantial tips.
Plot twist: But wait a minute! The tip was only recorded if paid by card and they probably paid in cash. Watch out for empty columns in data tables.
4.- Hashed email: in practice, it is much easier
The authors purchased (well, asked for free) samples from four data brokers: more than 6 million hashed emails. They re-identified more than half using only “rainbow tables” of made-up emails, 88% by cross-referencing them with public breaches, and 97% by applying password cracking techniques. Nice.
But aaaaall these examples are more or less striking, because the “attacker” who wants to re-identify the dataset is playing with one hand tied behind their back.
That is not the case for data brokers in their day-to-day activity.
Data brokers have attribute tables for each person, with many data points.
One example: all mobile phones have an identification number for advertising purposes (“Android Advertising ID” (AAID) and “ID for Advertisers” (IDFA) on the iPhone).
Both Android and iPhones allow you to change their respective IDs at will.
Is it useful for anything? For very little.
Data brokers sell tables with these device identifiers paired with the hashed email. Thus, when you reset the advertising ID, the email links you back to the old ID and to all your other devices.
Furthermore, companies, aware of the possibility, extract their own device identifier upon installation (much easier and cheaper than fingerprinting, which is also possible) and, even if you change the advertising ID, in the next “event” the app communicates your “new” advertising ID with the long-standing installation ID, and so much for your attempts to protect your privacy.
After this example, it is now easier to understand that it is trivial to reidentify the majority of users, through the device they use, or four geolocation data points (where you work / where you sleep and a couple more).
5.- What practical relevance does this have?
When you sign up to receive your store purchase receipt in digital format, they ask for two pieces of data (mobile phone number and then email address -on top of that via Meta’s WhatsApp-) where only the email was strictly necessary, which is where they are going to send you the receipt for every purchase.
Any company interested in purchasing advertising provides their customer or user lists to companies like Meta or Google to “determine audiences,” which, explained in a very simple way, means that, for example, if those department stores hire a personalized advertising campaign on Instagram, that advertisement (i) will not be shown to you, Mari Pili, because you are already a customer (they know because you have already provided your email and mobile number) and (ii) it will be shown to “lookalikes”: to Marta and Pedro, who are people with the same profile as Mari Pili (iii) it will be shown tirelessly to Mari Pili, if it knows she has looked at a certain product and has not bought it. As we all know.
The issue is that the privacy policy of the company, and Meta’s, and Google’s, and Criteo’s, LiveRamp’s, and other data brokers will tell you “we hash your users’ or customers’ identifiers” before performing our magic and that is how we manage to improve the relevance of our ads and measure the impact of our campaigns without identifying the data subject.without identifying the data subject.
The truth is that, as has already been made clear, the department stores, Meta, Google, Criteo, and all those ahem “partners” or “identity providers” know perfectly well that this hash of letters and numbers is Mari Pili.
Footnote: this post could have been oriented toward the suggestive implications of all this when applying the SRB/Scania doctrine, but in my interpretation, there is no need: If what the store, Meta, or Criteo want is for a personalized store ad to be shown to Marta and Pedro, but not to Mari Pili, considering the content, purpose, and effects of the processing, it is completely irrelevant whether the information available identifies them by their full names (Marta, Pedro, Mari Pili) or by unique hashes, because the purpose (hitting them with personalized ads) is achieved regardless.
6.- A couple of real-life examples
Nasdaq, Inc. In the list of third parties linked to its privacy policy, it says it may share “hashed and de-identified email addresses” with LiveRamp, along with IP and advertising identifiers. In the same sentence, it adds that LiveRamp uses this information to link the user’s device with its databases and to perform targeted advertising.
Google. In the Customer Match help for advertisers, it states that “Google doesn’t receive actual email addresses”, because it transforms the emails from its accounts into one-way SHA-256 hashes. We now know that if anyone says this, it’s false, but for Google to say it is laughable, huh? HUH?
LiveRamp. In a corporate post about RampID claims that the pseudonymized approach makes it possible to consolidate, enrich, and segment data “without exposing or sharing a customer’s private information”. Yeah baby, yeah.
Procter & Gamble: In “Addressable media“: encrypt the data or use UID2 and upload “a pseudonymized version (replaced with artificial numbers or letters)” of the email, phone number, or advertising ID to platforms like Facebook, YouTube, Instagram, or TikTok. It also includes a list of platforms with which it can share a hashed version of the email.
Footnote 2: Of course, all of this can be done properly: the solution is to add a “salt” before hashing the data. The salt protects against third parties (but among those who share the key), the data remains personal… but this post has already become too long.
Footnote 3: perhaps you need to give your privacy policy a fresh coat of paint if it classifies a hashed email as “anonymous” or “de-identified,” or rethink the legal basis for all that stuff you upload to Customer Match or Custom Audiences.







