AI Can Now Doxx Your Anonymous Accounts for as Little as $1, at Scale

Internet anonymity has never looked more fragile.

by · ZME Science
AI-generated image. Illustration: ZME Science.

It can cost less than a cup of coffee to try to strip the pseudonymity from an online account. Researchers from ETH Zurich, MATS, and Anthropic have shown that an AI agent can hunt for the person behind a pseudonymous profile for roughly $1 to $4 per attempt.

Deanonymization isn’t new. Neither is doxxing, the related practice of exposing someone’s private identifying information online. But finding the person behind an account has traditionally required time, skill, and a lot of tedious searching. Now, much of that work can potentially be automated.

“The core discovery is that LLMs can identify anonymous internet users at scale. What once required a skilled human investigator hours of manual labour can now be done by AI in minutes. We found that LLM-based approaches massively outperform classical methods that rely on structured data,” says Joshua Swanson, one of the study authors.

It can cost just a few dollars to hunt for the person behind an account

The team’s most direct experiment started with Hacker News users whose real identities were already known because they had linked their profiles to LinkedIn. The researchers removed names, URLs, handles, and other obvious identifiers, then asked an AI agent with web access to work out who each person was.

The researchers deliberately chose accounts with known identities so they could tell whether the AI was right, and they acknowledge that these profiles may be easier to identify than genuinely pseudonymous accounts.

The AI correctly identified 226 of 338 people, giving the system 67% recall at 90% precision. In other words, it found about two-thirds of the targets, and nine out of ten identities it proposed were correct.

In other experiments, it wasn’t quite as effective. On 25 identifiable Reddit academics, after stripping names and paper titles, the agent identified 13 of 25, or 52%, at 72% precision. On 36 software-engineering career posters whose identities were known from LinkedIn links, it identified 9 of 36, or 25%, at 90% precision.

It’s not like AI can now reliably type a username into a box and reveal anyone’s identity.
But it doesn’t have to be perfect to change the privacy equation.

×

Get smarter every day...

Stay ahead with ZME Science and subscribe.

Daily Newsletter
The science you need to know, every weekday.

Weekly Newsletter
A week in science, all in one place. Sends every Sunday.
No spam, ever. Unsubscribe anytime. Review our Privacy Policy.

Thank you! One more thing...

Please check your inbox and confirm your subscription.

The more you post, the more vulnerable you are

The system works because people constantly leak tiny pieces of themselves online. A comment might mention a university course, and that’s a bit of useful information. Another reveals a city. Somewhere else, you talk about your profession, a niche hobby, a programming tool, a favorite movie, or the place you used to live.

Individually, that’s not nearly enough to identify anyone. But put enough bits together, and you end up with a “fingerprint,” says Swanson.

“It’s rarely just one thing. It’s the combination of facts: the city you live in, your job, any niche hobby. Individually, these might be harmless. Together, they form a full fingerprint.” Even something as mundane as movie taste could give people away.

For their larger-scale matching experiments, the researchers developed a four-stage pipeline. First, an LLM extracts potentially identifying features from someone’s posts. Next, embeddings (numerical representations of meaning) narrow a huge candidate pool down to profiles that look similar. A stronger model then reasons over the most promising matches. Finally, the system calibrates its confidence so it can abstain when the evidence isn’t strong enough.

This is where LLMs have an advantage over older deanonymization techniques. They don’t need every useful clue to arrive neatly formatted in a database. They can pull information from messy human language and reason about combinations of clues that would otherwise take someone a long time to investigate manually.

Could AI link your new account to your old one?

The team also tested whether AI could recognize people across time.

RelatedPosts

Apple May Be Planning AirPods Cameras for Siri Alongside a 20th Anniversary iPhone
Federal Workers Say They’re Being Watched by AI for Saying Anything Bad about Trump or Musk
3D printing living things might revolutionize pastry
AI companies plan to use nuclear energy for their energy-hungry data centers. But should they?

They split Reddit users’ histories into an earlier and later profile and deliberately removed a one-year window between them. That prevented the system from simply matching people because they happened to be discussing the same news event or temporary obsession.

The strongest version of the system still linked 38.4% of the users at 99% precision. In other words, when it chose to make matches at that threshold, almost all of them were correct, while it successfully linked nearly four in ten of the available users.

This doesn’t mean anonymity can’t exist on the internet anymore. But it does mean that one of the main practical obstacles towards doxxing someone is starting to fade.

For years, pseudonymity has benefited from what the researchers call “practical obscurity.” Enough clues might exist to identify you, but finding, reading, and connecting them could require many hours of skilled work. That naturally limited how many people anyone could investigate.

LLMs can attack that bottleneck.

Ordinary commercial AI systems can now make this kind of privacy attack much cheaper and easier to scale.

We already know that we leave small traces of our identity online, even when we try to be careful. Until recently, the saving grace was that those clues could be painfully difficult to piece together.

AI is changing that equation. Online pseudonymity isn’t dead, but maintaining it may be becoming much harder.

The study was published in the pre-print arXiv.