LocateBaltimore
No Result
View All Result
No Result
View All Result
LocateBaltimore
No Result
View All Result
Home Technology

How to verify a data breach

Pauline Wright by Pauline Wright
March 16, 2024
in Technology
0
326
SHARES
2.5k
VIEWS
Share on FacebookShare on Twitter


Over the years, TechCrunch has extensively coated data breaches. In truth, a few of our most-read tales have come from reporting on large data breaches, comparable to revealing shoddy safety practices at startups holding delicate genetic info or disproving privateness claims by a fashionable messaging app.

It’s not simply our delicate info that may spill on-line. Some data breaches can comprise info that may have important public curiosity or that’s extremely helpful for researchers. Last 12 months, a disgruntled hacker leaked the inner chat logs of the prolific Conti ransomware gang, exposing the operation’s innards, and a large leak of a billion resident data siphoned from a Shanghai police database revealed a few of China’s sprawling surveillance practices.

But one of many largest challenges reporting on data breaches is verifying that the data is genuine, and never somebody making an attempt to sew collectively pretend data from disparate locations to promote to patrons who’re none the wiser.

Verifying a data breach helps each firms and victims take motion, particularly in instances the place neither are but conscious of an incident. The sooner victims learn about a data breach, the extra motion they’ll take to defend themselves.

Author Micah Lee wrote a ebook about his work as a journalist authenticating and verifying massive datasets. Lee lately revealed an excerpt from his ebook about how journalists, researchers and activists can verify hacked and leaked datasets, and the way to analyze and interpret the findings.

Every data breach is totally different and requires a distinctive method to decide the validity of the data. Verifying a data breach as genuine would require utilizing totally different instruments and methods, and in search of clues that may assist determine the place the data got here from.

In the spirit of Lee’s work, we additionally needed to dig into a few examples of data breaches we have now verified prior to now, and the way we approached them.

How we caught StockX hiding its data breach affecting tens of millions

It was August 2019 and customers of the sneaker promoting market StockX acquired a mass e-mail saying they need to change their passwords due to unspecified “system updates.” But that wasn’t true. Days later, TechCrunch reported that StockX had been hacked and somebody had stolen tens of millions of buyer data. StockX was compelled to admit the reality.

How we confirmed the hack was partially luck, but it surely additionally took a lot of labor.

Soon after we revealed a story noting it was odd that StockX would drive probably tens of millions of its clients to change their passwords with out warning or rationalization, somebody contacted TechCrunch claiming to have stolen a database containing data on 6.8 million StockX clients.

The particular person mentioned they have been promoting the alleged data on a cybercrime discussion board for $300, and agreed to present TechCrunch a pattern of the data so we may verify their declare. (In actuality, we might nonetheless be confronted with this similar state of affairs had we seen the hacker’s on-line posting.)

The particular person shared 1,000 stolen StockX consumer data as a comma-separated file, primarily a spreadsheet of buyer data on each new line. That data appeared to comprise StockX clients’ private info, like their identify, e-mail tackle, and a copy of the shopper’s scrambled password, together with different info believed distinctive to StockX, such because the consumer’s shoe measurement, what machine they have been utilizing, and what foreign money the shopper was buying and selling in.

In this case, we had an thought of the place the data initially got here from and labored underneath that assumption (except our subsequent checks steered in any other case). In concept, the one individuals who know if this data is correct are the customers who trusted StockX with their data. The higher the quantity of people that verify their info was legitimate, the higher likelihood that the data is genuine.

Since we can not legally verify if a StockX account was legitimate by logging in utilizing a particular person’s password with out their permission (even when the password wasn’t scrambled and unusable), TechCrunch had to contact customers to ask them instantly.

StockX’s password reset e-mail to clients citing unspecified “system updates.” Image Credits: file picture.

We will usually search out individuals who we all know will be contacted shortly and reply immediately, comparable to via a messaging app. Although StockX’s data breach contained solely buyer e-mail addresses, this data was nonetheless helpful since some messaging apps, like Apple’s iMessage, enable e-mail addresses instead of a cellphone quantity. (If we had cellphone numbers, we may have tried contacting potential victims by sending a textual content message.) As such, we used an iMessage account arrange with a @techcrunch.com e-mail tackle so the folks we have been contacting knew the request was actually coming from us.

Since that is the primary time the StockX clients we contacted have been listening to about this breach, the communication had to be clear, clear and explanatory and had to require little effort for recipients to reply.

We despatched messages to dozens of individuals whose e-mail addresses used to register a StockX account have been @icloud.com or @me.com, that are generally related to Apple iMessage accounts. By utilizing iMessage, we may additionally see that the messages we despatched have been “delivered,” and in some instances relying on the particular person’s settings it mentioned if the message was learn.

The messages we despatched to StockX victims included who we have been (“I’m a reporter at TechCrunch”), and the explanation why we have been reaching out (“We discovered your info in an as-yet-unreported data breach and wish your assist to verify its authenticity so we will notify the corporate and different victims”). In the identical message, we offered info that solely they may know, comparable to their username and shoe measurement that was related to the identical e-mail tackle we’re messaging. (“Are you a StockX consumer with [username] and [shoe size]?”). We selected info that was simply confirmable however nothing too delicate that would additional expose the particular person’s personal data if learn by another person.

By writing messages this manner, we’re constructing credibility with a one who could don’t know who we’re, or could in any other case ignore our message suspecting it’s some form of rip-off.

We despatched related customized messages to dozens of individuals, and heard again from a portion of these we contacted and adopted up with. Usually a chosen pattern measurement of round ten or a dozen confirmed accounts would counsel legitimate and genuine data. Every one who responded to us confirmed that their info was correct. TechCrunch offered the findings to StockX, prompting the corporate to strive to get forward of the story by disclosing the huge data breach in a assertion on its web site.

How we discovered leaked 23andMe consumer data was real

Just like StockX, 23andMe’s current safety incident prompted a mass password reset in October 2023. It took 23andMe one other two months to verify that hackers had scraped delicate profile data on 6.9 million 23andMe clients instantly from its servers — data on about half of all 23andMe’s clients.

TechCrunch discovered pretty shortly that the scraped 23andMe data was possible real, and in doing so discovered that hackers had revealed parts of the 23andMe data two months earlier in August 2023. What later transpired was that the scraping started months earlier in April 2023, however 23andMe failed to discover till parts of the scraped data started circulating on a fashionable subreddit.

The first indicators of a breach at 23andMe started when a hacker posted on a identified cybercrime discussion board a pattern of 1 million account data of Ashkenazi Jews and 100,000 customers of Chinese descent who use 23andMe. The hacker claimed to have 23andMe profile, ancestry data, and uncooked genetic data on the market.

But it wasn’t clear how the data was exfiltrated or even when the data was real. Even 23andMe mentioned on the time it was working to verify whether or not the data was genuine, an effort that might take the corporate a number of extra weeks to verify.

The pattern of 1 million data was additionally formatted in a comma-separated spreadsheet of data, revealing reams of equally and neatly formatted data, every line containing an alleged 23andMe consumer profile and a few of their genetic data. There was no consumer contact info, solely names, gender, and delivery years. But this wasn’t sufficient info for TechCrunch to contact them to verify if their info was correct.

The exact formatting of the leaked 23andMe data steered that every report had been methodically pulled from 23andMe’s servers, one after the other, however possible at excessive velocity and appreciable quantity, and arranged into a single file. Had the hacker damaged into 23andMe’s community and “dumped” a copy of 23andMe’s consumer database instantly from its servers, the data would possible current itself in a totally different format and comprise further details about the server that the data was saved on.

One factor instantly stood out from the data: Each consumer report contained a seemingly random 16-character string of letters and numbers, referred to as a hash. We discovered that the hash serves as a distinctive identifier for every 23andMe consumer account, but additionally serves as a part of the net tackle for the 23andMe consumer’s profile once they log in. We checked this for ourselves by creating a new 23andMe consumer account and in search of our 16-character hash in our browser’s tackle bar.

We additionally discovered that loads of folks on social media had historic tweets and posts sharing hyperlinks to their 23andMe profile pages, every that includes the consumer’s distinctive hash identifier. When we tried to entry the hyperlinks, we have been blocked by a 23andMe login wall, presumably as a result of 23andMe had mounted no matter flaw had been exploited to allegedly exfiltrate large quantities of account data and worn out all public sharing hyperlinks within the course of. At this level, we believed the consumer hashes could possibly be helpful if we have been ready to match every hash in opposition to different data on the web.

When we plugged in a handful of 23andMe consumer account hashes into serps, the outcomes returned net pages containing reams of matching ancestry data revealed years earlier on web sites run by family tree and ancestry hobbyists documenting their very own household histories.

In different phrases, among the leaked data had been revealed partially on-line already. Could this be outdated data sourced from earlier data breaches?

One by one, the hashes we checked from the leaked data completely matched the data revealed on the family tree pages. The key factor right here is that the 2 units of data have been formatted considerably in another way, however contained sufficient of the identical distinctive consumer info — together with the consumer account hashes and matching genetic data — to counsel that the data we checked was genuine 23andMe consumer data.

It was clear at this level that 23andMe had skilled a large leak of buyer data, however we couldn’t verify for certain how current or new this leaked data was.

A family tree hobbyist whose web site we referenced for wanting up the leaked data advised TechCrunch that that they had about 5,000 relations found via 23andMe documented meticulously on his web site, therefore why among the leaked data matched the hobbyist’s data.

The leaks didn’t cease. Another dataset, purportedly on 4 million British customers of 23andMe, was posted on-line within the days that adopted, and we repeated our verification course of. The new set of revealed data contained quite a few matches in opposition to the identical beforehand revealed data. This, too, appeared to be genuine 23andMe consumer data.

And in order that’s what we reported. By December, 23andMe admitted that it had skilled a large data breach attributed to a mass scrape of data.

The firm mentioned hackers used their entry to round 14,000 hijacked 23andMe accounts to scrape huge quantities of different 23andMe customers’ account and genetic data who opted in to a characteristic designed to match relations with related DNA.

While 23andMe tried to blame the breach on the victims whose accounts have been hijacked, the corporate has not defined how that entry permitted the mass downloading of data from the tens of millions of accounts that weren’t hacked. 23andMe is now going through dozens of class-action lawsuits associated to its safety practices prior to the breach.

How we confirmed that U.S. navy emails have been spilling on-line from a authorities cloud

Sometimes the supply of a data breach — even an unintentional launch of non-public info — just isn’t a shareable file filled with consumer data. Sometimes the supply of a breach is within the cloud.

The cloud is a fancy time period for “another person’s pc,” which will be accessed on-line from wherever on the planet. That means firms, organizations and governments will retailer their recordsdata, emails, and different office paperwork in huge servers of on-line storage typically run by a handful of the Big Tech giants, like Amazon, Google, Microsoft, and Oracle. And, for his or her extremely delicate clients like governments and militaries, the cloud firms supply separate, segmented and extremely fortified clouds for additional safety in opposition to probably the most devoted and resourced spies and hackers.

In actuality, a data breach within the cloud will be so simple as leaving a cloud server related to the web with out a password, permitting anybody on the web to entry no matter contents are saved inside.

It occurs, and greater than you may assume. People really discover them! And some of us are actually good at it.

Anurag Sen is a good-faith safety researcher who’s well-known for locating delicate data mistakenly revealed to the web. He’s discovered quite a few spills of data through the years by scouring the net for leaky clouds with the aim of getting them mounted. It’s a good factor, and we thank him for it.

Over the Presidents Day federal vacation weekend in February 2023, Sen contacted TechCrunch, alarmed. He discovered what regarded just like the delicate contents of U.S. navy emails spilling on-line from Microsoft’s devoted cloud for the U.S. navy, which ought to be extremely secured and locked down. Data spilling from a authorities cloud just isn’t one thing you see fairly often, like a rush of water blasting from a gap in a dam.

But in actuality, somebody, someplace (and by some means) eliminated a password from a server on this supposedly extremely fortified cloud, successfully punching a large gap on this cloud server’s defenses and permitting anybody on the open web to digitally dive in and peruse the data inside. It was human error, not a malicious hack.

If Sen was proper and these emails proved to be real U.S. navy emails, we had to transfer shortly to make sure the leak was plugged as quickly as doable, fearing that somebody nefarious would quickly discover the data.

Sen shared the server’s IP tackle, a string of numbers assigned to its digital location on the web. Using a web-based service like Shodan, which mechanically catalogs databases and servers discovered uncovered to the web, it was simple to shortly determine a few issues concerning the uncovered server.

First, Shodan’s itemizing for the IP tackle confirmed that the server was hosted on Microsoft’s Azure cloud particularly for U.S. navy clients (also referred to as “usdodeast“). Second, Shodan revealed particularly what utility on the server was leaking: an Elasticsearch engine, typically used for ingesting, organizing, analyzing and visualizing large quantities of data.

Although the U.S. navy inboxes themselves have been safe, it appeared that the Elasticsearch database tasked with analyzing these inboxes was insecure and inadvertently leaking data from the cloud. The Shodan itemizing confirmed the Elasticsearch database contained about 2.6 terabytes of data, the equal of dozens of exhausting drives filled with emails. Adding to the sense of urgency in getting the database secured, the data contained in the Elasticsearch database could possibly be accessed via the net browser just by typing within the server’s IP tackle. All to say, these navy emails have been extremely simple to discover and entry by anybody on the web.

By this level, we ascertained that this was virtually actually actual U.S. navy e-mail data spilling from a authorities cloud. But the U.S. navy is gigantic and disclosing this was going to be tough, particularly throughout a federal vacation weekend. Given the potential sensitivity of the data, we had to work out shortly who to contact and make this their precedence — and never drop emails with probably delicate info into a faceless catch-all inbox with no assure of getting a response.

Sen additionally supplied screenshots (a reminder to doc your findings!) displaying uncovered emails despatched from a variety of U.S. navy e-mail domains.

Since Elasticsearch data is accessible via the net browser, the data inside will be queried and visualized in a variety of methods. This can assist to contextualize the data you’re coping with and supply hints as to its potential possession.

A screenshot displaying how we queried the database to rely what number of emails contained a search time period, comparable to an e-mail area. In this case, it was “socom.mil,” the e-mail area for U.S. Special Operations Command. Image Credits: TechCrunch

For instance, most of the screenshots Sen shared contained emails associated to @socom.mil, or U.S. Special Operations Command, which carries out particular navy operations abroad.

We needed to see what number of emails have been within the database with out taking a look at their probably delicate contents, and used the screenshots as a reference level.

By submitting queries to the database inside our net browser, we used the in-built Elasticsearch “rely” parameter to retrieve the variety of instances a particular key phrase — on this case an e-mail area — was matched in opposition to the database. Using this counting approach, we decided that the e-mail area “socom.mil” was referenced in additional than 10 million database entries. By that logic, since SOCOM was considerably affected by this leak, it ought to bear some accountability in remediating the uncovered database.

And that’s who we contacted. The uncovered database was secured the next day, and our story revealed quickly after.

It took a 12 months for the U.S. navy to disclose the breach, notifying some 20,000 navy personnel and different affected people of the data spill. It stays unclear precisely how the database grew to become public within the first place. The Department of Defense mentioned the seller — Microsoft, on this case — “resolved the problems that resulted within the publicity,” suggesting the spill was Microsoft’s accountability to bear. For its half, Microsoft has nonetheless not acknowledged the incident.


To contact this reporter, or to share breached or leaked data, you will get in contact on Signal and WhatsApp at +1 646-755-8849, or by e-mail. You also can ship recordsdata and paperwork through SecureDrop.



Source hyperlink

Tags: breachCybersecuritydatadata breachdata exposedata leaksdata protectionprivacyVerify
Previous Post

The TikTok divestment bill's progress slows down in the Senate, as Majority Leader Chuck Schumer has not decided whether to bring the bill to the floor (New York Times)

Next Post

Patton Oswalt is coming to Baltimore at The Lyric in July

Next Post
Patton Oswalt is coming to Baltimore at The Lyric in July

Patton Oswalt is coming to Baltimore at The Lyric in July

No Result
View All Result

Categories

  • Construction (53)
  • Food (977)
  • Local News (1,995)
  • Local Sports (1,999)
  • Technology (4,000)

Recent.

How to Make Powdered Sugar (Without Cornstarch Option)

How to Make Powdered Sugar (Without Cornstarch Option)

August 25, 2026
Cream of Asparagus Soup with White Wine

Cream of Asparagus Soup with White Wine

August 25, 2026
Easy Whole Wheat Penne With Broccoli (18-Minute Base)

Easy Whole Wheat Penne With Broccoli (18-Minute Base)

August 24, 2026

We bring you the best Premium WordPress Themes that perfect for news, magazine, personal blog, etc. Check our landing page for details.

Category

  • Construction (53)
  • Food (977)
  • Local News (1,995)
  • Local Sports (1,999)
  • Technology (4,000)

Tags

2024 Draft 2024 Draft News Air apple Baltimore bridge Chicken Clifton Brown day Derrick Henry draft Easy Experiments Game Gameday Gameday News General Google Heres home Homepage Centerpiece Homepage Latest Headlines iPhone Jackson Key Lamar Lamar Jackson Late For Work Maryland NFL offseason OpenAI Ravens Recipe recipes Ryan Mink Savory season shopping tech TikTok users video Watch week
  • About
  • Home

© 2026 JNews - Premium WordPress news & magazine theme by Jegtheme.

No Result
View All Result
  • About
  • Home

© 2026 JNews - Premium WordPress news & magazine theme by Jegtheme.