Search Engines Uncover Compromising Documents

Using a search engine and free software tools, it's possible to dig up hidden -- even deleted -- information in documents posted to public web sites.

Using a search engine and free software tools, it’s possible to dig up hidden — even deleted — information in documents posted to public web sites.

Many search engines allow you to restrict your search results to non-HTML documents, such as Microsoft office documents, PDF files, and others. In addition to the text stored in these files, these types of documents often contain other types of information not intended to be seen by users.

This information includes metadata such as author name, organization, editing history, and can also include custom data such as the names of document reviewers, who the document was received from, and so on.

In addition to this metadata, many programs also store recently deleted text, allowing you to “undo” unwanted changes. Using simple, freely available software tools, much of this hidden metadata and seemingly deleted text can be converted into visible plain text.

Simon Byers, an AT&T security researcher, used a search engine to find more than 100,000 Microsoft Word files on the web, including business documents and resumes. He then used the free software tools “antiword” and “catdoc” to convert them to plain text.

Byers found deleted text and information including names, email headers, network paths and text from related documents — potentially compromising information that people publishing the documents to the web likely did not realize was included.

Byers suggested that job seekers, in particular, may not realize that even if they delete their social security number from a resume posted to the web, that the number may still be included in the file and accessible to someone intent on identity theft.

The New Scientist has an excellent report on Byers’ research, which has been submitted for publication in the IEEE journal Security and Privacy.

If you post non-HTML documents to the web, how can you make sure potentially compromising information is not included?

The safest way is to convert the document to plain text, then paste the text into a new document. Then, use the “File, Properties” command to see what metadata has been included. This method isn’t foolproof — to be absolutely certain a document doesn’t contain information you don’t want revealed, publish it as a simple HTML file.

More in Industry

View more
Industry

Solving the agency search intelligence gap

How can agencies fill the voids in their search intelligence? Ian O’Rourke, CEO, Adthena, and Stephen Davis, Global Product Leader for Media Intelligence, Kantar shed light on this dilemma.

Industry

What to expect from SEO in 2021?

Want to learn the most important trends to be on the lookout for in 2021? Read our article to find out the top SEO trends for 2021.