The Era of Information Junk

By: Stephanie Shepard

Anyone who has cleaned out the house of an elderly relative knows that finding something and learning something are not necessarily the same thing. You open a closet expecting family history and discover thirty years of thrift store clothes. The next cabinet contains old utility bills, extension cords, broken electronics, and a collection of NASCAR commemorative plates that apparently represented a serious financial commitment at some point.

You cannot immediately throw everything away, either. Somewhere underneath the clothes might be Grandpa’s address book from the war, an old photograph nobody has seen, or a letter that explains a piece of family history everyone has wondered about for fifty years. So you keep digging while asking the question that eventually occurs to anyone cleaning a pack rat’s house: Why did you save all of this?

 

The information age has created essentially the same problem.

When Information Was Hard to Get

A century ago, Herbert O. Yardley faced almost the opposite problem. Yardley was one of the pioneers of American signals intelligence and ran the Cipher Bureau, commonly remembered as the American Black Chamber, after World War I. His problem wasn’t an abundance of information. The information he wanted was deliberately hidden from him.

Foreign governments sent diplomatic communications through telegraph and cable networks, but important messages were encrypted. Yardley’s cryptanalysts therefore had to acquire the communications, identify the cryptographic systems protecting them, break the codes, translate the messages, and determine which information mattered enough to pass along to American officials.

The Washington Naval Conference demonstrated the value of that process. While diplomats publicly negotiated limitations on naval power, Yardley’s operation was secretly reading Japanese diplomatic traffic. Information that Japanese negotiators believed was concealed could become intelligence for American decision makers.

Yardley’s world was built around a fundamental problem of information scarcity: How do we obtain information we aren’t supposed to have?

Then everyone started saving things.

The House Starts Filling Up

The intelligence system Yardley helped pioneer didn’t disappear when his Black Chamber closed. It grew enormously. World War II generated mountains of military and intelligence records, while the Cold War created new agencies, new classification systems, new surveillance technologies, new investigations, and new filing cabinets to contain everything they produced.

Technology made copying easier at the same time. A document could have an original, a carbon copy, a photocopy, a microfilm copy, a scanned copy, a PDF copy, a database copy, an archived web copy, a screenshot, and eventually a social media post showing a photograph of the screenshot.

Historians wrote books about the documents. Journalists wrote articles about the books. Government investigators wrote reports about the events described in the documents, and later congressional committees wrote reports examining the earlier government investigations. Eventually those records were declassified, digitized, uploaded, copied, discussed, quoted, reposted, and summarized.

The house kept filling up.

That accumulation creates a strange problem whenever governments announce major declassification releases. The public understandably expects classified records to contain secrets and sometimes genuinely important new information does emerge. But classified and unknown are not synonyms.

A large archival release can contain previously released documents, different versions of the same document, administrative correspondence, handwritten notes, routing slips, degraded reproductions, and the archival classic: a photocopy of a photocopy of something somebody typed decades ago. A release containing tens of thousands of pages can therefore contain considerably fewer than tens of thousands of pieces of new information.

This is where we need another category.

Information Junk

Misinformation is inaccurate information. Disinformation is deliberately deceptive information, while noise makes useful signals harder to identify. Information junk is something slightly different because the information itself can be completely authentic.

Information junk is authentic or potentially useful material whose presence adds little additional informational value to the question being investigated.

The NASCAR plates are real. Grandpa genuinely bought them, displayed them, and apparently believed future generations would appreciate receiving them. We simply don’t need forty seven plates to establish that Grandpa liked NASCAR.

The same problem occurs in historical research. Ten documents repeating the same allegation can initially appear to represent ten pieces of evidence. Once their provenance is reconstructed, however, researchers might discover that nine documents ultimately derive from the tenth. The information multiplied, the evidence didn’t.

This becomes especially troublesome after decades of accumulated commentary. One newspaper quotes an anonymous government official. Another newspaper cites the first report, while a historian later cites both newspapers. A documentary interviews the historian, an article summarizes the documentary, and someone eventually posts the article online as confirmation that several independent sources support the original claim.

The house now contains seven NASCAR plates and they’re all reproductions of the same Dale Earnhardt plate.

Don’t Throw Away Grandpa’s Address Book

The obvious response would be to start throwing things away, but that creates another problem. Information junk is contextual. An FBI routing slip might contribute almost nothing to determining whether a historical event occurred. If the question changes to how quickly information traveled between FBI field offices and headquarters, that previously useless routing slip could suddenly become important evidence.

The same principle applies to handwritten annotations, duplicate documents, envelopes, timestamps, mailing addresses, and administrative records. Something that looks like junk while answering one question might contain the signal needed to answer another.

That means cleaning the information house cannot mean indiscriminately filling a dumpster. Someone has to inventory it first. This is where artificial intelligence may eventually become particularly useful to historians and researchers. One of AI’s most valuable contributions to an age of information abundance may not be creating more information. We have managed that perfectly well on our own.

AI can help sort what already exists.

A sufficiently capable research system can compare thousands of documents, identify likely duplicates, trace citations backward, construct timelines, recognize recurring names, distinguish original reports from derivative accounts, locate contradictions, and help determine whether ten apparently independent claims actually originated with one source.

Instead of arriving at Grandpa’s house with a dumpster, AI arrives carrying boxes.

ORIGINAL. DUPLICATE. DERIVATIVE. CONTRADICTORY. ADMINISTRATIVE. PROVENANCE UNKNOWN. POSSIBLY IMPORTANT.

And somewhere near the garage there is inevitably another box labeled NASCAR PLATES.

The human researcher still has to decide what matters. AI can tell us that seventeen documents appear to contain the same information, but deciding whether that information changes our understanding of history requires judgment. It can help us locate Grandpa’s address book, but it cannot decide why Grandpa’s address book matters to the story we’re trying to reconstruct.

The objective isn’t to throw history away. It’s to make the architecture of the house visible again.

Yardley’s Problem in Reverse

This creates an amusing symmetry between Yardley’s information world and ours.

In the 1920s, Yardley sat in a relatively small intelligence operation trying to extract hidden information from an emerging global communications network. He needed access to messages other people didn’t want him to read and cryptanalysis helped solve the problem.

A century later, we have accumulated government archives, intelligence records, congressional investigations, memoirs, newspapers, academic histories, digital databases, declassified files, websites, social media, and practically limitless commentary about them. The information that Yardley’s generation struggled to acquire has been joined by a hundred years of additional information about the information.

His problem was scarcity. Ours is abundance.

The central intelligence question has therefore undergone an interesting reversal. Yardley had to ask how he could get the information, while the twenty first century researcher increasingly has to ask how to see through all of it. Cryptanalysis helped solve the first problem. Artificial intelligence may become one of the tools that helps with the second, provided we remember that sorting information is not the same thing as understanding it.

Yardley entered the room when it was nearly empty. Over the next hundred years, governments, wars, intelligence agencies, historians, journalists, photocopiers, computers, and the internet kept carrying things inside until the stacks reached the ceiling.

Now AI has arrived at the door carrying boxes, while the rest of us stand in the hallway trying to remember where we left the floor.

The first question is obvious.

Where do you want the NASCAR plates?

••••

The Liberty Beacon Project is now expanding at a near exponential rate, and for this we are grateful and excited! But we must also be practical. For 7 years we have not asked for any donations, and have built this project with our own funds as we grew. We are now experiencing ever increasing growing pains due to the large number of websites and projects we represent. So we have just installed donation buttons on our websites and ask that you consider this when you visit them. Nothing is too small. We thank you for all your support and your considerations … (TLB)

••••

Comment Policy: As a privately owned web site, we reserve the right to remove comments that contain spam, advertising, vulgarity, threats of violence, racism, or personal/abusive attacks on other users. This also applies to trolling, the use of more than one alias, or just intentional mischief. Enforcement of this policy is at the discretion of this websites administrators. Repeat offenders may be blocked or permanently banned without prior warning.

••••

Disclaimer: TLB websites contain copyrighted material the use of which has not always been specifically authorized by the copyright owner. We are making such material available to our readers under the provisions of “fair use” in an effort to advance a better understanding of political, health, economic and social issues. The material on this site is distributed without profit to those who have expressed a prior interest in receiving it for research and educational purposes. If you wish to use copyrighted material for purposes other than “fair use” you must request permission from the copyright owner.

••••

Disclaimer: The information and opinions shared are for informational purposes only including, but not limited to, text, graphics, images and other material are not intended as medical advice or instruction. Nothing mentioned is intended to be a substitute for professional medical advice, diagnosis or treatment.

Be the first to comment

Leave a Reply

Your email address will not be published.


*