ALL IS NAUGHTARCHIVE
ENTRY 004 / ESSAY / APPLIED MODEL

Human-Information Consumption Threshold

When did the sum of written information surpass one person’s lifetime capacity—and when might it surpass humanity’s collective ability to read everything in existence?

Consume information – learn more… Learn – ha. Are we collectively getting smarter, or is knowledge being extracted and reserved from the droves drifting in doom-scrolling doldrums. What demonic conspiracy is afoot to control you? Subscribe to my channel to learn more!

At what point in time could a single human consume all available information? When did the sum of available information surpass humanity’s collective ability to consume it within a lifetime?

Everything extrapolates, the turtles do, in fact, go all the way down; to answer, we must constrain. Here, our construct of information is physically recorded media artifacts and confined to humans. In other words, written work – that we know about. Pam’s passive-aggressive note about cleaning the microwave does… count. Because that is a reference to The Office, where the script is digitally recorded on the Internet, which this work will scrape and include in the total corpus sum. Otherwise, your diary is off limits.

I will leave the fine continuation of this work, as it pertains to artificial intelligence, for my benevolent and immortal digital avatar who will be graciously resurrected out of cosmic necessity some thousand years from now.

Now, do you wish to slide further down those turtles? Kowabu…naught! Kowabunganaught? More constraints!

Physically recorded media objects do not include pictures, architecture, burial mounds, cave paintings, secret terraformed Mars habitats, or Humankind.

Humankind itself is information; we are records of 300,000 years of homo sapiens and each copulatory result. Continue imaging humankind as information and how homo sapiens came to be – what information infinitely precedes the unfathomable epoch culminating in your reading of this line? Do we even have the capacity to comprehend the proverbial beginning?

THE BEGINNING is information I do not have.

We are counting written words.

A BRIEF ETYMOLOGICAL ASIDE ON… INFORMATION

That’s a splash-brief on the word information, but we’re opening our arms to corpora in its entirety

First – a note on Zipf’s Law; which tells us that there is an inverse relationship between proportion and rank, expressed as:

(f ∝ 1/r)

Zipf, a linguist, observed this phenomenon: the most common item will occur twice as often as second most common and so on. When examining an extensive corpus we find, unsurprisingly, grammatically functional words (e.g. articles) dominate.

Thus expanded, what are the mode-words, translated to English, for available digitized corpora?

The following analyses were conducted with a dataset of our own query and comprised of 142 languages and 43 major language groupings, spanning from 2690 BC to present 2026 AD. The records examine each group’s earliest and latest attestation window and identified the mode-word translated to English. The data includes a validation cycle handled through a tiered confidence system (High to Low), supported by corpus references, historical dictionaries, epigraphic sources, available corpora, and caution flags for sparsity or bias.

Mode words were found to be: King, Person, and Say in this dataset.

Mode-Word Analysis

Mode-word analysis across language-family attestation dates
Mode-word analysis: person dominates genealogical diversity; say dominates exposure by speaker population.
1

Person, Say, King… Man, TV… Real Stable Genius :/

We’ve created a graph. We make deductions.

In the top six world languages: English, Mandarin Chinese, Hindi, Spanish, Arabic, and French “say” is mode word. This is confounding as the visual conclusion to Mode Word Analysis is that there was a shift to person becoming the dominant word.

Although the branch-level distribution suggests that person is the most common mode-word, this result is driven by the large number of independently sampled language branches exhibiting person-centered lexical frequency patterns. If we weight by speaker-population rather than lineage-count, the result reverses.

The world's largest contemporary language groupings, including Indo-Aryan, Romance, Germanic, Slavic, and Arabic, are all "say" systems. Thus, person dominates genealogical diversity, whereas say dominates human linguistic exposure.

<Screams in Occam>

Person is the modal word across language families; say is the modal word across speakers.

A BRIEF ETYMOLOGICAL ASIDE ENDS

Alright, you’re 729 words in at 238 WPM = 3.06min. Then we’ll add a 10% Oddity Penalty and 10% allowance to read the graph = 3.06*1.2 = 3.6min.

Man, it took me way longer to write all this.

Further approaching our initial question let us examine total corpora over time.

Unique Languages and Sum Corpora

Unique languages and cumulative corpus over time
Unique languages and estimated cumulative written corpus by first attestation.

Where then, respectively, did a single human and humanity-collective lose their ability to read the cumulative corpus in their lifetime?

We will answer this question with an applied model.

First, a review of variables.

t year

P(t) population

L(t) literacy rate (0–1)

E(t) life expectancy

A(t) reading start age

R(t) reading speed (words/min)

H reading hours/day

Individual Capacity: K(t)

Words one literate person could read in an entire lifetime.

K(t) = (E(t) − A(t)) × 365.25 × H × 60 × R(t)

Simplified — with H = 12, the three constants collapse to 365.25 × 12 × 60 = 262,980

K(t) = 262, 980 • (E(t) − A(t)) • R(t)

 

Collective Capacity: G(t)

Aggregate lifetime reading capacity of all living literate people. 

G(t) = P(t) • L(t) • (E(t) − A(t)) × 365.25 × H × 60 × R(t)

Compact form, expressed through the individual capacity K(t)

G(t) = P(t) • L(t) • K(t)

Corpus: C(t) 

Constructed using empirical estimates and log-linear interpolation between anchor points. For any year t between anchors (tᵢ, Cᵢ) and (tᵢ₊₁, Cᵢ₊₁)

C(t) = 10log10 Ci + (log10 Ci+1 − log10 Ci) · ttiti+1ti

 

This is a frictionless model: There are no geographic, cultural, linguistic, or technological constraints. In reality – the Egyptians and Mesopotamians were existing simultaneously (cool, dude) and writing, simultaneously, though they were not cross-consuming. Boundaries and access erode with technology, though today not every written work is digitized and universally translated… Further, estimates of populations’ internet connectivity – much less what people are using it for – is between 70-80%. That means some number beyond 2 billion people cannot reach www.allisnaught.com to receive the answer to this important question.

Human Reading Horizon

Human Reading Horizon comparing corpus size with individual and collective reading capacity
Human Reading Horizon: modeled individual crossover near 467 CE and projected collective crossover around 2049.

  1. Around the year 467 we lost the ability for an individual to consume all written text

  2. We are within a lifetime of losing our collective ability to read everything in existence

Practically speaking, the individual succumbed to the corpora centuries before based on geographic and linguistic constraints. Collectively, we’d have to make deeper considerations for language specific abilities, corpora, and so on.

We salute those humans, nay, memoric ghost-titans, who came to be from the same eternal spark from which you and I derive our consciousness.

In closing, our title and aside point at information, though we’ve answered in terms of reading. The two aren't identical, but they rhyme: reading is just the throughput valve on the pipe marked information. Widen the flow to include image, audio, code, DNA, whatever my benevolent digital avatar decides to count a thousand years from now — and the crossover dates will shift, but the shape of the story won't. Consumption capacity grows arithmetically, production grows exponentially, and each curve only ever gets to cross once.

We already lost the individual race somewhere between the last scroll finished at Alexandria and the rest of us not noticing. The collective forfeit around 2049 is fast approaching unless we can figure out how to read faster, or stop posting meaningless content – I’m thinking it too, don’t worry. Person, Say, and King were never going to read it all either; they just got quoted the most while the rest of us tried.

So: subscribe, or whatever. Learn more. You've got, generously, until 2049.


  1. Distribution of first attestations by semantic category among identified corpus mode words. Histograms show observed frequencies, rug marks indicate individual language-family attestations, and kernel density estimates illustrate overall temporal trends. Person-related mode words are concentrated in later attestations, whereas say and king are disproportionately represented in ancient literary and inscriptional corpora.↩︎