← Back to Press

The Alphabet Code: What Happens When You Run a Simple Rule on the Letter “a”

Bahaa Budargham · Engineering of Alphabets project

Montréal, Canada · October 2025

Here is a string of letters:

gzhlvzggjzyihjgheqvzgzhlvzhlvbjzgyhzizygjzbjzhlvgjzehjdzzgheqv...

It looks like a cat walked across a keyboard. It isn’t. It’s the start of a sequence that runs to thousands of characters, and every letter is fixed in advance by one rule applied to the letter “a.” No randomness, no guessing, no training data. Run the same rule again tomorrow, or on another computer, and you get exactly the same string.

The sequence for the letter “a” drawn as a path: each letter tells a virtual pen how to turn and move.
Figure 1. The sequence for the letter “a” drawn as a path: each letter tells a virtual pen how to turn and move, so a fixed rule produces this shape.

That simple question — what happens when one fixed rule is applied to “a” again and again? — led me into years of studying sequences like this one. I want to show you what I’ve found, what I still haven’t figured out, and why I think a curious person might care.

The idea in one minute

Start with the letter “a” and apply the same fixed rule to it repeatedly. Each round replaces the existing letters with new strings of letters, so the sequence grows. At the same time, the earlier rounds remain embedded within the later ones.

The first seven rounds have 1, 4, 14, 43, 138, 429, and 1361 letters. Mathematicians call this a substitution system.

The unusual feature of the Engineering of Alphabets project is that the rule is applied to the alphabet itself. When I say “alphabet” here, I mean the 26 English letters together with their spoken names. A different language, or a different naming convention, would give a different sequence. I wanted to see whether the letters we use for writing, treated as a mathematical system, produce structures that are interesting, unexpected, or perhaps useful.

How I did it. I wrote a short computer program that applies the rule and records the output. For the results below, I generated several thousand letters — enough to count blocks up to length 30 and to estimate the growth factor. I used the same program to compute the growth factor for each of the 26 starting letters. The degree-13 equation came from fitting the term lengths to a linear recurrence, then factoring the resulting polynomial. A colleague or reader can rerun the program on the public data; the exact rule itself is withheld for now.

What I found

Three observations stand out. I’ll be careful to say how solid each one is.

Distinct blocks p(k) by block length up to 30, shown on a logarithmic vertical axis.
Figure 2. Distinct blocks p(k) by block length up to 30, shown on a logarithmic vertical axis. Dashed orange is a linear reference; dashed blue is the Sturmian minimum (k + 1).

What I am not claiming

I think this matters as much as the findings, so here it is plainly:

Open Questions

Here is my challenge to you. See whether you can answer any of these. If you crack one, or show that I’ve got it wrong, I want to hear about it.

  1. What is the equation? The growth factor is the dominant root of a degree-13 polynomial with integer coefficients. What does that polynomial mean? Is it new, or does it match something already known?
  2. Is there a simpler rule hiding inside? Certain short stretches keep coming back. Do those recurring pieces follow a smaller, simpler rule of their own?
  3. Does the straight line hold? The number of distinct blocks grows roughly in a straight line over the lengths I’ve checked. Does it keep going, or does it bend at longer lengths? Only deeper computation can say.
  4. Why 13 of 26? Thirteen starting letters — a, d, e, g, h, i, k, l, q, u, w, y, z — have the same degree-13 minimal polynomial:
    4x13 − 21x12 + 45x11 − 74x10 + 98x9 − 34x8 − 116x7 + 210x6 − 152x5 + 26x4 + 38x3 − 30x2 + 9x − 1
    The rest of the letters have a degree-14 polynomial, which is exactly this one multiplied by (x − 1). So all 26 share a common structure. The 13 listed above are the ones where the extra factor cancels out. Why? Does that split mean something about the alphabet, or is it an artifact of how the rule was set up?
  5. Why is the growth factor so close to π? The closeness is striking, but the number is algebraic and cannot be π. Is there a mathematical reason for the resemblance, or is it a coincidence?

Why This Might Matter Beyond Pure Math

These are possible directions, not demonstrated applications. None has been tested yet.

You might be thinking: this is interesting, but what does it have to do with anything? Here are the directions I think are worth exploring.

No compression, OCR, speech, or language-modeling experiment is reported here.

The Bigger Picture

The Engineering of Alphabets is a research program that studies what happens when you apply deterministic rules to alphabets. The letter “a” is just the beginning. The project includes 13 main sequence families, plus several related cases. Each one has its own structure, properties, and open questions.

The alphabet can be treated as a mathematical object, not only as a tool for writing. Whether that view leads anywhere is what the open questions are for.

Try to break it

Science works when people try to prove each other wrong, so here is what I’d ask of you.

The data and part of the code are public. The code that generated the sequence is not public, but the code that counted the blocks and computed the growth factor is. The exact letter-by-letter table behind the rule is shared under a short written agreement, because a first formal paper is still in preparation. That is a timing decision, not a permanent lock, and the open questions can be worked from the public data alone.

You can reach me at Bahaa Budargham. ORCID: 0009-0005-2955-3703. Zenodo · GitHub.