Research Program · Neurolabs R&D · Symbolic Dynamics & Representation Theory
Open Questions
The EOA program poses 43 open questions across 10 domains. They are genuine open problems, not rhetorical questions. Some may have short answers. Others may require new mathematical frameworks. A few may turn out to be ill-posed, revealing that our underlying assumptions are wrong. All of them are real.
Invitation. These questions are not solvable by a single researcher or a single discipline. We invite mathematicians, phoneticians, computer scientists, cryptographers, and representation theorists to collaborate. See the home page for contact information.
Domain 2 · Operator
The class, structure, and dynamics of the operator β.
Q2.1 What operator class produces this convergence: substitution system, linear recurrence, contraction mapping, or none of these?
Q2.2 Why does a least-squares linear fit of the recurrence matrix fail? Where does the nonlinearity actually live?
Q2.3 Would random symbols converge to a shared ratio, or is alphabet-specific structure required?
Q2.4 Is the transient regime monotonic, oscillatory, or chaotic?
Domain 3 · Limiting Ratio
The limiting constant ratio (LCR) and its dependence on the phonetic encoding.
Q3.1 Is λ a computable function of the encoding? Can one predict λ without running β?
Q3.2 What properties of the encoding determine λ: mean letter-name length, character-frequency distribution, phonetic-feature profile, or graph-theoretic structure?
Q3.3 Which phonetic-feature profile is most predictive of λ, and does it vary across languages?
Q3.4 The English modified #1 encoding shifts λ from 3.306 to 3.037. Is there a continuous mapping from phonetic-space to λ-space? Can we design an encoding whose LCR approximates a target value to arbitrary precision?
Q3.5 Is there a universal ceiling on LCR, or can an encoding push it arbitrarily high?
Domain 4 · Sequence-Group Structure
The collapse from 26 letters to 16 distinct sequence groups and its algebraic origin.
Q4.1 Why does the 26-letter alphabet collapse to exactly 16 distinct sequence groups? Is this number a function of the alphabet size?
Q4.2 What property of the letter-name strings, not the phonemes, determines membership in a group?
Q4.3 If we added a 27th letter to English, would the group structure grow to 17, or would the new letter slot into an existing group?
Q4.4 Are the 13 sequences that form the basis for Q₂ (English modified #5) also a basis under other encodings?
Domain 5 · Dimensionality Collapse
The apparent rank-1 collapse of word vectors built from the sequence groups.
Q5.1 What is the precise, metric-backed definition of the apparent rank-1 collapse? Exact rank 1, numerical rank 1 under a specified tolerance, or an artifact of the construction formula?
Q5.2 Is the collapse a bug or a design feature? A limitation to overcome or a structural property to exploit?
Q5.3 What information survives in the small deviations from rank 1, and can that information be recovered?
Q5.4 Can the collapse be exploited for adversarial robustness, where perturbations orthogonal to the line are harmless?
Domain 6 · Transient Structure
The letter-specific early terms (1–38) and what information they carry.
Q6.1 What information survives in the transient regime, and how much of it is letter-specific?
Q6.2 Can the transient be detrended by λⁿ to isolate residual letter-specific structure?
Q6.3 Is the transient monotonic, oscillatory, or chaotic across different encodings?
Q6.4 Could the transient structure serve as a deterministic fingerprint for individual letters?
Domain 7 · Cross-Linguistic Variation
How the phenomenon behaves across scripts, languages, and synthetic alphabets.
Q7.1 Do non-alphabetic scripts (syllabaries, abugidas, logographies) show analogous convergence? Can β even digest the names of hiragana, katakana, or Devanagari characters?
Q7.2 Is the tight Semitic cluster (Arabic 2.751, Hebrew 2.710) a structural property of abjads, or just a small-sample artifact? Could it come from consonantal roots or vowel omission?
Q7.3 Would a synthetic alphabet with no phonetic history still produce a basis and an LCR? If yes, the phenomenon is purely formal. If no, the phonetic grounding actually matters.
Q7.4 Is there a relationship between alphabet size and λ (English 26, Hebrew 22, Greek 24)? How would we test it properly with a larger sample?
Domain 8 · Applications
Downstream uses of the EOA representation in NLP, compression, cryptography, and machine learning.
Q8.1 Can the EOA representation be used in low-resource settings where training data is unavailable?
Q8.2 Is there any downstream task where a low-rank, semantically empty, but algebraically rigid representation is actually useful (e.g., adversarial robustness, reproducible benchmarking)?
Q8.3 Can the 13-dimensional basis be used as a fixed codebook for deterministic source coding?
Q8.4 Can the LCR be exploited for deterministic pseudorandom number generation?
Q8.5 Can the sequence-group structure be used for lossless compression of letter-name dictionaries?
Q8.6 Can EOA-43 act as a zero-training structural prior or fixed initialization for learned embeddings, reducing trainable parameters?
Q8.7 How does the EOA representation compare to character trigram hashing and random vector baselines on standard benchmarks?
Domain 9 · Reverse Design and Targetability
Using the encoding-to-LCR mapping to design encodings with desired properties.
Q9.1 Can we search for encodings whose λ approximates π, e, or the golden ratio to arbitrary precision?
Q9.2 What constraints govern inverse design? Must the encoding stay phonetic and bijective?
Q9.3 Can an encoding be designed so its sequence terms encode primes at specific positions?
Q9.4 How large is the space of bijective encodings for a 26-letter alphabet, and can we search it efficiently?
Q9.5 What would a targetable λ actually enable? A representation tuned to π might possess periodic properties; a representation tuned to the golden ratio might exhibit self-similarity.
Domain 10 · Epistemology of Staged Disclosure
The philosophy of science and priority questions raised by disclosing the phenomenon but not the mechanism.
Q10.1 Is it legitimate to publish empirical observations of a phenomenon whose generative mechanism remains undisclosed?
Q10.2 What precedent exists for staged disclosure in mathematics and the sciences?
Q10.3 What precedent exists for staged disclosure in mathematics (RSA-129, AlphaFold)?
Q10.4 How should priority be adjudicated when the phenomenon is public but the mechanism is not?
Domain 11 · Structural Hints and Unseen Dynamics
Preliminary observations of a recursive, tree-like geometry that may underpin the sequence-group collapse.
Q11.1 Does the recursive generation exhibit a self-similar branching structure that converges to a finite closed set yet generates infinite descendants, and can its geometry be formally characterized?
Q11.2 Can the tree-like propagation be characterized by a finite grammar (e.g., L-system), and would such a grammar explain the 16-element group structure and the dimensionality collapse?
Summary
43 open questions across 10 domains. They range from operator classification (Domain 2) through encoding dependence (Domain 3), group structure (Domain 4), rank collapse (Domain 5), transients (Domain 6), cross-linguistic variation (Domain 7), applications (Domain 8), reverse design (Domain 9), disclosure (Domain 10), and structural hints (Domain 11).
The questions are open. We invite the community to answer them.