Analyzing Etruscan Language Origins Using Persistent Homology

AJAS · 2026

Overview

Language data is often difficult to analyze because of its high-dimensional nature, which causes an exponential growth of the size of the space in which the data lives as the dimension increases. This, coupled with sparsity of data (texts) from which to draw, has presented long-term challenges in comparing ancient languages concretely and holistically. This study analyzes a framework for overcoming this barrier through a quantitative method, topological data analysis (TDA), while also applying it to ongoing investigations of the isolate ancient Etruscan in an effort to understand its much-debated origin. TDA provides higher-level information about the inherent structure of a language’s phonology, and with the aid of one of its tools, persistent homology, one can compare different phonologies to determine which languages are most closely related. Results indicate that TDA successfully detects patterns unique to language data and is able to provide a greater distinguishability between languages than other methods. Comparing Etruscan to both Indo-European and other languages suggests that it may have been a relative of the Sanskrit or Semitic languages, explaining its singularity compared to the neighboring Latin and Greek.

Competition history

  • AJAS 2026 Category not listed

Related projects

Closest projects by meaning, across every fair and year in the corpus.

Browse more like this

Source: AAAS Annual Meeting (Confex) / American Junior Academy of Science

Save projects to your library

Sign in with Google to keep track of projects you find interesting, organized into folders. An account also raises your daily allowance for “Has this been done?”, and lets you create a key for the MCP server with a much higher limit than anonymous use. Browsing stays public.

Continue with Google