Full chapter text
Recognition
The technological history of Maithili cannot be reduced to the arrival of computers or the spread of
mobile phones. For a language with a historic script, the decisive question is whether text can be represented
as text rather than merely displayed as pictures of letters. Tirhuta, also called Mithilakshara and Maithili
script, illustrates the difference. Before standard encoding, users could design attractive fonts and produce
print or PDF pages, yet the underlying data often depended on private mappings to Latin, Devanagari or
other code positions. Such material could look correct on one machine while remaining unsearchable,
unexchangeable and fragile. Unicode recognition changed the conditions of possibility by assigning stable
identities to Tirhuta characters, but it did not automatically create fonts, keyboards, literacy, software
support or archives.
This chapter therefore treats Unicode encoding as one stage in a longer chain of technological
recognition. The official Unicode block for Tirhuta occupies U+11480–U+114DF, and the script entered
the Unicode Standard in Version 7.0 in 2014. The mature encoding proposal prepared by Anshuman Pandey
documented both the script’s historical independence and contemporary digital use, including periodical
production in Tirhuta and the limitations of legacy encodings. The proposal itself is a valuable historical
document because it records a moment when a regional writing tradition had already entered desktop
publishing but had not yet obtained interoperable plain-text representation. The history after encoding is
equally important: fonts, shaping engines, keyboards, web standards, converters, teaching and preservation
determine whether encoded characters become a living digital medium.
116.1 Script, language and technology are different historical categories
A language, a script and a technical encoding are related but not identical. Maithili can be written in
Devanagari or Tirhuta; Tirhuta can be studied as a script even when many contemporary Maithili readers use
Devanagari; and Unicode can encode a script without determining which script a community should prefer.
Confusing these levels creates false narratives in which technological recognition is treated as linguistic
recognition or in which script revival is assumed to require rejection of another script. A more precise history
asks what each layer enables. Constitutional or educational recognition affects institutions; script practice
affects literacy and cultural memory; character encoding affects digital interchange. Their timelines overlap
but do not coincide.
116.2 Tirhuta as a writing tradition of Mithila
Tirhuta belongs to the historical graphic culture of Mithila and carries associations with manuscripts,
genealogical records, ritual documents, scholarship and literary transmission. Its historical importance is
therefore larger than the number of people who type it daily. Scripts store visual habits and conventions:
letter forms, conjuncts, punctuation, numerals, scribal abbreviations and relationships between handwriting
and print. Digital revival that preserves only an alphabet chart but ignores these practices risks converting a
writing tradition into a decorative emblem. Conversely, a usable encoding allows the script to move beyond
symbolism into searchable editions, correspondence, teaching materials and born-digital composition.
11711171
GAJENDRA THAKUR
116.3 Devanagari dominance and the problem of technological substitution
The twentieth-century expansion of Devanagari for Maithili created a large modern reading public and
simplified access to printing systems already built for Hindi and other languages. Technologically, however,
this success could make Tirhuta invisible because software vendors tend to support scripts in proportion to
existing demand. The result is a feedback loop: readers use the script that devices support, and companies
interpret that use as proof that unsupported scripts are unnecessary. Unicode breaks only part of this loop by
giving Tirhuta a standard identity. Practical parity still requires fonts, keyboards and applications.
116.4 Early digitisation through legacy fonts
Before Unicode encoding, users developed Tirhuta computer fonts by assigning Tirhuta glyphs to code
positions intended for other characters. This was a common strategy for minority scripts in early desktop
publishing because it allowed documents to be composed and printed without waiting for international
standards. The visual achievement was real: books, periodicals and PDFs could appear in Tirhuta. Yet the
method tied meaning to a particular font file. Remove the font, paste the text elsewhere or index it with a
search engine, and the apparent Tirhuta could turn into unrelated Latin or Devanagari characters.
Digitisation had occurred at the image layer without full semantic interoperability.
Figure 460 — From legacy fonts to interoperable Tirhuta text.
116.5 The hidden cost of visually correct but semantically wrong text
A legacy-font document may be typographically beautiful while remaining computationally misleading.
Copying a word can produce nonsense; alphabetical sorting follows the substituted code points; screen
readers pronounce the wrong characters; search cannot find expected words; and archives cannot reliably
identify the script. The problem becomes acute during migration because a future curator may possess the
document but not the exact font or mapping table. Unicode’s central historical contribution is therefore not
aesthetic. It separates abstract character identity from the visual glyph chosen to display it, allowing the same
text to survive changes in fonts and platforms.
116.6 Governmental and community interest in standard encoding
The path to encoding combined user-community pressure, scholarly documentation and institutional
interest. The 2011 Tirhuta proposal records that India’s 2004 scheduling of Maithili renewed attention to its
HISTORY OF MITHILA, VAJJI & ANGA — VOLUME II
traditional script and notes a 2005 presentation to the Unicode Technical Committee by Om Vikas of the
Government of India’s Department of Information Technology. Such episodes matter because international
standards do not arise automatically from cultural antiquity. Someone must document the repertoire,
establish usage, resolve naming and encoding questions, and carry the proposal through technical review.
Standardisation is thus a form of cultural labour as well as engineering.
116.7 The Unicode proposal as historical documentation
An encoding proposal is unusually rich evidence for historians of technology. It must show how a script
behaves, not merely what letters look like. The Tirhuta proposal documents vowels, consonants, signs, digits,
conjunct behavior, similarities and differences with neighboring scripts, and examples of contemporary use.
It also records the technological problem that motivated encoding: legacy fonts could display Tirhuta but
could not provide stable plain-text representation. For cultural history, this makes the proposal a bridge
source linking manuscripts and traditional orthography to desktop publishing and international character
standards.
116.8 Why Tirhuta could not simply be unified with Bengali
Visual resemblance does not justify treating two scripts as one encoding. The Tirhuta proposal
specifically argued against unification with Bengali even though several forms look similar. Characters that
resemble one another can behave differently in vowel combinations, conjuncts and orthographic sequences.
Unicode encoding must represent character identity and text behavior, not merely graphic appearance. This
principle is culturally important because careless unification can erase distinctions that users themselves
consider meaningful. Technical classification therefore participates in the recognition of a script as an
independent system.
116.9 Character repertoire, orthography and conjunct behavior
A workable encoding requires decisions about the core repertoire and the behavior of combining signs.
Tirhuta, like other Brahmic scripts, cannot be represented adequately by listing isolated alphabet forms.
Vowels may be independent or dependent; consonants combine; virama-mediated sequences create
conjuncts; and marks interact spatially with base letters. Font shaping then converts these character
sequences into appropriate glyph arrangements. This separation between encoded sequence and rendered
form is what makes multiple typefaces possible without changing the underlying text. It also explains why a
font can contain all nominal characters yet still render real Maithili poorly if shaping rules are incomplete.
116.10 Unicode 7.0 and the 2014 encoding milestone
Tirhuta entered the Unicode Standard with Version 7.0 in 2014, a release that added twenty-two scripts
and substantially expanded the standard. For Maithili digital history, 2014 is a major interoperability date
rather than a date of script invention or revival. The script had centuries of history and decades of modern
advocacy before encoding. What changed was that software could now identify Tirhuta characters by
internationally stable code points. This enabled standards-conforming fonts, text exchange, databases and
web content to treat Tirhuta as text in its own right.
11731173
GAJENDRA THAKUR
Figure 461 — The technological recognition stack.
116.11 The Tirhuta block U+11480–U+114DF
The Unicode Tirhuta block spans U+11480 through U+114DF in the Supplementary Multilingual
Plane. The chart includes letters, vowel signs, marks and Tirhuta digits. Code-point ranges are not merely
technical trivia: they make it possible to validate data, detect script usage, create keyboard mappings and
construct archival search tools. Because Tirhuta lies outside the Basic Multilingual Plane, older software
designed around sixteen-bit assumptions sometimes created additional implementation problems. Modern
Unicode-aware systems normally handle supplementary-plane characters correctly, but legacy software can
still expose the historical layers of the computing environment.
116.12 Encoding is not the same as a font
Unicode assigns characters; fonts draw them. This distinction is one of the most important practical
lessons of script technology. A device may store perfectly valid Tirhuta text and still show blank rectangles if
no suitable font is available. Conversely, a legacy font may display Tirhuta-shaped glyphs while storing the
wrong characters. Sustainable support therefore requires both correct encoding and quality fonts. Testing
must include common words, conjuncts, vowel signs, punctuation and mixed-script contexts rather than
only isolated alphabet charts.
116.13 Open fonts and the importance of shaping quality
Openly distributable fonts reduce one of the largest barriers to script use. A font that can be embedded in
websites, EPUBs and documents allows readers to see Tirhuta without manually installing proprietary
software. Yet font availability is not enough: shaping quality, vertical metrics, line spacing and rendering
consistency across browsers matter. The emergence of Unicode Tirhuta fonts, including the Noto family,
converts encoding into visible access. For preservation, font licensing is also significant because an archive
should be able to retain the typeface needed to reproduce a document without depending on an inaccessible
commercial license.
HISTORY OF MITHILA, VAJJI & ANGA — VOLUME II
116.14 Keyboards and input methods as cultural infrastructure
Typing is where many encoded scripts encounter their next bottleneck. Users need mappings that are
learnable on desktop and mobile devices, and different communities may prefer phonetic transliteration,
InScript-like layouts or traditional character positions. A keyboard is not culturally neutral: it determines
which signs are easy to produce and can encourage simplified spelling if difficult sequences are buried. Good
input systems therefore need documentation, visible layouts and testing by fluent users. Training materials
should teach both the script and the digital method of entering it.
116.15 Mobile typing and the politics of practical availability
For many younger users, the phone rather than the desktop is the primary writing device. A script that
works only through a specialist computer keyboard remains technologically marginal even after Unicode
encoding. Mobile operating systems, third-party keyboards, messaging applications and browser fonts
determine whether Tirhuta can participate in everyday communication. Practical recognition occurs when a
reader can receive a Tirhuta message, reply to it, search it and paste it into another application without
corruption. Mobile support is therefore a social-history issue because it determines whether the script enters
routine interpersonal use or remains confined to ceremonial and scholarly contexts.
116.16 Conversion between Devanagari and Tirhuta
Automatic conversion between Devanagari Maithili and Tirhuta can dramatically enlarge the available
corpus because contemporary Maithili publishing is predominantly Devanagari. The task, however, is
transliteration between writing systems, not translation between languages. A reliable converter must map
signs and conjunct behavior while respecting Maithili orthography. It should preserve punctuation and
numerals deliberately and mark cases that require human review. Conversion tools are most useful when they
produce genuine Unicode Tirhuta rather than another private font encoding.
Figure 462 — Why Unicode text differs from a legacy-font document.
116.17 Conversion errors as philological errors
Conversion can create a false appearance of certainty. Historical spellings, manuscript abbreviations,
variant conjuncts and ambiguous modern input may not have one mechanically obvious equivalent. If a
converter silently normalises every form, it can destroy evidence that matters to philology. Scholarly
11751175
GAJENDRA THAKUR
workflows should therefore preserve the source text, record the conversion method and distinguish
automatic output from human-verified transcription. A reversible pipeline is preferable: researchers should
be able to trace a Tirhuta digital form back to the source sequence that generated it.
116.18 Search, indexing and machine-readable Tirhuta
Once Tirhuta is represented as genuine Unicode text, it becomes available to ordinary computational
operations: searching, indexing, frequency counting, concordance building and corpus comparison. These
capabilities change research questions. A scholar can locate recurring names across genealogical material,
compare orthographic forms, or build teaching dictionaries from a corpus. Yet search quality depends on
consistent encoding and normalisation. Visually identical sequences can sometimes be represented
differently, so archival systems need explicit data-cleaning rules rather than assuming that what looks the
same on screen is always the same underlying text.
116.19 Web publishing and browser interoperability
Web publication tests the entire technological stack at once. The HTML must declare Unicode correctly;
the page must use a font with Tirhuta coverage; the browser and shaping engine must support the script; and
fallback behavior must not replace characters with empty boxes. Search engines must also crawl
supplementary-plane characters without corruption. Web fonts can solve part of the distribution problem,
but a durable page should remain meaningful even if the original stylesheet disappears. Separating textual
content from presentation is therefore both a web-standard principle and a preservation strategy.
116.20 EPUB, PDF and the difference between reflowable and fixed
representation
Fixed-layout PDF was crucial in the pre-Unicode and transitional period because it could preserve the
intended visual appearance of a Tirhuta page. Its weakness is that text may be image-only or encoded
incorrectly underneath. Reflowable EPUB and semantic HTML demand cleaner Unicode text but provide
better resizing, search and accessibility. Preservation should not force a choice between appearance and
semantics. For historically significant works, the best package may include a faithful page image or PDF
together with verified Unicode transcription and metadata linking the representations.
116.21 OCR and handwritten-text recognition
OCR for Tirhuta remains more difficult than typing already transcribed text because printed and
handwritten sources vary widely in letter forms, ligatures, ink quality and page condition. General OCR
engines are usually trained on larger scripts and languages, so minority-script recognition requires curated
training data and human correction. Handwritten-text recognition is harder still. A responsible archive
should report confidence and preserve page images alongside machine transcription. Incorrect OCR
presented without provenance can turn technological convenience into a new source of textual corruption.
116.22 Tirhuta in education and script revitalisation
Script revitalisation depends on learners, not only standards. Unicode makes it possible to create
interactive primers, typing exercises, searchable readers and digital flashcards in genuine Tirhuta. These tools
can connect palaeographic learning with contemporary communication: the learner moves from recognizing
manuscript forms to typing messages and reading web text. Education also supplies the feedback needed to
HISTORY OF MITHILA, VAJJI & ANGA — VOLUME II
improve fonts and keyboards. If technological development remains separated from teaching, engineers may
solve abstract problems while users still lack a practical path into the script.
116.23 Digital archives, Panji materials and manuscript description
Digital archives can use Tirhuta encoding to describe and transcribe manuscript and genealogical material
without flattening every source into Devanagari. Panji-related records are particularly instructive because
names, place forms, abbreviations and scribal conventions can be historically significant. A digital edition
should distinguish the manuscript image, diplomatic transcription, normalised reading and translation or
transliteration. Unicode enables these layers to coexist, but metadata must record which layer a user is
viewing. The script therefore becomes part of archival provenance rather than an ornamental overlay.
116.24 Videha and user-driven script technology
Videha’s script work belongs within this wider user-driven history of digital adaptation. Long before
universal device support could be assumed, small-language publishers had to assemble fonts, PDFs,
conversion practices and web pages themselves. Later Unicode fonts, browser tools and script-conversion
utilities made it possible to move toward interoperable text. The historical importance of such initiatives lies
in their experimental role: users identify what formal standards still fail to make easy and build local bridges
between cultural demand and general-purpose software. These projects should be documented with version
histories and sample outputs so future researchers can reconstruct technological change.
116.25 Accessibility and script fallback
Accessibility introduces another dimension to script support. Screen readers may recognise Unicode code
points yet lack a high-quality Tirhuta voice; operating systems may fall back to a font with poor shaping; and
users with low vision depend on scalable text rather than image-only pages. A robust publication should
therefore offer semantic headings, selectable text, sufficient contrast and alternative representations where
speech technology is weak. Accessibility does not mean replacing Tirhuta with Devanagari. It means
designing a pathway in which the historic script remains available while users can also access equivalent
content through other representations when needed.
116.26 Unicode normalisation, data hygiene and preservation
Digital preservation requires stable character sequences. Unicode normalisation, duplicate encodings,
invisible characters and accidental mixing of legacy data can create files that render acceptably but behave
inconsistently. Archives should validate script ranges, retain source copies, calculate checksums and record
conversion histories. Text normalization should be applied cautiously because a mechanically ‘cleaner’
sequence may erase distinctions in historical sources. Data hygiene is thus not cosmetic correction; it is the
documentation of how a digital text came to have its present form.
116.27 Names, catalogues and bibliographic interoperability
Catalogues must be able to store Tirhuta titles and names while also supporting readers who search in
Devanagari or Roman transliteration. The solution is not to choose one representation as authentic and
discard the others. A bibliographic record can preserve the title in the script of the item, supply
transliteration, and include normalized access points for discovery. Such parallel metadata is especially
important for cross-border scholarship in India and Nepal and for global library systems. Unicode makes the
original-script field technically possible; cataloguing policy determines whether institutions actually use it.
11771177
GAJENDRA THAKUR
116.28 Technological recognition, prestige and symbolic politics
The arrival of a script in Unicode often acquires symbolic meaning beyond software engineering.
Communities may read encoding as international recognition of historical identity, while governments can
cite it as evidence of cultural support. This symbolism is understandable but should not obscure uneven
material realities. A script can be internationally encoded while remaining absent from school computers,
public forms and ordinary phone keyboards. Technological recognition is therefore both real and
incomplete. Its historical significance lies in opening a standardised space that institutions and users must still
populate.
Figure 463 — A sustainable Tirhuta digital ecosystem.
116.29 What remains incomplete after Unicode
Several tasks remain after encoding: broad mobile input support, mature OCR, tested open fonts,
searchable corpora, teaching materials, archival standards and sustained software maintenance. Minority-
script tools can disappear when a single developer stops updating them or when a hosting service closes.
Durable infrastructure requires open specifications, downloadable source data and more than one
maintainer. It also requires documentation for users who inherit the system later. A project that works only
because its creator remembers undocumented steps has not yet become a stable institution.
116.30 From encoded script to sustainable digital culture
Tirhuta’s digital history shows how cultural preservation and technical standards intersect. Legacy fonts
proved that users would adapt available technology even before formal support existed. Unicode 7.0 supplied
stable character identities; fonts and keyboards converted those identities into readable and writable text;
converters and archives expanded the corpus; and education can turn technical possibility into renewed
literacy. The next stage is not simply more software. It is the creation of a sustainable ecosystem in which
Tirhuta text can be authored, exchanged, searched, taught, cited and preserved without dependence on one
machine, one proprietary mapping or one generation of specialists.
HISTORY OF MITHILA, VAJJI & ANGA — VOLUME II
Table 116.1 — From technological recognition to durable Tirhuta access
Layer What recognition What can still fail Preservation response
provides
Unicode encoding stable character identity no font or shaping support store valid Unicode +
versioned text
Font visible glyphs and missing glyphs / bad metrics archive open font + license
conjuncts
Keyboard / IME practical text entry desktop-only or hard-to- document mappings;
learn layout support mobile
Converter rapid corpus expansion silent transliteration errors retain source; log
conversion; review
Web application publication and discovery browser/font dependencies semantic HTML +
downloadable data
OCR / HTR machine transcription of high error on rare retain images + confidence
scans fonts/handwriting + corrections
Catalogue metadata original-script discovery one-script-only access store Tirhuta + Devanagari
points + transliteration
Archive long-term cultural memory single-server loss / checksums, replicas,
undocumented formats provenance, open formats
PART XIII
THE CONTEMPORARY ECONOMY
Chapters 117–131
11791179