Full chapter text
Chapter 136 examines social media as a linguistic institution rather than merely a communications
technology. Facebook, WhatsApp, YouTube, Instagram, short-video platforms, messaging groups and
emerging AI-assisted interfaces now sit between speech, writing, performance, migration and public debate.
They make some language forms easy to record and circulate, others easy to search, and still others difficult to
type or discover. The result is not a simple shift from “offline” to “online” language. Digital media reorganise
the relative visibility of scripts, spellings, registers, accents, songs, jokes and political vocabularies while
connecting speakers who may live in different districts, states or countries.
The regional frame remains essential. Maithili, Angika and Bajjika do not occupy identical institutional
positions, and Mithila, Vajji and Anga do not form one current administrative unit. India’s Census 2011 C-
16 table makes Maithili directly visible as a named mother-tongue category, while the Bihar table does not
give Angika and Bajjika the same headline statistical visibility. Nepal’s 2021 census, by contrast, separately
records Maithili and Bajjika in Madhesh Province. Platform data introduce another asymmetry: commercial
services may know enormous amounts about users while releasing little language-specific regional evidence.
This chapter therefore combines official language statistics, telecommunications evidence, national internet-
use research and qualitative scholarship without inventing platform-specific counts for the historical-cultural
region.
136.1 Social media is now an infrastructure of everyday language change
Language change once depended heavily on face-to-face interaction, schools, print, broadcasting,
migration and institutions. Social media does not replace those mechanisms; it connects them at much higher
speed. A phrase spoken in a village can be recorded on a phone, circulated through a family group, remixed
into a short video, commented on by migrants in Delhi or the Gulf and returned to the village with a new
spelling or comic association. The crucial historical change is the compression of distance and delay.
Linguistic innovations can travel through weak ties as well as kinship networks, and archived or searchable
posts allow forms to remain visible after the original conversation has ended. Social media therefore operates
as infrastructure: it changes the pathways through which language moves.
136.2 Platforms do not simply carry pre-existing languages
Every platform has affordances and constraints. Text boxes reward what can be typed; search engines
reward what can be spelled predictably; voice notes reduce the need for literacy; video privileges performance,
facial expression and sound; recommendation systems rank some material above other material; monetisation
rewards attention rather than linguistic representativeness. These design features affect language without
mechanically determining it. Users adapt platforms, invent abbreviations, switch scripts, attach subtitles, use
hashtags and circulate screenshots. The correct unit of analysis is therefore the relationship between platform
design and social practice, not a claim that technology “causes” a single linguistic outcome.
13991399
GAJENDRA THAKUR
Figure 540 — Social media and language form a feedback system: platform affordances shape visible forms,
audience response and algorithmic selection, which feed back into community norms and everyday usage.
136.3 The regional language field is multilingual before it becomes digital
Digital multilingualism in Mithila, Vajji and Anga begins with an already multilingual social world.
Maithili, Bajjika, Angika, Hindi, Urdu, Nepali, Bhojpuri, English and other languages overlap through
locality, schooling, administration, marriage, work and migration. Census of India C-17 explicitly treats
bilingualism and trilingualism as measurable features, while Nepal’s census documents a highly plural
mother-tongue landscape. A social-media post that mixes Maithili and Hindi, or Bajjika and Nepali, should
not automatically be read as digital corruption of a previously pure code. It may represent an established
repertoire made newly visible in writing. The historian must distinguish long-standing multilingualism from
changes introduced by digital media.
136.4 Maithili has strong demographic presence but unequal digital resources
across domains
Maithili benefits from constitutional recognition in India, a large speaker base on both sides of the border
and substantial literary production. Yet recognition does not guarantee equal digital support in keyboards,
speech recognition, spell-checking, moderation, search, subtitles or machine translation. A language can be
demographically large and still function as a low-resource language in particular computational systems.
Social media partly compensates because video and audio allow creators to publish without sophisticated
language technology. At the same time, discoverability often depends on metadata typed in Devanagari,
Roman script, Hindi or English. The digital strength of Maithili is therefore domain-specific rather than
uniform.
136.5 Angika and Bajjika reveal how digital visibility can diverge from official
classification
Angika and Bajjika demonstrate why institutional statistics cannot be treated as complete maps of
linguistic life. In the Bihar Census 2011 C-16 presentation, neither appears with the same directly retrievable
headline status as Maithili, although both have strong regional identities and public cultural use. Nepal’s
2021 census separately records Bajjika and shows its major presence in Madhesh Province. Online, users can
name a language in a channel title, hashtag or video description even when official tables classify it differently.
HISTORY OF MITHILA, VAJJI & ANGA — VOLUME II
Social media thus creates a form of self-indexing: communities can assert names and boundaries in public
metadata. That visibility is culturally important, but it is not equivalent to a census count or legal
recognition.
Figure 541 — Madhesh Province mother tongues in Census 2021: Maithili 41.73%, Bhojpuri 18.81%, Bajjika
18.44%, Nepali 5.76%, Tharu 4.17%, Urdu 4.08%, Tamang 1.65%, with the remainder grouped here as other
languages.
136.6 Indic-language internet growth changes the economics of regional-
language audiences
The IAMAI–Kantar Internet in India Report 2024 estimated 886 million active internet users nationally,
with rural India accounting for 488 million. It also reported that 98 percent of internet users accessed
content in Indic languages and that 57 percent of urban users preferred regional-language content. These are
India-wide figures, not Bihar estimates, but they establish an important market context. Regional-language
content is no longer a marginal add-on to an English-dominant internet. Advertising, video distribution,
news, entertainment and commerce increasingly have incentives to address users in non-English languages.
For Maithili, Angika and Bajjika, the opportunity is real, but audience economics remain shaped by the
much larger markets for Hindi and other major languages.
136.7 Video lowers the script threshold and gives oral vernaculars new public
reach
Video is especially important for language varieties whose speakers may be more confident in speech than
in standardised writing. A singer, comic performer, farmer, teacher or political speaker can address an
audience in a local variety without choosing a formal orthography for every sentence. Subtitles and titles can
then be added in Devanagari, Roman script, Hindi or English. IAMAI–Kantar’s India-wide evidence shows
video as the most common activity undertaken in Indic languages, ahead of social networking and search. For
regional language history, this means that oral performance has become a major digital publication format.
The public record of language is no longer dominated by books, newspapers and institutional broadcasting.
14011401
GAJENDRA THAKUR
Figure 542 — India-wide Indic-language internet activities in the IAMAI–Kantar Internet in India Report
2024: video and music dominate, while social networking, shopping and search also occur substantially in Indic
languages; these are not Bihar-specific estimates.
136.8 Voice messages return spoken registers to the centre of digital
communication
Messaging applications changed digital writing by making audio ordinary. Voice notes allow elderly users,
people with limited literacy, migrants under time pressure and speakers unsure of standard spelling to
participate in digital networks through speech. They preserve intonation, hesitation, local vocabulary and
pronunciation that text messages flatten. Yet voice is also less searchable, less easily quoted and harder to
archive systematically. A language may therefore thrive in private audio circulation while remaining
underrepresented in searchable public text. Historians of digital language need to distinguish communicative
vitality from archival visibility.
136.9 Devanagari has become the practical common script for much regional-
language social media
For Maithili, Angika and Bajjika, Devanagari is widely usable across phones, operating systems and
mainstream platforms and overlaps with Hindi literacy acquired through schooling. This lowers the
technical cost of publishing regional-language text. The same convenience can blur linguistic boundaries
because identical script does not imply identical language. A Maithili sentence in Devanagari may be
mistaken by automated systems or casual readers for Hindi unless vocabulary, metadata or context mark the
distinction. Script convergence therefore supports access while making language identification more
dependent on linguistic and social cues.
136.10 Roman transliteration is a functional digital register, not automatic
evidence of language loss
Roman-script writing appears frequently in messaging, comments and short captions because keyboards
are familiar, code-switching with English is easy and users can avoid uncertainty about Devanagari spelling.
Such writing is often unstable: the same word may appear in several Roman spellings. That instability
reduces searchability and can make archiving difficult, but it also records pronunciation and informal style.
Treating Roman transliteration simply as linguistic decline misses its practical function. The more useful
HISTORY OF MITHILA, VAJJI & ANGA — VOLUME II
question is whether users can move between scripts and domains when required—for example, between a
Roman-script chat, a Devanagari public post and a formal Maithili publication.
136.11 Tirhuta or Mithilakshar has a different digital role: heritage, expertise and
identity
Unicode encodes Tirhuta at U+11480–U+114DF and identifies it as the traditional writing system of
Maithili. Digital encoding makes genuine text possible rather than image-only reproduction, yet encoded
existence is not the same as effortless everyday use. Fonts, keyboards, rendering support, user knowledge and
platform compatibility still matter. On social media, Tirhuta often functions as a marker of heritage,
scholarship, calligraphy and regional identity rather than as the default script for rapid conversation. This
specialised visibility can nevertheless be historically significant: a script that had become difficult to
reproduce in ordinary print can circulate globally as searchable text when the technical chain works.
136.12 Unicode, fonts and keyboards are part of language infrastructure
A user can only type what the input system makes practicable and only read what the device can render.
Unicode standardisation solved the fundamental problem of assigning stable code points, but practical
inclusion requires more: fonts, keyboard layouts, search indexing, line breaking, text-to-speech, optical
recognition and developer support. Chapter 116 treated Unicode and Tirhuta as technological recognition;
the present chapter follows the social consequence. When technical support is uneven, users route around it
through screenshots, images, Roman transliteration or another script. These workarounds preserve
communication but weaken machine readability and long-term digital preservation.
136.13 Code-switching is a resource for audience design
A single post may combine a Maithili or Bajjika spoken base with Hindi for a wider regional audience and
English for technical terms, hashtags or prestige. Such code-switching can index education, humour,
urbanity, migration or group membership. Creators also adjust language by platform: a family message may
use a local variety, a public caption Hindi, and a professional profile English. Networked multilingualism
therefore involves strategic audience design. The same person can inhabit several linguistic publics without
abandoning a mother tongue. Analysis should ask who is being addressed and what social work each language
performs.
136.14 Orthographic variation becomes publicly visible and therefore
contestable
Print institutions historically reduced variation through editors, proofreaders and house styles. Social
media exposes far more unedited writing, including local pronunciations, spellings, abbreviations and hybrid
forms. Comment threads can then become sites of correction: users debate whether a word is “proper”
Maithili, whether a form is Hindi-influenced, or whether a Bajjika or Angika expression belongs to another
label. This can strengthen metalinguistic awareness, but it can also convert normal variation into status
conflict. Digital visibility does not simply standardise language; it makes the politics of standardisation more
public.
136.15 Search and hashtags create pressure toward searchable spellings
Search systems operate on strings, metadata and probabilistic matching. A creator who wants discovery
has incentives to repeat widely recognised spellings, add alternative spellings or include Hindi and English
14031403
GAJENDRA THAKUR
keywords. Hashtags similarly reward consistency because a one-character difference can fragment an
audience. This may encourage practical standardisation around popular forms even without an academy or
grammar committee. It can also privilege the spellings used by the largest accounts. Searchability is therefore a
new source of linguistic authority—less formal than a dictionary, but powerful because it determines
whether content can be found.
136.16 Recommendation systems can expand small-language reach while
obscuring how selection occurs
Recommendation feeds can expose users to regional-language material they did not explicitly search for.
This is potentially transformative for small and dispersed linguistic publics: a song or comic clip can travel
beyond subscribers and locality. Yet recommendation systems are proprietary and constantly changing.
Researchers usually cannot see the full ranking logic, training data or language-identification errors. A
sudden increase in visibility may reflect content quality, engagement, platform experimentation or broader
trends. The historian should therefore treat algorithmic reach as a mechanism with uncertain internal
evidence rather than as a transparent measure of cultural popularity.
Figure 543 — Different platform affordances affect different linguistic dimensions: typing and script choice,
spoken registers, audiovisual performance, searchable spelling, algorithmic amplification and AI-assisted
conversion.
136.17 Monetisation can reward language mixing and scalable formats
Advertising and creator-revenue systems convert attention into income, but they do not necessarily
reward linguistic distinctiveness. A creator may mix Hindi with Maithili or Angika to reach a larger
monetisable audience, use English words in titles for search, or produce short repeatable formats favoured by
recommendation systems. This does not mean local languages disappear. It means linguistic choices are partly
embedded in platform economics. The strongest cultural output may come from creators who can move
between intimate local address and broader multilingual circulation. Economic incentives should therefore
be analysed alongside identity and expressive preference.
HISTORY OF MITHILA, VAJJI & ANGA — VOLUME II
136.18 Migration turns regional languages into translocal digital publics
Migration, discussed throughout Parts XI and XIII, is one of the most important engines of online
regional-language use. Migrants in Delhi, Punjab, Gujarat, Mumbai, the Gulf, Kathmandu and other
destinations maintain everyday contact with home through messaging and video. Social media allows
festivals, songs, local news, jokes and political debates to circulate between origin and destination in real time.
The audience for a Maithili or Angika video is thus not geographically confined to its historical region.
Digital language publics are translocal: they are organised by memory, kinship and interest as much as by
current residence.
136.19 The India–Nepal border is porous linguistically even when digital
regulation is national
Madhesh and north Bihar share Maithili and other linguistic networks, but telecommunications
regulation, data protection, platform law, education policy and official language regimes remain national.
The same video or message can cross the open border instantly while the institutional context differs. Nepal’s
2021 census records Maithili as 41.73 percent and Bajjika as 18.44 percent of Madhesh Province’s
population, alongside Bhojpuri, Nepali, Tharu, Urdu and other languages. Social media overlays this plural
census geography with a cross-border communication field. Cultural connectedness should not be mistaken
for administrative uniformity.
136.20 Gender and device ownership shape who speaks publicly and in which
register
Access to a household phone is not the same as autonomous control of a personal device. Earlier chapters
showed gender gaps in personal smartphone ownership among young people even where household access
was high. Shared devices can limit privacy, timing and willingness to participate in public discussion. Women
may use closed messaging groups more freely than public comment sections, while safety concerns can
influence profile names, photographs and language style. Social-media language data are therefore socially
selected: the most visible speakers are not automatically representative of everyone who speaks the language
offline.
136.21 Generational change is visible in format as much as vocabulary
Younger users often move quickly among short video, memes, emoji, voice notes, gaming vocabulary,
Hindi film references and English technical terms, while older users may rely more on voice calls, forwarded
messages or devotional and family content. These are tendencies, not fixed age rules. The linguistic
significance lies in format. A generation raised with searchable, editable and remixable media experiences
language through captions, comments and clips as well as conversation and books. This expands the range of
public writing while shortening many messages and increasing the role of audiovisual cues.
136.22 Memes and humour are laboratories of lexical innovation
Humour rewards recognition of local accent, stereotype, proverb, kinship term and double meaning.
Memes compress these into highly repeatable forms. A local word can become salient because it is attached to
a recurring image or catchphrase, while code-mixing creates jokes that depend on contrast between registers.
Humour can strengthen linguistic solidarity, but it can also caricature rural speech, caste, gender or district
identity. The researcher should preserve both creativity and power relations: virality is not neutral evidence of
acceptance.
14051405
GAJENDRA THAKUR
136.23 Music and short video create a powerful oral economy of language
Songs have always travelled across Mithila, Vajji and Anga through performance, cassettes, radio,
television and migration. Social platforms change the speed and granularity of circulation. A complete song
can coexist with fifteen-second hooks, dance clips, devotional excerpts and user remixes. Language travels
through melody even when listeners cannot read the script. This helps explain why oral and musical forms
can achieve digital reach larger than prose. It also creates copyright and attribution questions when
recordings are reposted without clear credit, linking the language economy back to the intellectual-property
issues examined in Chapter 128.
136.24 News and public debate create a new vernacular sphere but not a single
public sphere
Local journalists, institutions, political actors and ordinary users now publish directly through pages,
channels and groups. Regional languages can therefore enter public debate without waiting for a newspaper
column or broadcast slot. Yet online publics are fragmented by platform, locality, caste/community
networks, ideology and recommendation systems. A widely viewed Maithili video may reach one segment of
speakers while remaining invisible to another. The existence of more vernacular content should not be
equated with a unified democratic public sphere. The key change is lower publication cost combined with
greater fragmentation of attention.
136.25 Misinformation travels through trusted linguistic networks as efficiently
as accurate information
Regional-language intimacy can increase trust. A message forwarded by a relative or recorded in a familiar
accent may feel more credible than an anonymous English notice. This is valuable for public-health
communication, disaster warnings and civic information, but the same trust can accelerate rumours,
fabricated quotations and manipulated video. Fact-checking systems often have fewer resources for smaller
languages and local dialects than for major national languages. Digital literacy therefore includes the ability to
verify claims across languages and sources, not merely the ability to operate a phone.
136.26 Content moderation and automated language identification are uneven
for low-resource varieties
Large platforms use combinations of automated classifiers, user reports and human review. Performance
is generally strongest where abundant labelled data and commercial incentives exist. Regional varieties with
fluid spelling, Roman transliteration and code-mixing can be harder to classify. A system may mistake
Maithili or Bajjika for Hindi, fail to recognise abusive local vocabulary, or over-remove benign expressions if
context is poorly modelled. Because moderation systems are proprietary, regional error rates are rarely
available. Researchers should describe this as a plausible structural risk supported by the low-resource
character of many language technologies, not invent unsupported platform-specific accuracy figures.
136.27 AI translation and speech technologies create a new phase of digital
language politics
India’s BHASHINI programme illustrates the rapid institutional expansion of multilingual AI. The
Ministry of Electronics and Information Technology’s 2025–26 annual report states that the platform
supports more than 36 languages for text translation and more than 22 for voice, alongside a large
multilingual glossary. Such infrastructure can improve subtitles, government services, speech interfaces and
HISTORY OF MITHILA, VAJJI & ANGA — VOLUME II
content discovery. Yet support counts do not guarantee equal quality across languages or varieties. Training
data, dialect coverage, named entities, spelling variation and code-switching remain difficult. For Maithili,
Angika and Bajjika, AI can widen access while also amplifying errors if users treat automated output as
authoritative language.
136.28 Digital archives must confront ephemerality, deletion and platform
dependence
Social media produces enormous quantities of linguistic material but is an unstable archive. Accounts
disappear, links break, platforms change policies, comments are deleted and recommendation context is lost.
Screenshots preserve appearance but not searchable text or metadata; downloaded videos may lose captions
and discussion threads. Long-term language history therefore requires deliberate preservation outside the
platform: dated captures, transcripts, metadata, rights information and where possible original files. Digital
abundance should not be confused with permanent preservation.
136.29 Measurement is the central methodological problem
Official censuses measure mother tongue, not social-media language. Telecom regulators measure
subscriptions and traffic, not the language of every message. IAMAI–Kantar measures national internet
behaviour through surveys, while platform analytics are proprietary and creator-specific. Public search
counts change over time and can be distorted by duplicate, automated or cross-linguistic content. No single
dataset therefore answers “how much Maithili, Angika or Bajjika is on social media?” A rigorous account
triangulates population language data, connectivity evidence, platform affordances, samples of public
content and qualitative observation, and states clearly when figures are national rather than regional.
136.30 Conclusion: social media expands language domains without
guaranteeing language transmission
The major historical change is domain expansion. Regional languages now circulate through messaging,
video, music, comments, livestreams, digital commerce and transnational family networks in ways that print-
era institutions could not match. This can increase prestige, creativity and cross-border connection, especially
where oral forms are strong. But digital visibility is not the same as intergenerational transmission, literacy,
educational use or institutional security. A language can be highly visible in entertainment while weakening
in schooling or formal writing. The future of Maithili, Angika and Bajjika will therefore depend on the
interaction of homes, schools, publishing, public institutions and digital platforms rather than on social
media alone.
Table 136.1 — Evidence architecture for analysing social media and language
Evidence source What it establishes Use in Chapter 136 Main limitation
Census of India 2011 mother tongue plus baseline linguistic dated; does not
C-16 / C-17 bilingualism/trilingualism repertoire and measure social-media
in Bihar classification use
Linguistic Survey of census-linked description shows how official classification is not
India: Bihar of Bihar language classification groups identical to self-
categories and names varieties identification online
Nepal NPHC 2021 / mother-tongue counts cross-border census language is not
Languages in Nepal and multilingual comparison; Maithili platform language
2025 structure and Bajjika in
Madhesh
Madhesh Province province-level language regional visualisation administrative
of the multilingual province is not
14071407
GAJENDRA THAKUR
Evidence source What it establishes Use in Chapter 136 Main limitation
Data Portal counts from NPHC 2021 field identical to historical
Mithila
IAMAI–Kantar national active internet establishes scale of India-wide survey;
Internet in India users and Indic-language Indic-language digital not Bihar-specific
Report 2024 activities demand
TRAI performance current telecom and connectivity telecom subscriptions
indicators 2026 internet-system context in infrastructure do not identify
India surrounding language language
use
Nepal broadband and mobile- connectivity context national service data
Telecommunications service context on Nepal side are not Madhesh
Authority MIS language-use counts
Unicode Standard, encoded script repertoire technical basis for encoding does not
Tirhuta block machine-readable ensure fonts,
Tirhuta text keyboards or user
adoption
BHASHINI / MeitY multilingual AI models, institutional context support count does
2025–26 translation and voice for language not demonstrate
support technology equal quality by
variety
Public platform actual spellings, scripts, qualitative study of non-random;
samples / creator formats and audience language practice proprietary metrics;
analytics interaction unstable over time
Digital ethnography how users negotiate interprets social requires contextual
and CMC identity, code-switching meaning rather than sampling and cannot
scholarship and platform norms only counts produce a census
total