MACHINE ASR ACCESSIBILITY AID

Maithili’s_Fight_for_Digital_Survival.m4a

Not an editorially verified transcript. This text was generated automatically from the preserved recording and may contain recognition, language-detection, spelling, segmentation or name errors. Consult the source recording for authoritative content.
Collection
Part 7 · VIDEHA MITHILA MAITHILI DISCUSSION CRITICISM SERIES PART 7
Status
asr-draft
Human verified
No
Editorial review
not-reviewed
ASR model
small
Detected language
en (0.99735)
Duration
20:43
Source
Open preserved recording

Timestamped machine output

  1. 0:00–0:30आप आप आप आप आप आप आप आप आप आप आप आप आप आप आप आप आप आप आप आप आप आप आप आप आप आप आप आप आप आप आप आप आप आप आप आप आप आप आप आप आ� आ� आप आप आ� आ� आप आप आप आ� आप आ� आप आ� आ� आ� आ� आप �
  2. 0:30–0:34massive collision of eras. You have the cold binary reality of the modern internet, and
  3. 0:34–0:39then you have a language whose oldest known treatise dates back to the time of Geoffrey
  4. 0:39–0:44Chaucer. Yeah, a language that is actively fighting to survive in the digital world.
  5. 0:44–0:48Exactly. So our mission today for this Deep Dyes is to explore a truly fascinating document
  6. 0:48–0:53from 2009. It's the English-Mathilli Computer Dictionary Volume 1, and it was compiled
  7. 0:53–0:58by Gajendra Sakhur, Nagendra Kumar Shah, and Panjikar Vijayananjha. Right, and
  8. 0:58–1:02We're also going to be looking at the really deeply insightful forward to this dictionary,
  9. 1:02–1:04which was written by Professor Udaya Narayana Singh.
  10. 1:04–1:09Yeah, because we want to explore how an ancient language modernizes itself, you know, the
  11. 1:09–1:14monumental effort behind dictionary making and this massive socio-political battle over
  12. 1:14–1:19who actually gets to be counted in official census data.
  13. 1:19–1:23Because as it turns out, dictionaries aren't just about spelling, they're really about
  14. 1:23–1:24existence.
  15. 1:24–1:25Exactly.
  16. 1:25–1:26To be named is to be recognized.
  17. 1:26–1:27Right.
  18. 1:27–1:33really grasp the sheer scale of what these authors accomplished by building a modern computer
  19. 1:33–1:38dictionary for a traditional language like mathilly. We first need to understand how humans
  20. 1:38–1:40have historically built dictionaries in the first place.
  21. 1:40–1:44Right. Let's definitely go back in time because I'll admit when I thought of the first dictionary,
  22. 1:44–1:49my mind immediately went to the Western world. I pictured Samuel Johnson's famous English
  23. 1:49–1:57dictionary from 1755 or maybe Robert Codrie's A Table Alphabetical from 1604. I just figured
  24. 1:57–2:01It was you know some guy writing down words you heard in a tavern. Yeah, that's a pretty common misconception
  25. 2:01–2:08Honestly, but the source material points out that humanity's obsession with cataloging words is so much older and way more systematic than that
  26. 2:08–2:11Oh, really? How far back are we talking? Far older
  27. 2:11–2:18The forward actually notes that the oldest known Western dictionary efforts originated in the Akkadian Empire Wow
  28. 2:18–2:25Yeah, we have these bilingual Sumerian Akkadian word lists found in modern Syria that date back to roughly
  29. 2:25–2:322300 BC. That is incredible. It is. And the Chinese lexicographic tradition and like,
  30. 2:32–2:36you know, lexicography being the science of dictionary making, that goes back to the third
  31. 2:36–2:41century BC. So that's over four millennia of humans just trying to organize language. But
  32. 2:41–2:45what really grabbed my attention in the sources was the history of dictionary making in India
  33. 2:45–2:50because it wasn't just about defining words so people could read the daily news or whatever.
  34. 2:50–2:51No, not at all.
  35. 2:51–2:56Right, in the Indian tradition, lexicography was originally driven by this very specific,
  36. 2:56–3:00almost sacred need, which was the preservation of Vedic literature.
  37. 3:00–3:05Yes, and their approach to it was entirely structural.
  38. 3:05–3:10Ancient scholars in the Sanskrit grammatical traditions, like Panini, they had this rigorous
  39. 3:10–3:11methodology.
  40. 3:11–3:13Right, they weren't just making alphabetical lists.
  41. 3:13–3:14Exactly!
  42. 3:14–3:16They didn't just write down words and what they meant.
  43. 3:16–3:20They literally segmented ancient Vedic sentences
  44. 3:20–3:22into individual words,
  45. 3:22–3:24and then broke those words down further
  46. 3:24–3:26into their root and suffix components.
  47. 3:26–3:28Okay, when I read about them dissecting sentences
  48. 3:28–3:30into roots and suffixes,
  49. 3:30–3:33my mind immediately went to modern software development.
  50. 3:33–3:34That's interesting.
  51. 3:34–3:36Yeah, like these ancient Indian scholars,
  52. 3:36–3:39they were basically the original programmers
  53. 3:39–3:41debugging lines of human code.
  54. 3:41–3:44That is actually a brilliant way to look at it, yes.
  55. 3:44–3:45And by applying that logic,
  56. 3:45–3:48were looking for the source code of human language.
  57. 3:48–3:49Right.
  58. 3:49–3:54They realized that if you can break a word down to its absolute root light, its base
  59. 3:54–3:59variable, you can add prefixes or suffixes to build entirely new linguistic functions
  60. 3:59–4:02without breaking the rules of the language itself.
  61. 4:02–4:02Exactly.
  62. 4:02–4:06They were like pulling apart the syntax to figure out the underlying logic, testing
  63. 4:06–4:10the variables, basically just to make sure the cultural system wouldn't crash and
  64. 4:10–4:12the knowledge wouldn't be lost to time.
  65. 4:12–4:17And by doing that, they develop these incredibly complex theories on how sounds and word structures
  66. 4:17–4:18actually work.
  67. 4:18–4:19Right.
  68. 4:19–4:20Which we see in the text.
  69. 4:20–4:21Yeah.
  70. 4:21–4:25You see this in the Negantu from 700 BC and later in Amara Simha's Amarikosa from the
  71. 4:25–4:266th century AD.
  72. 4:26–4:30And that one arranged words by their synonyms and homonyms rather than just alphabetical
  73. 4:30–4:32order.
  74. 4:32–4:34Which sounds incredibly complicated.
  75. 4:34–4:35It is.
  76. 4:35–4:40Creating a dictionary requires phonetic marking, meticulously organizing definitions,
  77. 4:40–4:44and trying to predict how future users will actually search for a concept.
  78. 4:44–4:47So it's way more than a clerical task of just typing out a list.
  79. 4:47–4:48Oh, absolutely.
  80. 4:48–4:52If we connect this to the bigger picture, lexicography is really about structuring human
  81. 4:52–4:53thought.
  82. 4:53–4:55A dictionary is a cognitive map.
  83. 4:55–4:57I love that phrase, a cognitive map.
  84. 4:57–4:58It is.
  85. 4:58–5:03It tells you how a culture perceives reality and how it categorizes the physical and
  86. 5:03–5:05abstract world around it.
  87. 5:05–5:09So if you are building a dictionary for computing terms in a traditional language
  88. 5:09–5:14like Mathili, you are essentially drawing a brand new map for a very old territory.
  89. 5:14–5:18You are laying down digital highways over ancient landscapes.
  90. 5:18–5:20That's a great way to visualize it.
  91. 5:20–5:23Which brings us to the core mystery of this source material.
  92. 5:23–5:29Why did Mepheli specifically need this new technological map in 2009?
  93. 5:29–5:34Because reading the forward of this dictionary feels less like a dry academic introduction
  94. 5:34–5:37and more like, I don't know, a geopolitical thriller.
  95. 5:37–5:41It really does, because my belly wasn't just casually evolving in the background.
  96. 5:41–5:44It was locked in this massive existential battle for its identity.
  97. 5:44–5:49Right, and to give you listening some context, this language spans a really massive, vibrant
  98. 5:49–5:51cultural space.
  99. 5:51–5:55It's spoken across the Genetic Plain in the Indian state of Bihar, the Terai region,
  100. 5:55–5:59at the Himalayan foothills in Churkand, and it makes up a significant chunk, about
  101. 5:59–6:0114% of the entire population of Nepal.
  102. 6:01–6:04Right, it's not a small, isolated language.
  103. 6:04–6:05Not at all.
  104. 6:05–6:09If you look at the official Indian census data over the 20th century, the number of
  105. 6:09–6:11make-lea speakers is absolutely wild.
  106. 6:11–6:13It makes no logical sense whatsoever.
  107. 6:13–6:14No, it's completely erratic.
  108. 6:14–6:19Yeah, the source notes that between 1911 and 1921, the population of speakers supposedly
  109. 6:19–6:24decreased by 0.77%.
  110. 6:24–6:30But then, between 1951 and 1961, the official census says the number of speakers suddenly
  111. 6:30–6:33rocketed up by 22.35%.
  112. 6:33–6:35And the wildly erratic fluctuations just continue from there.
  113. 6:35–6:42The 2001 census officially counted about 12.1 million Methylese speakers.
  114. 6:42–6:46But Professor Singh, who wrote the foreword, he breaks down historical data, geographical
  115. 6:46–6:49expansions and normal population growth over the decades.
  116. 6:49–6:53And he estimates the actual number of speakers is closer to 40 million.
  117. 6:53–6:5440 million.
  118. 6:54–6:58Let me stop you right there because how on earth can tens of millions of people
  119. 6:58–7:01simply vanish and reappear in official government data?
  120. 7:01–7:02It's pretty shocking.
  121. 7:02–7:06I mean, looking at these census numbers, it's like looking at a volatile stock market chart.
  122. 7:06–7:11It's like the stock being traded as an entire community's linguistic identity, and political
  123. 7:11–7:14record keepers are just mipulating the market.
  124. 7:14–7:16How does a government misplace tens of millions of speakers?
  125. 7:16–7:21Well, it comes down to how data is categorized, and more importantly, the politics of that
  126. 7:21–7:23categorization.
  127. 7:23–7:28According to the source material, these erratic numbers were the result of a very specific
  128. 7:28–7:31structural spread of disinformation.
  129. 7:31–7:34information, like what? Well, there was a narrative pushed that
  130. 7:34–7:39Methili was an exclusive language spoken only by the Brahmin cast in the
  131. 7:39–7:42Mithila region. Okay, but why push that narrative? What's the goal there?
  132. 7:42–7:47The goal, according to the foreword, was to classify Mithili not as a distinct
  133. 7:47–7:53independent language, but merely as a regional dialect of Hindi. Oh, I see.
  134. 7:53–7:56Yeah, by officially counting Mithili speakers as Hindi speakers,
  135. 7:56–8:00it artificially inflated the demographic power and official count of the
  136. 8:00–8:01Hindi language on the census.
  137. 8:01–8:07Wow. So they just relabeled millions of people to boost another language's numbers.
  138. 8:07–8:12But the actual demographic data completely contradicts that Brahman only narrative, right?
  139. 8:12–8:17Completely. The census returns that accurately recorded methilies show massive overwhelming
  140. 8:17–8:21support across all demographics, which completely shatters the caste-exclusive myth.
  141. 8:21–8:23Right. The numbers just don't back it up.
  142. 8:23–8:29Exactly. For instance, the source notes that up to 46.84% of the population in some
  143. 8:29–8:31if the last speaking districts are Muslims.
  144. 8:31–8:34Which means you couldn't possibly have those high returns for
  145. 8:34–8:39Maitrely unless a massive portion of the Muslim population was also registering
  146. 8:39–8:41Maitrely as their mother tongue.
  147. 8:41–8:46Precisely. The data proves it is a language of the broader region deeply embedded across
  148. 8:46–8:49different communities, not just a single cast.
  149. 8:49–8:53And I want to clarify for you listening, our goal here isn't to weigh in on the
  150. 8:53–8:58historical political disputes or, you know, take sides on Indian state politics.
  151. 8:58–8:59Of course not.
  152. 8:59–9:04We are simply unpacking the demographic realities reported in our source material to understand
  153. 9:04–9:10exactly why this dictionary was created. The historical stakes for this language were incredibly
  154. 9:10–9:15high. They really were. And what's fascinating is how the community responded to all of this.
  155. 9:15–9:20The forward describes this long period of being denied constitutional rights as a boon in disguise.
  156. 9:20–9:25Which sounds completely counterintuitive. I mean, how is being erased from the sense as a boon?
  157. 9:25–9:31because the political friction actually fueled a fierce cultural and literary vigor among the speakers.
  158. 9:31–9:33Oh, so it pushed them to fight back.
  159. 9:33–9:40Exactly. The resistance to being reclassified created a powerful sense of unity and advocacy.
  160. 9:40–9:45They wrote more. They published more. They organized. And that resilience eventually
  161. 9:45–9:50pan off, leading to Maitreya Lee being officially included in the eighth schedule of the Indian
  162. 9:50–9:54Constitution. Right. And for those who might not know, being included in the eighth schedule
  163. 9:54–9:58basically means the government legally recognizes the language.
  164. 9:58–10:00Yes, it's a huge milestone.
  165. 10:00–10:04It grants it official status, meaning it can be used in government exams,
  166. 10:04–10:09it receives federal funding for development, and it gets institutional patronage.
  167. 10:09–10:11It is a massive victory for a language of survival.
  168. 10:11–10:14It is the ultimate official validation.
  169. 10:14–10:19But as the source points out, this victory immediately created a new massive problem.
  170. 10:19–10:20Always a new problem.
  171. 10:20–10:21Right.
  172. 10:21–10:26And this brings us to why a computer dictionary matters so much to a community of 40 million people.
  173. 10:27–10:30The forward points out a crucial statistic.
  174. 10:30–10:37While many Mathili speakers are multilingual, you know, they can navigate Hindi, Bhujpuri, Magahi and Bengali,
  175. 10:37–10:43about 25 to 30 percent of Mathili speakers are completely monolingual.
  176. 10:43–10:45Meaning they only speak Mathili.
  177. 10:45–10:45Right.
  178. 10:45–10:48and only a tiny fraction of the total population,
  179. 10:48–10:52maybe three to 5%, can speak English effectively.
  180. 10:52–10:54Okay, let's put ourselves in there shoes for a second.
  181. 10:54–10:56Imagine you are one of those 10 million
  182. 10:56–10:58monolingual mythology speakers.
  183. 10:58–11:01You've just won this massive constitutional victory
  184. 11:01–11:02for your language.
  185. 11:02–11:03A huge moment of pride.
  186. 11:03–11:05Yeah, but then you pick up a smartphone
  187. 11:05–11:07or you sit at a computer terminal
  188. 11:07–11:08in a local government office
  189. 11:08–11:11and every single button, error message,
  190. 11:11–11:14and software setting is in English or Hindi languages
  191. 11:14–11:15you don't read.
  192. 11:15–11:16It's a wall.
  193. 11:16–11:19You don't have a word for internet or browser or download.
  194. 11:19–11:20You were just digitally stranded.
  195. 11:20–11:23You were locked out of the 21st century entirely.
  196. 11:23–11:25Exactly.
  197. 11:25–11:29Being constitutionally recognized is really only half the battle.
  198. 11:29–11:33If your language cannot interact with a microprocessor, it will eventually die out in the modern
  199. 11:33–11:34world.
  200. 11:34–11:39Having access to digital tools is a critical step in a language's survival.
  201. 11:39–11:41Which is what birthed this specific dictionary.
  202. 11:41–11:47authors Thakur and the Jaws took on the monumental task of bringing Mythili online.
  203. 11:47–11:49Oh wait, let me play devil's advocate here for a minute.
  204. 11:49–11:50Sure.
  205. 11:50–11:57I completely get the cultural pride, but wouldn't it be vastly easier to just use the English
  206. 11:57–11:58words?
  207. 11:58–12:01I mean, almost every other language just borrows tech terms, right?
  208. 12:01–12:04French people say le weekend and le computer.
  209. 12:04–12:08Why go through the grueling intellectual labor of inventing brand new Mythili
  210. 12:08–12:10words for things like algorithm?
  211. 12:10–12:13That's a very fair question and it goes back to what we discussed earlier about cognitive
  212. 12:13–12:14maps.
  213. 12:14–12:15Right, mapping the territory.
  214. 12:15–12:19Yes, if you just take an English word like bandwidth and drop it into a monolingual rural
  215. 12:19–12:23community in Bihar, it has no conceptual hook.
  216. 12:23–12:25It's just a meaningless sound.
  217. 12:25–12:26That makes sense.
  218. 12:26–12:29But if you can build a bridge between the new technology and the concepts they already
  219. 12:29–12:33intimately understand, the technology becomes accessible.
  220. 12:33–12:35It feels natural, not foreign.
  221. 12:35–12:38Okay, but that makes perfect sense.
  222. 12:38–12:41And to achieve that, the authors didn't just take the easy way out.
  223. 12:41–12:43Here's where it gets really interesting.
  224. 12:43–12:48They engaged in what I can only describe as linguistic alchemy.
  225. 12:48–12:49Oh, absolutely.
  226. 12:49–12:54They took these cold, binary tech concepts and fused them with classical roots.
  227. 12:54–12:58What's fascinating here is the lexical engineering, the actual building of these
  228. 12:58–13:01new words, is profound.
  229. 13:01–13:05Let's look at some specific translations from the text because they perfectly illustrate
  230. 13:05–13:07this bridge between the ancient and the digital.
  231. 13:07–13:09Yes, I really want to get into the actual words.
  232. 13:09–13:11Let's start with the basics.
  233. 13:11–13:12How do you say computer?
  234. 13:12–13:15In the dictionary, it is translated as sun-gunuk.
  235. 13:15–13:18It utilizes an ancient root related to counting
  236. 13:18–13:22and calculation, elegantly capturing the fundamental nature
  237. 13:22–13:24of a machine that computes data.
  238. 13:24–13:25Okay, sun-gunuk.
  239. 13:25–13:27Simple enough, it's a calculator.
  240. 13:27–13:30But what about something highly conceptual,
  241. 13:30–13:31like a cyber cafe?
  242. 13:31–13:33This is one of the most beautiful translations
  243. 13:33–13:35in the book, in my opinion.
  244. 13:35–13:42translated cybercafe as Sangunik Pangriha. I absolutely love this one. Pangriha literally
  245. 13:42–13:47translates to a traditional drink house or a tavern, right? Yes, exactly. So they took
  246. 13:47–13:52the modern concept of an internet cafe, a communal place where people gather, sit together,
  247. 13:52–13:57and consume data, and they mapped it perfectly onto the ancient cultural concept of a tavern.
  248. 13:58–14:02A drink house of computers. That is just brilliant. It really is. It highlights
  249. 14:02–14:07the sheer intellectual labor required by the lexicographers. Consider the term algorithm,
  250. 14:07–14:13a highly specific mathematical and computational concept that, honestly, most people struggle
  251. 14:13–14:18to define even in English. Oh, for sure. They translated it as abhukti vidi kalpa.
  252. 14:18–14:21It sounds so poetic, but what does it actually mean?
  253. 14:21–14:26Well, vidi relates to a rule, a method, or a procedure, and kalpa implies an order
  254. 14:26–14:31or a rule of practice. So they are framing an algorithm not as some invisible technological
  255. 14:31–14:36magic but as a formal ordered procedure. Which is exactly what an algorithm is.
  256. 14:36–14:39Right. Or take artificial intelligence. They translated that as
  257. 14:39–14:43Kretrem-Pragueh. Let's unpack that because I know Kretrem means artificial or
  258. 14:43–14:49constructed but what is Pragueh? Pragueh is a deep classical term. It relates to
  259. 14:49–14:53wisdom, supreme knowledge or intelligence in the classical almost
  260. 14:53–14:57spiritual sense. Oh wow. Yeah. By combining it with Kretrem they are
  261. 14:57–15:02elevating the technology. They are describing a neural network using vocabulary historically
  262. 15:02–15:08reserved for philosophy and theology. That is fascinating. But you know, it also works
  263. 15:08–15:12for the everyday frustrating parts of tech too, like a bug in the software. When your
  264. 15:12–15:18app crashes, they translated the concept of a bug as DOSHA. Now why use DOSHA? This
  265. 15:18–15:23is a perfect example of repurposing classical knowledge. In traditional Ayurvedic medicine
  266. 15:23–15:29in Indian philosophy, adosha is an imbalance or a fundamental flaw in a system's constitution.
  267. 15:29–15:30Oh my gosh.
  268. 15:30–15:34So a computer bug isn't just a literal insect inside the machine like we use in English.
  269. 15:34–15:38They've translated it to mean a fundamental imbalance in the software's harmony.
  270. 15:38–15:43You are basically diagnosing the code with a medical ailment.
  271. 15:43–15:45That is such a vivid way to understand a glitch.
  272. 15:45–15:48It is perfectly vivid and look at bandwidth.
  273. 15:48–15:50They translated it as vahini-vistara.
  274. 15:50–15:53Which means the expanse or the spread of the channel.
  275. 15:53–15:53Yes.
  276. 15:53–15:58Vahini being a channel or a river and Vistara being the expanse.
  277. 15:58–15:59OK, I see where this is going.
  278. 15:59–16:00Think about it.
  279. 16:00–16:03If you are a monolingual farmer trying
  280. 16:03–16:05to understand why your agricultural weather
  281. 16:05–16:08app is loading slowly, the English word
  282. 16:08–16:10bandwidth means absolutely nothing.
  283. 16:10–16:11Right, it's just a foreign sound.
  284. 16:11–16:13But Vahini Vistara.
  285. 16:13–16:15A farmer intimately understands the concept
  286. 16:15–16:17of a channel's width, restricting
  287. 16:17–16:19the flow of water to a field.
  288. 16:19–16:21By using that term, the dictionary
  289. 16:21–16:26maps the invisible flow of digital data onto the physical reality of agricultural irrigation.
  290. 16:26–16:32That is, they took a 21st century operating system and explained it using a 2,000 year old
  291. 16:32–16:34agrarian landscape. That is the magic of this dictionary.
  292. 16:34–16:38And you see this deep contextual mapping throughout the source material. For instance,
  293. 16:38–16:43the concept of Boolean logic, you know, the true and false one in zero foundation of all
  294. 16:43–16:46computing code. Right. The absolute basics of programming.
  295. 16:46–16:52They translated it as Boolean Tarka. They keep the name of the mathematician George Bool,
  296. 16:52–16:54but they attach it to Tarka.
  297. 16:54–17:00And Tarka is the classical Indian philosophical system of logic and rigorous debate, isn't it?
  298. 17:00–17:07Right. It immediately signals to the user. We are dealing with absolute structural logic here.
  299. 17:07–17:11It grounds the foreign mathematics in a familiar intellectual tradition.
  300. 17:11–17:16Or what about data mining? They translated it as Datash Uttkana,
  301. 17:16–17:21And Utkana is the ancient human act of physically excavating the earth, like digging a well
  302. 17:21–17:22or a mine.
  303. 17:22–17:24Yes, physical excavation.
  304. 17:24–17:28So they just applied the physical act of digging dirt to the digital excavation of
  305. 17:28–17:29data.
  306. 17:29–17:30It's a perfect metaphor.
  307. 17:30–17:31It truly is.
  308. 17:31–17:34And to make sure all of this was incredibly accessible, the authors didn't just stop
  309. 17:34–17:35at the words.
  310. 17:35–17:38They formatted the dictionary using multiple scripts.
  311. 17:38–17:41Right, because reading the word is just as important as knowing the word.
  312. 17:41–17:42Exactly.
  313. 17:42–17:46Every single entry includes the international phonetic alphabet, so a linguist from anywhere
  314. 17:46–17:49in the world knows exactly how to pronounce the Myphilic word.
  315. 17:49–17:50That's super helpful.
  316. 17:50–17:55It also includes the widely used Devanagari script, which is what Hindi and Sanskrit are
  317. 17:55–17:57usually written in today.
  318. 17:57–18:01And crucially, it includes the traditional Mithalakshara, also known as the Tirhuta
  319. 18:01–18:02script.
  320. 18:02–18:06Which is a huge deal, because they honored the language's ancient orthography, meaning
  321. 18:06–18:10they literally printed the book using the traditional written script that the Mithili
  322. 18:10–18:16people have been using for centuries rather than forcing them to only read a modernized alphabet.
  323. 18:16–18:21Yeah, this dictionary proves that traditional languages are not static museum pieces waiting
  324. 18:21–18:27to collect dust. They are highly elastic. They really are. When given the proper academic rigor
  325. 18:27–18:32and when they secure that official patronage, they can expand to encompass any modern concept.
  326. 18:32–18:36The challenge isn't that a language is too old to understand the internet. The challenge is
  327. 18:36–18:41merely finding the right lexicographers to do the heavy lifting of building the bridges.
  328. 18:41–18:45So if we zoom all the way out, what does this actually mean for us for you listening to
  329. 18:45–18:46this deep dive right now?
  330. 18:46–18:50I think it fundamentally changes how we are supposed to look at a dictionary.
  331. 18:50–18:52I completely agree.
  332. 18:52–18:56Because the English May Selly Computer Dictionary is far more than just a thick reference
  333. 18:56–18:58book sitting on a desk somewhere in New Delhi.
  334. 18:58–19:01It is an active tool of survival.
  335. 19:01–19:03It is an active defiance, honestly.
  336. 19:03–19:08It represents a vibrant population of tens of millions of people who flat out refuse
  337. 19:08–19:13to let their mother tongue fade into obscurity simply because the world decided to invent
  338. 19:13–19:16microchips and fiber optic cables.
  339. 19:16–19:21By defining these words, they are ensuring that their language has a permanent seat
  340. 19:21–19:22at the digital table.
  341. 19:22–19:28It's taking that ancient linguistic x-ray machine, the exact same logical framework
  342. 19:28–19:33that scholars used to dissect Vedic texts thousands of years ago and plugging it directly
  343. 19:33–19:35into a modern server farm.
  344. 19:35–19:36It really is.
  345. 19:36–19:40It proves that an ancient language can absolutely hold its own in a modern cybercafe or rather
  346. 19:40–19:43a Sangunik pan-gre-ha.
  347. 19:43–19:48And frankly, it serves as a master blueprint for thousands of other regional and indigenous
  348. 19:48–19:53languages around the globe that are currently facing the exact same digital threshold.
  349. 19:53–19:57Which leaves us with a final thought to mull over as we wrap up today.
  350. 19:57–20:02As artificial intelligence and global tech continue to advance at lightning speed, as
  351. 20:02–20:07neural networks get better and better at live translation, will local languages have to
  352. 20:07–20:12constantly race to invent new dictionaries to keep up with the ever-changing tech vocabulary?
  353. 20:12–20:14Oh, that's a fascinating prospect.
  354. 20:14–20:19Or will our technology eventually evolve to the point where it natively speaks and
  355. 20:19–20:23understands the nuance of every ancient language perfectly without us needing humans
  356. 20:23–20:25to build a bridge at all?
  357. 20:25–20:31Will the machine eventually learn the etymology of Doshua or Pragya so fluently that the burden
  358. 20:31–20:35of translation is finally lifted from the humans who speak it?
  359. 20:35–20:36Definitely something to think about.
  360. 20:36–20:40But until that day comes, we desperately need the dictionary makers.
  361. 20:40–20:41We need the map builders.
  362. 20:41–20:43Thanks for taking the deep dive with us.

Plain text

आप आप आप आप आप आप आप आप आप आप आप आप आप आप आप आप आप आप आप आप आप आप आप आप आप आप आप आप आप आप आप आप आप आप आप आप आप आप आप आप आ� आ� आप आप आ� आ� आप आप आप आ� आप आ� आप आ� आ� आ� आ� आप � massive collision of eras. You have the cold binary reality of the modern internet, and then you have a language whose oldest known treatise dates back to the time of Geoffrey Chaucer. Yeah, a language that is actively fighting to survive in the digital world. Exactly. So our mission today for this Deep Dyes is to explore a truly fascinating document from 2009. It's the English-Mathilli Computer Dictionary Volume 1, and it was compiled by Gajendra Sakhur, Nagendra Kumar Shah, and Panjikar Vijayananjha. Right, and We're also going to be looking at the really deeply insightful forward to this dictionary, which was written by Professor Udaya Narayana Singh. Yeah, because we want to explore how an ancient language modernizes itself, you know, the monumental effort behind dictionary making and this massive socio-political battle over who actually gets to be counted in official census data. Because as it turns out, dictionaries aren't just about spelling, they're really about existence. Exactly. To be named is to be recognized. Right. really grasp the sheer scale of what these authors accomplished by building a modern computer dictionary for a traditional language like mathilly. We first need to understand how humans have historically built dictionaries in the first place. Right. Let's definitely go back in time because I'll admit when I thought of the first dictionary, my mind immediately went to the Western world. I pictured Samuel Johnson's famous English dictionary from 1755 or maybe Robert Codrie's A Table Alphabetical from 1604. I just figured It was you know some guy writing down words you heard in a tavern. Yeah, that's a pretty common misconception Honestly, but the source material points out that humanity's obsession with cataloging words is so much older and way more systematic than that Oh, really? How far back are we talking? Far older The forward actually notes that the oldest known Western dictionary efforts originated in the Akkadian Empire Wow Yeah, we have these bilingual Sumerian Akkadian word lists found in modern Syria that date back to roughly 2300 BC. That is incredible. It is. And the Chinese lexicographic tradition and like, you know, lexicography being the science of dictionary making, that goes back to the third century BC. So that's over four millennia of humans just trying to organize language. But what really grabbed my attention in the sources was the history of dictionary making in India because it wasn't just about defining words so people could read the daily news or whatever. No, not at all. Right, in the Indian tradition, lexicography was originally driven by this very specific, almost sacred need, which was the preservation of Vedic literature. Yes, and their approach to it was entirely structural. Ancient scholars in the Sanskrit grammatical traditions, like Panini, they had this rigorous methodology. Right, they weren't just making alphabetical lists. Exactly! They didn't just write down words and what they meant. They literally segmented ancient Vedic sentences into individual words, and then broke those words down further into their root and suffix components. Okay, when I read about them dissecting sentences into roots and suffixes, my mind immediately went to modern software development. That's interesting. Yeah, like these ancient Indian scholars, they were basically the original programmers debugging lines of human code. That is actually a brilliant way to look at it, yes. And by applying that logic, were looking for the source code of human language. Right. They realized that if you can break a word down to its absolute root light, its base variable, you can add prefixes or suffixes to build entirely new linguistic functions without breaking the rules of the language itself. Exactly. They were like pulling apart the syntax to figure out the underlying logic, testing the variables, basically just to make sure the cultural system wouldn't crash and the knowledge wouldn't be lost to time. And by doing that, they develop these incredibly complex theories on how sounds and word structures actually work. Right. Which we see in the text. Yeah. You see this in the Negantu from 700 BC and later in Amara Simha's Amarikosa from the 6th century AD. And that one arranged words by their synonyms and homonyms rather than just alphabetical order. Which sounds incredibly complicated. It is. Creating a dictionary requires phonetic marking, meticulously organizing definitions, and trying to predict how future users will actually search for a concept. So it's way more than a clerical task of just typing out a list. Oh, absolutely. If we connect this to the bigger picture, lexicography is really about structuring human thought. A dictionary is a cognitive map. I love that phrase, a cognitive map. It is. It tells you how a culture perceives reality and how it categorizes the physical and abstract world around it. So if you are building a dictionary for computing terms in a traditional language like Mathili, you are essentially drawing a brand new map for a very old territory. You are laying down digital highways over ancient landscapes. That's a great way to visualize it. Which brings us to the core mystery of this source material. Why did Mepheli specifically need this new technological map in 2009? Because reading the forward of this dictionary feels less like a dry academic introduction and more like, I don't know, a geopolitical thriller. It really does, because my belly wasn't just casually evolving in the background. It was locked in this massive existential battle for its identity. Right, and to give you listening some context, this language spans a really massive, vibrant cultural space. It's spoken across the Genetic Plain in the Indian state of Bihar, the Terai region, at the Himalayan foothills in Churkand, and it makes up a significant chunk, about 14% of the entire population of Nepal. Right, it's not a small, isolated language. Not at all. If you look at the official Indian census data over the 20th century, the number of make-lea speakers is absolutely wild. It makes no logical sense whatsoever. No, it's completely erratic. Yeah, the source notes that between 1911 and 1921, the population of speakers supposedly decreased by 0.77%. But then, between 1951 and 1961, the official census says the number of speakers suddenly rocketed up by 22.35%. And the wildly erratic fluctuations just continue from there. The 2001 census officially counted about 12.1 million Methylese speakers. But Professor Singh, who wrote the foreword, he breaks down historical data, geographical expansions and normal population growth over the decades. And he estimates the actual number of speakers is closer to 40 million. 40 million. Let me stop you right there because how on earth can tens of millions of people simply vanish and reappear in official government data? It's pretty shocking. I mean, looking at these census numbers, it's like looking at a volatile stock market chart. It's like the stock being traded as an entire community's linguistic identity, and political record keepers are just mipulating the market. How does a government misplace tens of millions of speakers? Well, it comes down to how data is categorized, and more importantly, the politics of that categorization. According to the source material, these erratic numbers were the result of a very specific structural spread of disinformation. information, like what? Well, there was a narrative pushed that Methili was an exclusive language spoken only by the Brahmin cast in the Mithila region. Okay, but why push that narrative? What's the goal there? The goal, according to the foreword, was to classify Mithili not as a distinct independent language, but merely as a regional dialect of Hindi. Oh, I see. Yeah, by officially counting Mithili speakers as Hindi speakers, it artificially inflated the demographic power and official count of the Hindi language on the census. Wow. So they just relabeled millions of people to boost another language's numbers. But the actual demographic data completely contradicts that Brahman only narrative, right? Completely. The census returns that accurately recorded methilies show massive overwhelming support across all demographics, which completely shatters the caste-exclusive myth. Right. The numbers just don't back it up. Exactly. For instance, the source notes that up to 46.84% of the population in some if the last speaking districts are Muslims. Which means you couldn't possibly have those high returns for Maitrely unless a massive portion of the Muslim population was also registering Maitrely as their mother tongue. Precisely. The data proves it is a language of the broader region deeply embedded across different communities, not just a single cast. And I want to clarify for you listening, our goal here isn't to weigh in on the historical political disputes or, you know, take sides on Indian state politics. Of course not. We are simply unpacking the demographic realities reported in our source material to understand exactly why this dictionary was created. The historical stakes for this language were incredibly high. They really were. And what's fascinating is how the community responded to all of this. The forward describes this long period of being denied constitutional rights as a boon in disguise. Which sounds completely counterintuitive. I mean, how is being erased from the sense as a boon? because the political friction actually fueled a fierce cultural and literary vigor among the speakers. Oh, so it pushed them to fight back. Exactly. The resistance to being reclassified created a powerful sense of unity and advocacy. They wrote more. They published more. They organized. And that resilience eventually pan off, leading to Maitreya Lee being officially included in the eighth schedule of the Indian Constitution. Right. And for those who might not know, being included in the eighth schedule basically means the government legally recognizes the language. Yes, it's a huge milestone. It grants it official status, meaning it can be used in government exams, it receives federal funding for development, and it gets institutional patronage. It is a massive victory for a language of survival. It is the ultimate official validation. But as the source points out, this victory immediately created a new massive problem. Always a new problem. Right. And this brings us to why a computer dictionary matters so much to a community of 40 million people. The forward points out a crucial statistic. While many Mathili speakers are multilingual, you know, they can navigate Hindi, Bhujpuri, Magahi and Bengali, about 25 to 30 percent of Mathili speakers are completely monolingual. Meaning they only speak Mathili. Right. and only a tiny fraction of the total population, maybe three to 5%, can speak English effectively. Okay, let's put ourselves in there shoes for a second. Imagine you are one of those 10 million monolingual mythology speakers. You've just won this massive constitutional victory for your language. A huge moment of pride. Yeah, but then you pick up a smartphone or you sit at a computer terminal in a local government office and every single button, error message, and software setting is in English or Hindi languages you don't read. It's a wall. You don't have a word for internet or browser or download. You were just digitally stranded. You were locked out of the 21st century entirely. Exactly. Being constitutionally recognized is really only half the battle. If your language cannot interact with a microprocessor, it will eventually die out in the modern world. Having access to digital tools is a critical step in a language's survival. Which is what birthed this specific dictionary. authors Thakur and the Jaws took on the monumental task of bringing Mythili online. Oh wait, let me play devil's advocate here for a minute. Sure. I completely get the cultural pride, but wouldn't it be vastly easier to just use the English words? I mean, almost every other language just borrows tech terms, right? French people say le weekend and le computer. Why go through the grueling intellectual labor of inventing brand new Mythili words for things like algorithm? That's a very fair question and it goes back to what we discussed earlier about cognitive maps. Right, mapping the territory. Yes, if you just take an English word like bandwidth and drop it into a monolingual rural community in Bihar, it has no conceptual hook. It's just a meaningless sound. That makes sense. But if you can build a bridge between the new technology and the concepts they already intimately understand, the technology becomes accessible. It feels natural, not foreign. Okay, but that makes perfect sense. And to achieve that, the authors didn't just take the easy way out. Here's where it gets really interesting. They engaged in what I can only describe as linguistic alchemy. Oh, absolutely. They took these cold, binary tech concepts and fused them with classical roots. What's fascinating here is the lexical engineering, the actual building of these new words, is profound. Let's look at some specific translations from the text because they perfectly illustrate this bridge between the ancient and the digital. Yes, I really want to get into the actual words. Let's start with the basics. How do you say computer? In the dictionary, it is translated as sun-gunuk. It utilizes an ancient root related to counting and calculation, elegantly capturing the fundamental nature of a machine that computes data. Okay, sun-gunuk. Simple enough, it's a calculator. But what about something highly conceptual, like a cyber cafe? This is one of the most beautiful translations in the book, in my opinion. translated cybercafe as Sangunik Pangriha. I absolutely love this one. Pangriha literally translates to a traditional drink house or a tavern, right? Yes, exactly. So they took the modern concept of an internet cafe, a communal place where people gather, sit together, and consume data, and they mapped it perfectly onto the ancient cultural concept of a tavern. A drink house of computers. That is just brilliant. It really is. It highlights the sheer intellectual labor required by the lexicographers. Consider the term algorithm, a highly specific mathematical and computational concept that, honestly, most people struggle to define even in English. Oh, for sure. They translated it as abhukti vidi kalpa. It sounds so poetic, but what does it actually mean? Well, vidi relates to a rule, a method, or a procedure, and kalpa implies an order or a rule of practice. So they are framing an algorithm not as some invisible technological magic but as a formal ordered procedure. Which is exactly what an algorithm is. Right. Or take artificial intelligence. They translated that as Kretrem-Pragueh. Let's unpack that because I know Kretrem means artificial or constructed but what is Pragueh? Pragueh is a deep classical term. It relates to wisdom, supreme knowledge or intelligence in the classical almost spiritual sense. Oh wow. Yeah. By combining it with Kretrem they are elevating the technology. They are describing a neural network using vocabulary historically reserved for philosophy and theology. That is fascinating. But you know, it also works for the everyday frustrating parts of tech too, like a bug in the software. When your app crashes, they translated the concept of a bug as DOSHA. Now why use DOSHA? This is a perfect example of repurposing classical knowledge. In traditional Ayurvedic medicine in Indian philosophy, adosha is an imbalance or a fundamental flaw in a system's constitution. Oh my gosh. So a computer bug isn't just a literal insect inside the machine like we use in English. They've translated it to mean a fundamental imbalance in the software's harmony. You are basically diagnosing the code with a medical ailment. That is such a vivid way to understand a glitch. It is perfectly vivid and look at bandwidth. They translated it as vahini-vistara. Which means the expanse or the spread of the channel. Yes. Vahini being a channel or a river and Vistara being the expanse. OK, I see where this is going. Think about it. If you are a monolingual farmer trying to understand why your agricultural weather app is loading slowly, the English word bandwidth means absolutely nothing. Right, it's just a foreign sound. But Vahini Vistara. A farmer intimately understands the concept of a channel's width, restricting the flow of water to a field. By using that term, the dictionary maps the invisible flow of digital data onto the physical reality of agricultural irrigation. That is, they took a 21st century operating system and explained it using a 2,000 year old agrarian landscape. That is the magic of this dictionary. And you see this deep contextual mapping throughout the source material. For instance, the concept of Boolean logic, you know, the true and false one in zero foundation of all computing code. Right. The absolute basics of programming. They translated it as Boolean Tarka. They keep the name of the mathematician George Bool, but they attach it to Tarka. And Tarka is the classical Indian philosophical system of logic and rigorous debate, isn't it? Right. It immediately signals to the user. We are dealing with absolute structural logic here. It grounds the foreign mathematics in a familiar intellectual tradition. Or what about data mining? They translated it as Datash Uttkana, And Utkana is the ancient human act of physically excavating the earth, like digging a well or a mine. Yes, physical excavation. So they just applied the physical act of digging dirt to the digital excavation of data. It's a perfect metaphor. It truly is. And to make sure all of this was incredibly accessible, the authors didn't just stop at the words. They formatted the dictionary using multiple scripts. Right, because reading the word is just as important as knowing the word. Exactly. Every single entry includes the international phonetic alphabet, so a linguist from anywhere in the world knows exactly how to pronounce the Myphilic word. That's super helpful. It also includes the widely used Devanagari script, which is what Hindi and Sanskrit are usually written in today. And crucially, it includes the traditional Mithalakshara, also known as the Tirhuta script. Which is a huge deal, because they honored the language's ancient orthography, meaning they literally printed the book using the traditional written script that the Mithili people have been using for centuries rather than forcing them to only read a modernized alphabet. Yeah, this dictionary proves that traditional languages are not static museum pieces waiting to collect dust. They are highly elastic. They really are. When given the proper academic rigor and when they secure that official patronage, they can expand to encompass any modern concept. The challenge isn't that a language is too old to understand the internet. The challenge is merely finding the right lexicographers to do the heavy lifting of building the bridges. So if we zoom all the way out, what does this actually mean for us for you listening to this deep dive right now? I think it fundamentally changes how we are supposed to look at a dictionary. I completely agree. Because the English May Selly Computer Dictionary is far more than just a thick reference book sitting on a desk somewhere in New Delhi. It is an active tool of survival. It is an active defiance, honestly. It represents a vibrant population of tens of millions of people who flat out refuse to let their mother tongue fade into obscurity simply because the world decided to invent microchips and fiber optic cables. By defining these words, they are ensuring that their language has a permanent seat at the digital table. It's taking that ancient linguistic x-ray machine, the exact same logical framework that scholars used to dissect Vedic texts thousands of years ago and plugging it directly into a modern server farm. It really is. It proves that an ancient language can absolutely hold its own in a modern cybercafe or rather a Sangunik pan-gre-ha. And frankly, it serves as a master blueprint for thousands of other regional and indigenous languages around the globe that are currently facing the exact same digital threshold. Which leaves us with a final thought to mull over as we wrap up today. As artificial intelligence and global tech continue to advance at lightning speed, as neural networks get better and better at live translation, will local languages have to constantly race to invent new dictionaries to keep up with the ever-changing tech vocabulary? Oh, that's a fascinating prospect. Or will our technology eventually evolve to the point where it natively speaks and understands the nuance of every ancient language perfectly without us needing humans to build a bridge at all? Will the machine eventually learn the etymology of Doshua or Pragya so fluently that the burden of translation is finally lifted from the humans who speak it? Definitely something to think about. But until that day comes, we desperately need the dictionary makers. We need the map builders. Thanks for taking the deep dive with us.

← Transcript index