MACHINE ASR ACCESSIBILITY AID

Making_Videos_with_AI_Video.mp4

Not an editorially verified transcript. This text was generated automatically from the preserved recording and may contain recognition, language-detection, spelling, segmentation or name errors. Consult the source recording for authoritative content.
Collection
Part 9 · VIDEHA MITHILA MAITHILI DISCUSSION CRITICISM SERIES PART 9
Status
asr-draft
Human verified
No
Editorial review
not-reviewed
ASR model
small
Detected language
en (0.997737)
Duration
55:05
Source
Open preserved recording

Timestamped machine output

  1. 0:00–0:06Making videos with AI. Teach yourself series, written by Gage and Rathakar.
  2. 0:06–0:08Preface.
  3. 0:08–0:12This book is for everyone who wants to say something through video but has no camera,
  4. 0:12–0:18no studio, and no training in editing. Artificial intelligence AI has now made all three possible
  5. 0:18–0:22with an ordinary mobile phone or computer. What once required equipment worth hundreds of
  6. 0:22–0:27thousands and a team of dozens can now be done with a written instruction, called a prompt.
  7. 0:27–0:33but let one thing be clear at the very outset. AI is no magic wand, it is a tool, like an axe,
  8. 0:33–0:39like a pen, the axe cuts the wood, but which tree to fell and what house to build, the carpenter
  9. 0:39–0:44decides, in the same way, AI will make the visuals and generate the voice, but what to say,
  10. 0:44–0:49why to say it, and for whom, the answers to these three questions remain with you,
  11. 0:49–0:53this book will teach you to handle the tool, the craftsmanship will remain your own.
  12. 0:53–0:59This book is written in the teacher-self style, that is, no teacher or training institute is
  13. 0:59–1:04needed. Each chapter is one step, begin with the first chapter, move forward in order,
  14. 1:04–1:08and be sure to do the exercise give at the end of every chapter. Reading and doing must
  15. 1:08–1:14walk together, only then does learning happen. One more thing, all the illustrations in this
  16. 1:14–1:20book are original explanatory figures, not copies of any company's actual screens. There
  17. 1:20–1:25There are two reasons. First, the screens of AI tools change every few months, so a real
  18. 1:25–1:28screenshot would be outdated before the book left the press.
  19. 1:28–1:33Second, once you understand the principle, any new screen will feel familiar. If you
  20. 1:33–1:37memorize a screen, every change will leave you stranded. The figures in this book show
  21. 1:37–1:42the structure that applies to nearly every tool alike. Where the prompt box sits, what
  22. 1:42–1:47the settings contain, what a timeline looks like. Once you grasp this framework, you
  23. 1:47–1:52open any new tool and work it out on your own, and that is the true meaning of teach yourself.
  24. 1:53–1:58The absence of technical books in Mathaly has long been a sore point, and this book was first
  25. 1:58–2:03written in Mathaly. This English edition carries the same conviction, whatever you learn here,
  26. 2:03–2:09use it to put your own language, your own region, and your own stories on screen, the tales,
  27. 2:09–2:14psalms, paintings, history and festivals of Mithila, and of every homeland like it, are
  28. 2:14–2:17are all waiting for their videos.
  29. 2:17–2:21CHAPTER 1 WHAT IS AI VIDEO?
  30. 2:21–2:22Artificial Intelligence
  31. 2:22–2:28AI is the capacity of a computer to do human-like work, understanding language, recognizing images
  32. 2:28–2:34and now, creating images and video, when we write to an AI, show two children playing
  33. 2:34–2:38by a pond, and it produces a moving picture of that very scene on its own.
  34. 2:38–2:40This is called Generative AI.
  35. 2:40–2:42To generate means to bring forth.
  36. 2:42–2:47This AI does not fetch a video from somewhere, it composes a new video that never existed
  37. 2:47–2:49anywhere before.
  38. 2:49–2:51How is this possible?
  39. 2:51–2:56Understand it in simple terms, these AI models have been shown crores of images and videos,
  40. 2:56–3:01having seen so much, they have learned what an egret looks like, how water ripples, what
  41. 3:01–3:05color the morning light takes, and how a walking person's feet rise and fall, when
  42. 3:05–3:09you write an egret beside a pond at dawn, the model builds such a scene from its learned
  43. 3:09–3:14knowledge, just as a painter, who has seen thousands of ponds can now draw a pond from
  44. 3:14–3:17memory without one before their eyes.
  45. 3:17–3:20There are chiefly four routes to making video with AI.
  46. 3:20–3:25First, text to video, you only write, and the AI creates the entire scene.
  47. 3:25–3:30This is the most astonishing route, but it offers somewhat less control, the scene
  48. 3:30–3:33that arrives may not match your imagination exactly.
  49. 3:33–3:39Second, image a video, you supply one still image, a photograph or an AI generated picture,
  50. 3:39–3:41and the AI sets it in motion.
  51. 3:41–3:43Here the control is greater
  52. 3:43–3:46because you yourself chose the opening frame.
  53. 3:46–3:48Third, avatar video.
  54. 3:48–3:50A digital speaker avatar sits on screen
  55. 3:50–3:53and speaks your written text, like a news anchor,
  56. 3:53–3:55for lessons, announcements and lectures.
  57. 3:55–3:57This is extremely useful.
  58. 3:57–3:59Fourth, AI assisted editing.
  59. 3:59–4:01Here the footage is your own recording,
  60. 4:01–4:03but the cutting and joining,
  61. 4:03–4:06the subtitles, the background removal.
  62. 4:06–4:08All this the AI does.
  63. 4:08–4:12This book covers all four routes, because in practice a good video is usually a blend of
  64. 4:12–4:14them.
  65. 4:14–4:16Now understand what AI cannot do.
  66. 4:16–4:21AI does not yet produce long videos in one go, it typically makes short clips of 5 to
  67. 4:21–4:2420 seconds, which are joined to build a longer video.
  68. 4:24–4:29AI may seem sometimes contain errors, a hand grows too many or too few fingers, written
  69. 4:29–4:34letters come out garbled, a character's face changes between one clip and the
  70. 4:34–4:36next, and the biggest point of all.
  71. 4:36–4:41AI knows nothing of Mathila, of Mathili, or of what lies in your heart. It will give only as
  72. 4:41–4:46much as you know how to ask. That is why half of the spoke is devoted to how to ask, that is,
  73. 4:46–4:52planning and prompts. Exercise. On paper, write down three subjects for videos you would like
  74. 4:52–4:58to make, for each, write one line, who will watch this video and what will they gain from watching
  75. 4:58–5:02it. Chapter 2. Getting Ready. What Do You Need?
  76. 5:02–5:08No costly machinery is needed to make video with AI, all that is required is this.
  77. 5:08–5:15First, a device, a smartphone is sufficient, a computer or laptop adds convenience, writing
  78. 5:15–5:20prompts, managing files and editing are easier on a large screen, no specially powerful computer is
  79. 5:20–5:25needed because the video is made not on your device but on the company servers, your device
  80. 5:25–5:31merely sends the instruction and downloads the finished video. Second, the internet,
  81. 5:31–5:35But the better the speed, the shorter the weight, video files are large, so downloading will
  82. 5:35–5:37consume data.
  83. 5:37–5:39Budget a few gigabytes a month.
  84. 5:39–5:44Third, an email account, nearly every AI tool requires an account, and most allow direct
  85. 5:44–5:48sign-in with a Google email Gmail, one suggestion.
  86. 5:48–5:52Keep a separate email for this work, so that the newsletters of AI tools do not flood
  87. 5:52–5:53your main inbox.
  88. 5:53–5:56Fourth, an understanding of credits.
  89. 5:56–5:58Most AI video tools run on a credit system.
  90. 5:58–6:04A credit is a kind of coupon, making one video deduct some credits, a free account receives
  91. 6:04–6:09a small allowance, daily and some tools monthly and others, once only in a few.
  92. 6:09–6:16Paid plans bring more credits, higher quality such as 1080p or 4k, videos without watermarks
  93. 6:16–6:18and longer clips.
  94. 6:18–6:22Adapt a practical policy here, which I call cheap first, then deep, while learning a
  95. 6:22–6:27new tool, run small experiments on the free credits, learn to write prompts, learn
  96. 6:27–6:33the tools temperament when it feels time for serious work, say, running a YouTube channel,
  97. 6:33–6:35then take a paid plan on one tool.
  98. 6:35–6:39Buying plans on every tool is wasteful, one or two suffice.
  99. 6:39–6:405.
  100. 6:40–6:41A system for files.
  101. 6:41–6:45It sounds a small matter, but the experience know how work drowns without order.
  102. 6:45–6:49Make a separate folder for each video project on your computer or phone.
  103. 6:49–6:55Inside it keep four subfolders, script scripts and prompts, clips raw AI made video, voice
  104. 6:55–7:00Peace music voiceover and music and final the finished video give every file a meaningful
  105. 7:00–7:01name.
  106. 7:01–7:05A file called clip a one-pond on will still be recognizable six months later.
  107. 7:05–7:08Video three final new two will not.
  108. 7:08–7:14And finally, patience, AI tools are sometimes busy, sometimes give strange results, sometimes
  109. 7:14–7:17fail entirely to understand your perfectly good prompt.
  110. 7:17–7:18All this is natural.
  111. 7:18–7:23Those who persist through three or four attempts learn, those who quit at the first odd
  112. 7:23–7:25result are left behind.
  113. 7:25–7:30Exercise, create the folder system described above on your device, make a new email account
  114. 7:30–7:35if you wish, and open the website of any one AI video tool simply to see what the free plan
  115. 7:35–7:40offers, do not make anything yet, only look.
  116. 7:40–7:44Chapter 3, Planning the video, from idea to script
  117. 7:44–7:48The biggest mistake beginners make is to open the tool straight away and start pressing
  118. 7:48–7:52a video made without a plan looks exactly like a house built without a drawing.
  119. 7:52–7:58That is why this chapter comes before the tools. Before every video, write the answers to three
  120. 7:58–8:05questions. 1. What is the purpose? To teach say how Matubati painting is made, to tell a story,
  121. 8:05–8:11to show scenes of a village or to an else. 1. Video, 1 purpose. Hold to this rule.
  122. 8:11–8:17A video that contains everything contains nothing. 2. Who is the audience? Children,
  123. 8:17–8:23students, expatriates far from home, curious outsiders, or the general viewer, the audience decides
  124. 8:23–8:28how simple the language should be, how long the video should run, and what the visuals should
  125. 8:28–8:33look like, bright colors and quick movement for children, stillness and gravity for adults.
  126. 8:33–8:39Free, How Long, In The Beginning, Make Videos Of 30 Seconds To 2 Minutes. A short video is
  127. 8:39–8:45easier to make, easier to fix, and better liked by today's viewer, remember, AI produces clips
  128. 8:45–8:50of 5 to 10 seconds, so a 1 minute video means 6 to 12 clips.
  129. 8:50–8:55Now the script, a script is nothing complicated or literary, it is simply a two column table,
  130. 8:55–9:00in the left column, what is seen the visual, in the right column, what is heard the voice
  131. 9:00–9:05or subtitle, for a 30 second video, 5 or 6 rows are enough.
  132. 9:05–9:09Here AI can help in a second role, as script assistant, give your subject to a chatbot
  133. 9:09–9:15such as Claude, ChatGPT, or Gemini and AskIt to draft the script, but the AskInG needs
  134. 9:15–9:21skill, merely saying write a video script on Madhubani will fetch a generic, lifeless text,
  135. 9:21–9:27ask like this. I am making a 60 second YouTube video on Madhubani painting, audience, young
  136. 9:27–9:32Indian viewers who have heard the name but know little more, voice, simple and warm,
  137. 9:32–9:37I need a script in two columns, on the left, a description of each visual which I will
  138. 9:37–9:41give to an AI video tool. On the right, the voice of her text.
  139. 9:41–9:476 scenes, 10 seconds each. Notice, subject, duration, audience, language,
  140. 9:47–9:51format, and use are all stated. The clearer the ask, the better the result.
  141. 9:51–9:55This same principle will apply later to video prompts.
  142. 9:55–9:59Do not accept the chatbot's script with your eyes closed, read it, test it against
  143. 9:59–10:03your own knowledge. Are the facts right? Does the language sound like your own?
  144. 10:03–10:08any line that falls flat, or have it rewritten make the third scene more tender, shorten
  145. 10:08–10:13the last line, the script is your signature, the AI is only a scribe.
  146. 10:13–10:15Exercise.
  147. 10:15–10:19Take one of the three subjects you chose in chapter 1, write the answers to the three
  148. 10:19–10:22questions purpose, audience, duration.
  149. 10:22–10:27Then, using the detailed style of asking shown above, have a chatbot draft a two-column
  150. 10:27–10:32script and make at least two corrections to it with your own hand.
  151. 10:32–10:36Chapter 4. The Storyboard. A Map of Scenes.
  152. 10:36–10:41The script is a map of words, the storyboard is a map of scenes, a storyboard lays out one
  153. 10:41–10:46sketch, a rough drawing or description, for every scene of the video in order. In filmmaking
  154. 10:46–10:51this method is a century old, and in the AI age its importance has not shrunk but grown.
  155. 10:51–10:52Why?
  156. 10:52–10:57Because AI needs a separate, clear instruction for every clip, and the storyboard is precisely
  157. 10:57–10:59that list of instructions.
  158. 10:59–11:02Do not worry, no drawing skill is required.
  159. 11:02–11:07Round faces, stick figures, and arrows are enough, if you would rather not sketch at all,
  160. 11:07–11:12write three lines for each scene, what is seen, what the camera does, and how many seconds.
  161. 11:12–11:16Learn a little of the camera's language, for these very words will serve you later in
  162. 11:16–11:18prompts.
  163. 11:18–11:23Wide shot, the whole scene from a distance, for establishing the place, like a full view
  164. 11:23–11:24of the village.
  165. 11:24–11:29Close up, from near, for feeling and fine detail, like the steam over a cup of
  166. 11:29–11:34teal on the hearth. Tracking shot, the camera moves along with the subject, like following
  167. 11:34–11:40a child on the way to school. Drone view, from above as a bird sees, like a sweeping view
  168. 11:40–11:47of pond and fields. Dallied in slash out, the camera slowly draws near or pulls away,
  169. 11:47–11:53for emotional weight. Static shot, the camera stays in one place, for calm, settled scenes.
  170. 11:53–11:58Keep one simple formula for the order of scenes, establish, develop, resolve. The
  171. 11:58–12:02The first scene tells where we are established, usually a wide shot.
  172. 12:02–12:08The middle scenes bring the subject close develop, close ups, tracking, the last scene gathers
  173. 12:08–12:12it all, a feeling, a message, or title card resolve.
  174. 12:12–12:16Vigor 2 shows the storyboard of a 30-second video called My Village.
  175. 12:16–12:22See how six scenes travel from dawn to dusk, morning mist establish, tea, school and pond
  176. 12:22–12:24develop, lamps and title resolve.
  177. 12:24–12:28Beside each scene its duration and camera note are written, each of these scenes will, further
  178. 12:28–12:31on, become one prompt.
  179. 12:31–12:36While making the storyboard, attend to continuity, if the first scene is morning, the second must
  180. 12:36–12:40not suddenly be night, if a character wears a red kurta, the red kurta must appear in
  181. 12:40–12:45every scene, and this must be written into every prompt, because AI does not remember
  182. 12:45–12:50the previous clip, write the characters' description on a separate sheet, a character card, and
  183. 12:50–12:55pasted word for word into every prompt. The small trick is the simplest way to keep the
  184. 12:55–13:02clips consistent. Exercise. Turn your script into a 16 storyboard. For each scene, write
  185. 13:02–13:06the visual description, the camera, the duration. If there is a character, make a character
  186. 13:06–13:14card dress, age, appearance, in three lines. Chapter 5. The Art of Prompt Writing.
  187. 13:14–13:19A prompt is the written instruction you give to the AI. It is the most valuable skill
  188. 13:19–13:24of the AI age, and the happy news is that it demands no technical knowledge, only clarity
  189. 13:24–13:29of language and language is our home ground. See the difference between a poor prompt and
  190. 13:29–13:35a good one. Poor village scene, this tells that AI nothing, a village of which country,
  191. 13:35–13:40which season, day or night, what is happening, it will invent something from its own mind,
  192. 13:40–13:45most likely some European or placeless village, now the good one, a village in North India,
  193. 13:45–13:50Early morning, light missed over the fields, thatched and tiled houses.
  194. 13:50–13:55A banyan tree in the distance, camera panning slowly to the right, soft gold and light, realistic
  195. 13:55–13:57documentary style.
  196. 13:57–14:00Now the AI holds the complete picture.
  197. 14:00–14:03Keep in mind the six-part formula shown in Figure 3.
  198. 14:03–14:07Subdecked, who or what is central, two egrets.
  199. 14:07–14:09Action, what are they doing?
  200. 14:09–14:10Catching fish.
  201. 14:10–14:15Scene. Where and when, upon full of lotus, at dawn.
  202. 14:15–14:19Camera. What does the camera do? Slowly moving in.
  203. 14:19–14:22Light. What kind of light? Golden morning light.
  204. 14:22–14:26Style. What look? Documentary, realistic.
  205. 14:26–14:30Let every prompt carry all six parts and keep roughly this order.
  206. 14:30–14:35Subject and action first, style last, some useful words of style,
  207. 14:35–14:42realistic, cinematic, animation, like a watercolor painting, like old film, documentary.
  208. 14:42–14:48The question of language, most AI video tools understand English prompts best, some also
  209. 14:48–14:52follow Hindi and other languages, the practical path is this, they can script in your own
  210. 14:52–14:57language and shape the final prompt in English, here a chatbot helps, give it your scene
  211. 14:57–15:02description in your language and say, turn this into an English video prompt with
  212. 15:02–15:096-part subject, action, scene, camera, light, style, thus the imagination stays yours,
  213. 15:09–15:12only the translation is mechanical.
  214. 15:12–15:16There is also the practice of the negative prompt, where you state what you do not want, for
  215. 15:16–15:22example, no blur, no distorted hands, no text on screen, no watermark, some tools give
  216. 15:22–15:26it a separate box, in others it is added to the main prompt.
  217. 15:26–15:31Now the most important principle of all, the improvements cycle figure 4, the first
  218. 15:31–15:36prompt rarely yields the video of your wishes, and this is no failure, it is the method itself.
  219. 15:36–15:41Look at the result, name the fault, change just that much in the prompt, and generate again.
  220. 15:41–15:46Did the egret come out too large? At a small egret, does the scene look garish?
  221. 15:46–15:51Change to soft, gentle light, usually within 3-5 cycles a usable clip arrives.
  222. 15:51–15:56Change only one or two things per cycle. Change everything at once and you will never
  223. 15:56–16:00know which change did the work. Keep saving your good prompts in one file,
  224. 16:00–16:03The you of six months hence will thank the you of today.
  225. 16:04–16:09Exercise, for the first scene of your storyboard, write one complete prompt using the six-part
  226. 16:09–16:13formula in your own language, then make its English form yourself or through a chatbot.
  227. 16:13–16:15Keep both in your script folder.
  228. 16:16–16:20Chapter 6. Text-to-video tools, the first clip.
  229. 16:20–16:24The moment has come to take the tool in hand. In this chapter we understand the
  230. 16:24–16:28common structure of text-to-video tools and make the first clip.
  231. 16:28–16:33First, an introduction to the tools, this field changes at great speed.
  232. 16:33–16:37Every few months a new model arrives and the old ones grow stronger.
  233. 16:37–16:43At the time of writing 2026 the leading names are, Google's BO, OpenAI's Sora, Kling,
  234. 16:43–16:48Oneway, Pixverse, Seedance, Haleuo, Luma, and Pika.
  235. 16:48–16:53Each has its own temperament, one excels at realistic scenes, another at stylized or
  236. 16:53–16:57artistic ones, another is faster and cheaper, the names will keep changing, but the
  237. 16:57–17:03structure, described below, remains nearly the same in all of them, so learn the structure,
  238. 17:03–17:09not the names. Look at figure 5, in almost every tool you will find these 5 things.
  239. 17:09–17:151. The prompt box, a large empty field where you write your instruction, this is the heart
  240. 17:15–17:16of the tool.
  241. 17:16–17:222. Model selection. A single company offers several models, new and powerful costly,
  242. 17:22–17:26Older or fast cheap, choose the cheap model while learning.
  243. 17:26–17:29The good model for final, publishable clips.
  244. 17:29–17:313.
  245. 17:31–17:32Aspect Ratio
  246. 17:32–17:3816 colon 9 for YouTube and television wide, 916 for Reels, Shorts, and Status Upright,
  247. 17:38–17:431 colon 1 square, the side at the outset where the video will go, because changing
  248. 17:43–17:45the ratio later crops the scene.
  249. 17:45–17:464.
  250. 17:46–17:47Duration
  251. 17:47–17:50Usually options of 5, 8 or 10 seconds.
  252. 17:50–17:54A shorter clip costs fewer credits and carries fewer errors.
  253. 17:54–17:555.
  254. 17:55–18:00The generate button and the results area, press the button, wait from a few seconds to a few
  255. 18:00–18:03minutes and the clip appears in the results area.
  256. 18:03–18:05Download it from there.
  257. 18:05–18:08Now the method for the first clip, step by step.
  258. 18:08–18:091.
  259. 18:09–18:12Open the tools website and create an account with your email.
  260. 18:12–18:132.
  261. 18:13–18:16Check the free credits, how many, and when they renew.
  262. 18:16–18:173.
  263. 18:17–18:22Use the cheap slash fast model, 16 colon 9 ratio and the shortest duration.
  264. 18:22–18:234.
  265. 18:23–18:26Pace the prompt you built in chapter 5.
  266. 18:26–18:275.
  267. 18:27–18:28Press generate and wait.
  268. 18:28–18:296.
  269. 18:29–18:33Watch the finished clip in full, not once, but two or three times.
  270. 18:33–18:37Watch the hands, the faces, any lettering, the way things move.
  271. 18:37–18:387.
  272. 18:38–18:41Download it into your clip's folder under a meaningful name.
  273. 18:41–18:45Even if the clip is imperfect, it will serve for comparison.
  274. 18:45–18:468.
  275. 18:46–18:51the improvement cycle. Name the fault, refine the prompt, generate again.
  276. 18:51–18:56A few practical tricks. Generate two or three clips from the same prompt, AI gives a somewhat
  277. 18:56–19:01different result each time, and you pick the best. A clip whose main subject is right
  278. 19:01–19:06but whose edges carry faults can often be saved by cropping in the edit, do not discard
  279. 19:06–19:11it at once and set yourself a daily usage limit before the credits run dry or weeks
  280. 19:11–19:16credits will vanish in one enthusiastic evening. Exercise. Make the clip for the first scene
  281. 19:16–19:21of your storyboard, running at least three improvement cycles. Save the final clip together
  282. 19:21–19:28with the prompt that produced it. Chapter 7. From Enage to Video, The Road of Greater Control.
  283. 19:29–19:34Text to video has won in convenience. You have no hold over the opening scene. Whatever the AI
  284. 19:34–19:40makes, it makes. The remedy is image of video. First, prepare a still image that is exactly
  285. 19:40–19:45to your mind, then tell the AI to set this image in motion. Hear the look of the scene,
  286. 19:45–19:50the character's face, the clothing, all are fixed in advance. The AI only adds the movement.
  287. 19:51–19:57Where will the image come from? Free sources. First, your own photographs, your village,
  288. 19:57–20:02your festivals, nature, art, animating a photograph you took yourself is the most
  289. 20:02–20:07authentic or out. Remember, the photograph must be your own or used with the owner's permission,
  290. 20:07–20:12and before animating a photograph of a living person, be sure to take their consent.
  291. 20:12–20:16This is both courtesy and part of the ethics, described in chapter 15.
  292. 20:17–20:23Second. AI-generated images. There are separate tools for image generation and most video tools
  293. 20:23–20:28include an image-making feature. The image prompt follows the same six-part formula,
  294. 20:28–20:33only in place of camera movement, describe the composition. Images are cheap and quick to make.
  295. 20:33–20:39so run your improvement cycle on the image first. Generating 10 images and choosing the best is
  296. 20:39–20:45far cheaper than generating 10 videos. Third, scanning your own artwork. A work of Mathila
  297. 20:45–20:49painting if it is your own or you hold the rights can be scanned and given gentle motion.
  298. 20:49–20:54A fish stirring, the line ornament shimmering, exercise great restraint here, the dignity of
  299. 20:54–20:59traditional art lies in its stillness, so keep the motion extremely slight, right gentle,
  300. 20:59–21:06slow, subtle motion in the prompt. The method is simple. To use the image video option in the tool,
  301. 21:06–21:11upload your image and write the motion instructional on side. This prompt now carries less seen
  302. 21:11–21:17description and more motion description, what moves, in which direction, how fast, and what the camera
  303. 21:17–21:23does. For example, gentle ripples on the water, the egrets wings moving slowly, camera moving
  304. 21:23–21:29forward very slowly, nothing else changes, that last phrase, nothing else changes, matters,
  305. 21:29–21:34without it the AI will sometimes transform the whole scene. This is also the best remedy for the
  306. 21:34–21:40problem of character consistency. If the same character appears again and again in your video,
  307. 21:40–21:45first create or choose one excellent image of that character. And for every scene animate that
  308. 21:45–21:50same image with different motion instructions, some advanced tools offer a feature called character
  309. 21:50–21:56reference or consistent character where the character's image, given once, appears in every clip.
  310. 21:56–22:02Look for this feature in your tool. Exercise. Take any image or own photograph or AI generated
  311. 22:02–22:08and make three different motion versions of it. In one, only the camera moves. In the second,
  312. 22:08–22:14only some element of the scene moves. In the third, both, compare the three, which feels most natural.
  313. 22:16–22:22Chapter 8 Voice, Voice over and text to speech. The soul of a video lives not in the visuals
  314. 22:22–22:28but in the voice. A viewer will forgive a blurry scene, but will close the video at a bad voice,
  315. 22:28–22:33so read this chapter with care, and for speakers of mathily and other less-served languages,
  316. 22:33–22:38it holds some special advice. There are two ways to add voice, record your own,
  317. 22:38–22:44or have AI generate a text-to-speech, TTS for short. First, your own voice,
  318. 22:44–22:49because for mathily and languages like it, this remains the best route, the pure pronunciation,
  319. 22:49–22:55the natural cadence, the rise and fall of feeling. No machine yet renders these as well as a native
  320. 22:55–23:00speaker and no studio is needed. A smartphone microphone today is quite good enough, follow
  321. 23:00–23:07a few rules, record in a quiet room fan off. Windows shut. Night or early morning is best.
  322. 23:07–23:12Hold the phone about a hand span from your mouth, speak standing or sitting upright.
  323. 23:12–23:17The voice stays open. Keep the script before you, but speak as if telling, not reading,
  324. 23:17–23:22and record paragraph by paragraph rather than all in one take, when you slip, you
  325. 23:22–23:29re-speak only that much. AI can polish a recorded voice. Many tools, usually named enhanced voice,
  326. 23:29–23:34or studio sound in editing apps strip the background, noise and give the voice a studio finish,
  327. 23:34–23:39use it without fail, the difference between a plain recording and an enhanced one will astonish you.
  328. 23:40–23:45Now text to speech figure six, here you type the text, choose the language and the speaker
  329. 23:45–23:52female or male, young or mature, adjust pace and pitch, and download an MP3 file for Hindi,
  330. 23:52–23:57English, and other major languages this facility is very mature, for mathily the situation is
  331. 23:57–24:03improving. Some tools have begun to offer a mathily voice, and tools built for Indian languages are
  332. 24:03–24:08the most likely place to find one. Search in your tool. If mathily appears, first test it with
  333. 24:08–24:14a short passage to hear how pure the pronunciation is. If no mathily voice is available, two remedies,
  334. 24:14–24:20The first and best. Your own voice by the method above. The second. Making do with the hindi voice.
  335. 24:20–24:25Write the text phonetically, listen and adjust the spelling until it sounds right.
  336. 24:25–24:30Keep the pace a little slow. Use short sentences. The result will not carry a fully
  337. 24:30–24:35mathal cadence, but it will serve. Remember, this is a compromise, not an ideal.
  338. 24:35–24:40Wherever feeling at purity matter poetry, stories, children's material, give your own voice.
  339. 24:40–24:47A word on a newer facility, voice-cloning, some tools, from a few minutes of your recording,
  340. 24:47–24:52build a digital replica of your voice, which will then read any text in your own tones,
  341. 24:52–24:56for content in a less-served language this is attractive, teach the tool your voice
  342. 24:56–25:01wants, and the voice-overs of many videos can be made, but two iron rules, clone only
  343. 25:01–25:06your own voice, imitating anyone else's voice without written permission is absolutely
  344. 25:06–25:12forbidden, and where a cloned voice is used in a video, disclosing it is good practice.
  345. 25:12–25:16The joining of voice and visuals will happen at the edit Chapter 11, so keep the voice file
  346. 25:16–25:22separately in the voice music folder, make the voiceover first and the video clips after.
  347. 25:22–25:26This order is wise, because hearing the length of the voice tells you how many seconds each
  348. 25:26–25:27scene needs.
  349. 25:27–25:28Exercise.
  350. 25:28–25:33Make the voiceover of your script both ways, once recorded in your own voice, once through
  351. 25:33–25:38a TTS tool, listen to both with your eyes closed, which sounds more like you.
  352. 25:38–25:40Why?
  353. 25:40–25:41Chapter 9.
  354. 25:41–25:44Avatar Videos, The Digital Speaker.
  355. 25:44–25:49Imagine, every fortnight you must make an announcement video for a journal's new issue,
  356. 25:49–25:51or 50 lectures for a course.
  357. 25:51–25:55Camera, lighting, dress, recording, every single time.
  358. 25:55–25:56Impossible.
  359. 25:56–25:58The remedy is the Avatar video.
  360. 25:58–26:03A digital human sits on screen and speaks your written text, lips moving, eyes blinking,
  361. 26:03–26:06hands gesturing, like a news anchor.
  362. 26:06–26:11The well-known tools of this class are Hei-jen, Synthesia, and others, and the class itself
  363. 26:11–26:12is growing fast.
  364. 26:12–26:15The structure is nearly the same in all figure 7.
  365. 26:15–26:161.
  366. 26:16–26:18Choose the avatar.
  367. 26:18–26:22Tools carry hundreds of ready-made avatars, of different ages, dress, and bearing.
  368. 26:22–26:27Some tools also let you build your own avatar from a photograph or a short video.
  369. 26:27–26:32That is, you remain on screen without recording each time, if you make your own avatar.
  370. 26:32–26:37the same consent rule given for voice cloning, only your own likeness, never another's.
  371. 26:37–26:392.
  372. 26:39–26:44Give the script, it can take two forms, written text which the tool will speak through TTS,
  373. 26:44–26:49or your own recorded audio file which the avatar will lip sync, for mathily the second
  374. 26:49–26:54road is usually better, upload your own mathily voiceover, and the avatar speaks it, the
  375. 26:54–26:56lip sync you get is surprisingly good.
  376. 26:56–26:573.
  377. 26:57–27:02Choose the background and layout, a library, an office, a plain color or an image of your
  378. 27:02–27:03own.
  379. 27:03–27:08Choose the aspect ratio 16 colon 9 or 916, and generate.
  380. 27:08–27:14The beauty of the avatar video lies in its practicality, not in spectacle, news-style
  381. 27:14–27:20presentation, journal announcements, lesson explanations, introductions of an institution.
  382. 27:20–27:24Information videos, in all these it is excellent for the feeling-laden delivery of story and
  383. 27:24–27:27poetry, a human is still better.
  384. 27:27–27:28One tip on presentation.
  385. 27:28–27:32Do not keep the avatar on screen for the whole video without relief.
  386. 27:32–27:37In between, show related scenes, images or text cards called B-roll while the avatar's
  387. 27:37–27:39voice runs beneath.
  388. 27:39–27:40The video comes alive.
  389. 27:40–27:45This weaving happens at the edit, keep the avatar clip and the B-roll clips as separate
  390. 27:45–27:46files.
  391. 27:46–27:50And yes, when the avatar in a video looks human, the viewer has a right to know it
  392. 27:50–27:52is a digital speaker.
  393. 27:52–27:54Write one line in the description.
  394. 27:54–28:01stays intact and trust as a channel's real capital. Exercise, on the free plan of any avatar tool,
  395. 28:01–28:06make a 30-second introduction video, subject, an introduction to my village or an introduction
  396. 28:06–28:11to a favorite book, try both methods, type text and uploaded voice.
  397. 28:12–28:19Chapter 10 Music and sound effects. Visuals for the eye, voice for the ear and music,
  398. 28:19–28:23for the heart, the same scene feels lifeless without music and comes alive with the right
  399. 28:23–28:27score, but with music comes the greatest danger of all.
  400. 28:27–28:31Copyright, put someone's song in your video without permission and YouTube can block the
  401. 28:31–28:36video, others can claim its earnings, and the channel can be penalized, so rule one,
  402. 28:36–28:41never a famous film's song or commercial recording, unless you hold written permission.
  403. 28:41–28:43Then where will the music come from?
  404. 28:43–28:45Free lawful sources.
  405. 28:45–28:50First, AI generated music, there are now tools that compose music from a written
  406. 28:50–28:56description. Write slow, tender, flute-led, 60 seconds and the music is ready. Some tools
  407. 28:56–29:01even build a full song, voice included, from your lyrics. Music made this way for your
  408. 29:01–29:06own video is generally safe to use, but read each tool's license once, especially whether
  409. 29:06–29:10commercial use including YouTube monetization is permitted.
  410. 29:10–29:16Second, copyright-free music libraries. YouTube's own audio library inside YouTube studio is
  411. 29:16–29:22free and safe. Beyond it, many websites offer freely licensed music, some entirely free,
  412. 29:22–29:27some on the condition of attribution. If attribution is required, do not forget to write the musician's
  413. 29:27–29:31name in the video description. Third, your own recorded music,
  414. 29:31–29:37Mathilla has its own rich musical tradition, and so does every region, if you or someone you know,
  415. 29:37–29:42sings or plays, then a folk tune recorded by yourselves is the most authentic source of all,
  416. 29:42–29:47and it gives your video an identity no AI can. The tune of a folk's song is traditional,
  417. 29:47–29:53but a particular recording or arrangement belongs to its maker. Keep this distinction in mind.
  418. 29:53–29:59The craft of laying music, under a voiceover keep the music low, 20 to 30% of the main voice.
  419. 29:59–30:05Where there is no voiceover opening, close, scene changes the music may rise, let the mood of the
  420. 30:05–30:10music match the mood of the video, brightness for a morning scene, tenderness for a farewell,
  421. 30:10–30:16and at the end let the music sink away slowly fade out, music cut off abruptly jolts the ear.
  422. 30:16–30:21Sound effects are the small sounds, birdsong, the splash of water, the rustle of wind,
  423. 30:21–30:26they make a scene believable, some newer video models generate sound along with the scene native
  424. 30:26–30:32audio, if your tool has this, keep it on, if not take sounds from a free library and add them
  425. 30:32–30:38at the edit. Exercise. Gather background music of two different styles for your video,
  426. 30:38–30:42One AI generated, one from a free library. Play each behind the video in your mind's
  427. 30:42–30:45eye at least and consider which mood fits better.
  428. 30:46–30:49Chapter 11. Editing, turning clips into a video.
  429. 30:51–30:57Now you hold all the ingredients, video clips, voiceover, music, editing is the kitchen where
  430. 30:57–31:02these ingredients become the dish and the good news, editing skill, once learned,
  431. 31:02–31:06Serves in every video, it does not keep changing the way AI tools do.
  432. 31:06–31:10Shoes in editor, free editors exist for both mobile and computer.
  433. 31:10–31:15CapCite is at present the most popular and the simplest, on the computer.
  434. 31:15–31:18DaVinci Resolve is professional grade even in its free form.
  435. 31:18–31:21Shoes either, the structure figure 8 is the same in all.
  436. 31:21–31:25The media area, where you bring and import all your files.
  437. 31:25–31:27The preview, where the video plays as you work.
  438. 31:27–31:30The timeline, the most important of all.
  439. 31:30–31:35The line of time on which the clips are arranged in order, the timeline has several strips tracks,
  440. 31:35–31:42one for video, one for voice, one for music, one for subtitles, stacked one above another,
  441. 31:42–31:45all playing together. Keep the basic order of editing thus.
  442. 31:46–31:501. First lay the voice over on the timeline, this is the spine of the video,
  443. 31:50–31:55the visuals will be arranged upon it. 2. Listening to the voice.
  444. 31:55–31:59Place the video clips in order, let the scenes show what the words are saying,
  445. 31:59–32:04Cut the clips, keep the best portion of each, remove the rest. AI clips are often awkward at the
  446. 32:04–32:08very start and the very end. Trim both edges and the clip cleans up.
  447. 32:09–32:14Pre, add transitions, the manner of passing from one clip to the next, the rule, the fewer,
  448. 32:14–32:19the better, the plane cut is the purest, a light fade at a change of mood.
  449. 32:19–32:21Thinning, twirling, color transitions are the mark of the beginner.
  450. 32:22–32:26Four, lay the music on its track and to bring its level down chapter 10.
  451. 32:26–32:315. Color correction Most editors have a one-click filter or enhance.
  452. 32:31–32:36If your AI clips came from different tools, put the same filter on all, the colors fall
  453. 32:36–32:39into step, and the video feels stitched of one cloth.
  454. 32:39–32:446. Add a title card at the start 3-4 seconds and a closing card at the end the channel's
  455. 32:44–32:47name, a word of thanks.
  456. 32:47–32:48Export settings
  457. 32:48–32:541080p, MP4 format, 30 frames per second, the standard for YouTube.
  458. 32:54–32:57Keep the exported file in the final folder.
  459. 32:57–33:02One suggestion after the video is done, watch it once from beginning to end without stopping
  460. 33:02–33:07as a viewer, wherever your attention drifts, know that a cut is needed there, then show
  461. 33:07–33:11it to someone at home, the reaction of one first viewer teaches more than a hundred
  462. 33:11–33:12critics.
  463. 33:12–33:13Exercise.
  464. 33:13–33:19Join all your clips, voice, and music into your first complete video with title card
  465. 33:19–33:23and closing card, export it, show it to a family member and write down their first
  466. 33:23–33:25reaction.
  467. 33:25–33:27Chapter 12.
  468. 33:27–33:28Subtitles.
  469. 33:28–33:30A video that can be read.
  470. 33:30–33:35Most viewers today watch video without sound, on the bus, in the office, in bed at night,
  471. 33:35–33:40no subtitles, no viewers, and for content in a language like Mathalie the importance
  472. 33:40–33:42of subtitles is doubled.
  473. 33:42–33:46Subtitles in their original language build the habit of reading it, while Hindi or English
  474. 33:46–33:51subtitles bring in viewers who do not know the language at all, that is, your story
  475. 33:51–33:53travels the whole world.
  476. 33:53–33:58The standard format of subtitles is SRT, a plain text file in which three things repeat
  477. 33:58–34:04over and over Figure 9, a serial number, a timeline from which second to which second,
  478. 34:04–34:09and the text, this file can be made and corrected even in an ordinary text editor Notepad.
  479. 34:09–34:15But matching the timing by hand is laborious, and here AI helps again, two ways.
  480. 34:15–34:19First, automatic transcription, the auto captions feature in an editor such as
  481. 34:19–34:25cap cut listens to the video's voice and writes the subtitles itself. Timing included. In Hindi
  482. 34:25–34:31and English this is very accurate. A mathily voice it will usually hear as Hindi and write accordingly,
  483. 34:31–34:36then you correct the text it made, even so, correcting is far faster than writing from scratch,
  484. 34:36–34:42because the timing arrives ready made. Second, from the script, you already have the script
  485. 34:42–34:47written, give a chatbot your script and the video's total duration and say, divide this into SRT
  486. 34:47–34:53format, each subtitle at most two lines, at a comfortable reading pace, load the file in
  487. 34:53–34:58the editor and nudge the timings forward or back. The craft rules of subtitling, at most
  488. 34:58–35:05two lines at a time, roughly 32 to 40 characters per line. Each text stays on screen at least
  489. 35:05–35:10one second, a sentence breaks where the meaning allows I went slash to the market, no, I
  490. 35:10–35:15went to the market together, letters in white with a light dark shadow or strip behind,
  491. 35:15–35:20So they can be read even over a bright scene, place them at the lower middle of the screen,
  492. 35:20–35:23but not so low that the real format cuts them off.
  493. 35:23–35:28On YouTube, subtitles can be given in two ways, burned into the video joint at export from
  494. 35:28–35:33the editor or upload it as a separate SRT file which the viewer can switch on and off.
  495. 35:33–35:38The best method, give the original language subtitles as a separate file, and add Hindi
  496. 35:38–35:41and English as separate SRT files too.
  497. 35:41–35:46A chatbot will help with the translation the duty of checking it remains yours, thus one
  498. 35:46–35:49video reaches the viewers of three languages.
  499. 35:49–35:54Exercise, make the SRT of your video in its own language by either method, then make its
  500. 35:54–35:59Hindi or English translation file, run both with the video and check.
  501. 35:59–36:00Is the timing right?
  502. 36:00–36:02Is any line too long?
  503. 36:02–36:04Chapter 13.
  504. 36:04–36:05Publishing on YouTube.
  505. 36:05–36:10The video was made, now carry it to the viewer, YouTube remains the largest and the most
  506. 36:10–36:16Lasting platform, a video placed here keeps being watched year upon year, while on real-format
  507. 36:16–36:21platforms a video's life is a few days, so make YouTube the main house, let reels and
  508. 36:21–36:23short speed its windows.
  509. 36:23–36:28Making a channel is simple, sign into YouTube with your Google account and create one, choose
  510. 36:28–36:33the channel's name with thought, short, easy to say, and suggestive of the subject.
  511. 36:33–36:38The channel picture logo and banner too can be made with an AI image tool.
  512. 36:38–36:41And up o' time keep the checklist of figure 10 before you.
  513. 36:41–36:42A few points in detail.
  514. 36:42–36:44Title.
  515. 36:44–36:48The main matter in the first three or four words, with a title in your own language,
  516. 36:48–36:52adding Hindi or English in brackets helps the video surface in search, because seekers
  517. 36:52–36:54search in every language.
  518. 36:54–36:55Description.
  519. 36:55–36:58The first two lines are the most valuable.
  520. 36:58–36:59These appear in search results.
  521. 36:59–37:01Write the video's essence here.
  522. 37:01–37:02Key words included.
  523. 37:02–37:07Below them, the chapter list what comes at which minute, the list of sources, and
  524. 37:07–37:13channel introduction. Thumbnail. The viewer sees the thumbnail before the title, the rule,
  525. 37:13–37:18one image, one feeling, at most three words, and words large enough to be read on a small
  526. 37:18–37:23mobile screen make an attractive thumbnail with an AI image tool, but never a misleading
  527. 37:23–37:27one. What is in the thumbnail must be in the video where the viewer feels cheated
  528. 37:27–37:32and trust in the channel is gone. The AI disclosure, YouTube now expects that realistic
  529. 37:32–37:37looking AI generated or AI altered content be declared at upload in answer to the altered
  530. 37:37–37:42content question, this is not mere rule keeping, it is honesty with the viewer, for plainly
  531. 37:42–37:48imaginary styles such as animation the duty usually does not arise, but when in doubt,
  532. 37:48–37:50declaring is always the better course.
  533. 37:50–37:51Language setting
  534. 37:51–37:53Choose the video's language.
  535. 37:53–37:58Mathely is in YouTube's list, as are many others, this helps the video reach those
  536. 37:58–38:02searching in that language, and it strengthens the statistics of the language's content
  537. 38:02–38:08besides. After the upload, what then? Watch the response of the first hours and days, reply to
  538. 38:08–38:13comments, the early conversation waters the channel's roots, and keep regularity.
  539. 38:13–38:19One video a fortnight makes 24 in a year, and this bears more fruit than 100 videos at random.
  540. 38:19–38:23Viewers attach themselves to a program, not to scattered surprises.
  541. 38:23–38:28For reels and shorts, cut the most engaging 30-60 seconds of your main video into the
  542. 38:28–38:349-16 ratio with the editor's reframe or crop feature, and write at the end.
  543. 38:34–38:38Full video on the channel, this is the window that leads new viewers to the house.
  544. 38:38–38:40Exercise.
  545. 38:40–38:43Upload your video, completing all seven points of the checklist.
  546. 38:43–38:46Also cut a shorts version and upload it separately.
  547. 38:46–38:51After one week, compare the figures of the two views, watch time.
  548. 38:51–38:55After 14, special considerations for Mathili content.
  549. 38:55–38:57This chapter is the heart of this book.
  550. 38:57–39:02The tools are universal, but our purpose is particular, a Mathili for Mathila, and readers
  551. 39:02–39:06working in any less served language will find the same principles apply to their own.
  552. 39:06–39:12First, purity of language, AI tools, when writing Mathili usually let the shadow of
  553. 39:12–39:17Hindi fall across it, because they have learned far more Hindi, scripts, translations,
  554. 39:17–39:19Subtitles made by a chatbot.
  555. 39:19–39:21Check every one with your own eyes.
  556. 39:21–39:22The plain rule.
  557. 39:22–39:26The A.I.'s mathily is a draft, not an authority, where in doubt.
  558. 39:26–39:27Trust your ear.
  559. 39:27–39:29Speak the line aloud.
  560. 39:29–39:32Whatever grates on the ear is the shadow of Hindi.
  561. 39:32–39:34Second, the question of script.
  562. 39:34–39:36Mathily is written in Devanagari,
  563. 39:36–39:38and it also has its own ancient script,
  564. 39:38–39:39Turhuta mythilakshara.
  565. 39:39–39:43Keep the video subtitles and text cards in Devanagari.
  566. 39:43–39:45The most people will be able to read them.
  567. 39:45–39:51used Turhuta for beauty and identity, entitled cards, in the logo, in a watermark, thus the
  568. 39:51–39:57script stays before the eye, and curiosity awakens too, remember. AI image tools cannot
  569. 39:57–40:02yet write Devanagari or Turhuta letters correctly, always add written, text yourself at the
  570. 40:02–40:05edit, never have the AI write it.
  571. 40:05–40:10Third, authenticity of the visuals. Tell an AI an Indian village and it will produce
  572. 40:10–40:14a generalized North Indian scene, which is not Mithila, give the prompt Mithila's
  573. 40:14–40:21particular signs, the pond, the banyan, the mango orchard, the patty field, fish, pond,
  574. 40:21–40:25the thatched house, the tall-sea platform in the courtyard, wall ornament in the manner
  575. 40:25–40:31of Arapan. Even then, what comes will be Mathila-like, not Mathila, so wherever possible, blend in
  576. 40:31–40:34real photographs and footage by the method of Chapter 7.
  577. 40:34–40:39The mixture of AI scenes and real scenes gives the most authentic result of all.
  578. 40:39–40:44Fourth, the honor of Madhubani's slash Mathila painting, the AI can be told to generate
  579. 40:44–40:48in Madhubani style, and it will imitate the colors and the line.
  580. 40:48–40:49But pause here and think.
  581. 40:49–40:53Methila painting is a living tradition, the livelihood of thousands of artists rests
  582. 40:53–40:58on it and each of its manners Barney, Kachni, Godna, Gober carries its own lineage.
  583. 40:58–41:03An AI made Madhubani like image lifts the traditions appearance without its labor and its
  584. 41:03–41:04knowledge.
  585. 41:04–41:09My counsel, in your videos show real works by real artists with permission and credit.
  586. 41:09–41:12This honors the artist and strengthens your video at once.
  587. 41:12–41:17If you do use AI ornament in the Madhubani manner, say plainly that it is an AI mediation,
  588. 41:17–41:19not authentic Madhubani.
  589. 41:19–41:245. An inexhaustible store of subjects The field of mathily video is still nearly
  590. 41:24–41:30empty. Whoever makes makes first. Some directions. Children's material songs,
  591. 41:30–41:36tales, letters, the child audiences the fastest growing of all. Festival explainer C. H.
  592. 41:36–41:43Chathai, Samachakpa, Jhursital, Huat, Wai, Hau, recipes, folk tales and the verses of
  593. 41:43–41:48Vidya Patti presented with images, the vocabulary of village and home a visual dictionary of
  594. 41:48–41:52the words now slipping away, introductions to the places of Mathila.
  595. 41:52–41:55Each direction could be a channel in itself.
  596. 41:55–41:59And the last word, the patience of quality, in Methili the audience will be smaller than
  597. 41:59–42:00in Hindi.
  598. 42:00–42:04This is natural, but the loyalty of the Methili viewer is greater.
  599. 42:04–42:08They will comment, they will share, they will return again and again, look not at numbers
  600. 42:08–42:14but at relationships, a hundred devoted viewers are worth more than 10,000 indifferent ones.
  601. 42:14–42:19Exercise, choose one of the six directions above and plan three consecutive videos on
  602. 42:19–42:22its subject plus a one-line summary each.
  603. 42:22–42:27Three, because one video is an experiment, three are a direction.
  604. 42:27–42:28Chapter 15.
  605. 42:28–42:32Copyright, ethics, and responsibility.
  606. 42:32–42:34The powerful tool comes responsibility.
  607. 42:34–42:38This chapter is short, but bring its every line into practice.
  608. 42:38–42:43Copyright, the root principle, what you did not make, you do not use without permission,
  609. 42:43–42:49film songs, portions of others' videos, the text of books, other people's photographs,
  610. 42:49–42:51the rule covers them all.
  611. 42:51–42:56Everyone does it is no argument, YouTube's automatic system content ID catches it, and
  612. 42:56–42:57the channel bears the penalty.
  613. 42:57–43:03When freely licensed material to read the conditions, one says a tribution required, another no
  614. 43:03–43:05commercial use.
  615. 43:05–43:08Know also the question of rights over your own AI made material.
  616. 43:08–43:14In most tools terms, permission for commercial use of generated video and images comes with
  617. 43:14–43:17the paid plan, at a sometimes limited on the free one.
  618. 43:17–43:22For any tool you work with regularly, read its terms of use once, especially the ownership
  619. 43:22–43:24and commercial use sections.
  620. 43:24–43:30The dignity of persons, three iron prohibitions, no imitation of anyone's face or voice without
  621. 43:30–43:34written consent living or departed, both.
  622. 43:34–43:37Never show a person in a scene that lowers their honor, nor put in their mouth words
  623. 43:37–43:43they never spoke the deep ache, and take special care with the images and voices of children,
  624. 43:43–43:46before posting a video even of your own child.
  625. 43:46–43:48Consider that it is becoming public.
  626. 43:48–43:49Toward truth.
  627. 43:49–43:53A eye can make scenes that look real, and therefore it can deceive.
  628. 43:53–43:56The rule, let fiction be called fiction and news be news.
  629. 43:56–44:02When you make a video on a historical or cultural subject, check the facts against your own trusted
  630. 44:02–44:03sources.
  631. 44:03–44:07An AI chatbot speaks its errors with full confidence this is called hallucination.
  632. 44:07–44:11Wherever an AI scene looks real, declare it Chapter 13.
  633. 44:11–44:15A viewer's trust takes years to earn and one video to lose.
  634. 44:15–44:16And toward yourself.
  635. 44:16–44:19Let AI's convenience never turn into laziness.
  636. 44:19–44:21Take work from AI, not thought.
  637. 44:21–44:26The day you hand the script, the choices, and the final judgment over to AI, that day
  638. 44:26–44:30the video ceases to be yours, and viewers can smell the difference.
  639. 44:30–44:34Keep the tool in your hand, let not the hand become the tool.
  640. 44:34–44:35Exercise.
  641. 44:35–44:40Run a responsibility check on the first video you made, the music's license, permission
  642. 44:40–44:45for the image sources, the facts verified, the AI disclosure, keep the four answers
  643. 44:45–44:46in writing.
  644. 44:46–44:52Repeat this exercise with every video, within days it will become second nature.
  645. 44:52–44:56Chapter 16 The Practice Project, a first video from start
  646. 44:56–44:58to finish.
  647. 44:58–45:03Now all the threads in one place, this chapter is the model of one complete project, a 92nd
  648. 45:03–45:08video called CHathai, the Great Festival of Sun Worship, read it, then repeat the
  649. 45:08–45:11same sequence on a subject of your own.
  650. 45:11–45:15Day 1 Planning Chapters 3, 4, Purpose, to explain simply
  651. 45:15–45:21the spirit and the observance of Sirei Chathai, audience, Mathal families living away from
  652. 45:21–45:27home, the new generation especially, duration, 90 seconds, that is, 9 or 10 clips, a two-column
  653. 45:27–45:32script drafted by a chat pot, corrected by hand in three places, local detail added
  654. 45:32–45:37to the description of Nahikai, the feeling of the final line deepened, 10 scenes on
  655. 45:37–45:42the storyboard, the river-gat at dawn established, the courtyard where Thekua is being made,
  656. 45:42–45:47the supendora of the offering, the evening argia, the morning argia, the peran, and the
  657. 45:47–45:49closing card resolve.
  658. 45:49–45:55Day 2, Voice Chapter 8, the voiceover recorded in my own voice, at night, in a quiet room,
  659. 45:55–46:01paragraph by paragraph, polished with the editor's enhanced voice, total 82 seconds,
  660. 46:01–46:03which means the scene plan sits right.
  661. 46:03–46:09These three four, visuals chapters 5-7, a six-part prompt built for every scene thought out and
  662. 46:09–46:11mathily, translated into English.
  663. 46:11–46:17Methylas signs in each, a river ghat in north Bihar, banana stems, bamboo kania, fruit in
  664. 46:17–46:21the soup and aura, three clips came out well at the first attempt, four clips took two
  665. 46:21–46:24or three improvement cycles.
  666. 46:24–46:28For two scenes the tekua and the soup the image of video method was used, first the
  667. 46:28–46:31image perfected, then subtle motion given.
  668. 46:31–46:36And in one scene a real photograph of my own was used, the crowd at the gate, which no
  669. 46:36–46:38AI could ever have made.
  670. 46:38–46:41This mixture is the essence of Chapter 14.
  671. 46:41–46:42Day 5.
  672. 46:42–46:45Music and Assembly Chapters 10-11.
  673. 46:45–46:46AI Music.
  674. 46:46–46:52Slow, devotional, flute and emridang, 90 seconds, the license checked, the editing order,
  675. 46:52–46:57voice first, then clips, plain cuts, a fade in two places, one and the same light warm
  676. 46:57–47:02filter on every clip, on the title card CH-athai in Thirhuda and the full title in Devanagari.
  677. 47:03–47:09Day 6, subtitles and publication chapters 12-13, the Mathili SRT from the script,
  678. 47:09–47:14then Hindi and English translation files translated by Chatbot, checked by hand,
  679. 47:14–47:18the thumbnail, the scene of the evening argh, three words,
  680. 47:18–47:25hui chathai, ad upload, language Mathili, the AI disclosure switched on, music credit and sources
  681. 47:25–47:29in the description. Alongside, a 35-second shorts version.
  682. 47:30–47:32Day 7 Review Chapter 15,
  683. 47:32–47:36the responsibility check on all four points shown to the family,
  684. 47:36–47:41the suggestion came that the morning argia scene feels short, noted down for the next video.
  685. 47:42–47:46See, seven days, an hour or two a day, and from nothing to a published video,
  686. 47:46–47:51the first project will take longer than this, let it, by the third or fourth video the sequence
  687. 47:51–47:58becomes habit and the time falls by half. A last word. This book ends here, but the learning begins
  688. 47:58–48:03here. The tools will change. The models will change. The screens will change. But the framework you
  689. 48:03–48:10have learned plan, prompt, improvement cycle, assembly, responsibility will stand. Now go and
  690. 48:10–48:14tell the story of your Mathila in your own voice, in your own language.
  691. 48:14–48:16The Pentax A.
  692. 48:16–48:17Blossary.
  693. 48:17–48:23AI artificial intelligence, the capacity of a computer to learn and understand in a human-like
  694. 48:23–48:24way.
  695. 48:24–48:25Generative AI.
  696. 48:25–48:30AI that creates new material, text, images, video, voice.
  697. 48:30–48:31Prompt.
  698. 48:31–48:34The written instruction given to an AI.
  699. 48:34–48:35Negative prompt.
  700. 48:35–48:38The list of what is not wanted.
  701. 48:38–48:39Text to video.
  702. 48:39–48:42The method of making video from a written description.
  703. 48:42–48:43Image of video.
  704. 48:43–48:46Method of setting a still image in motion.
  705. 48:46–48:51TTS text-to-speech, the method of producing a spoken voice from written text.
  706. 48:51–48:55Avatar, a digital speaker who delivers text on screen.
  707. 48:55–48:59Lip sync, the matching of voice and lip movement.
  708. 48:59–49:03Voice cloning, making a digital replica of a voice.
  709. 49:03–49:08Credit, the usage currency of AI tools, each generation deducts some.
  710. 49:08–49:12Model, a particular version or engine of an AI.
  711. 49:12–49:16Clip, a short segment of video usually 5-20 seconds.
  712. 49:16–49:23Aspect Ratio, the ratio of a screen's width to its height, 16 colon 9 wide, 9-16 upright.
  713. 49:23–49:30Resolution, the fineness of the image, 1080p standard, 4K extra fine.
  714. 49:30–49:33Watermark, an identifying mark printed on content.
  715. 49:33–49:36Storyboard, the scene by scene sketch plan.
  716. 49:36–49:40B-roll, supporting footage beyond the main speaker.
  717. 49:40–49:43Timeline, the line of time in an editor on which clips are arranged.
  718. 49:44–49:51Track, a layer of the timeline, video, voice, music, and so on. Cut, trimming a clip, or passing
  719. 49:51–49:57directly from one clip to the next. Transition, the manner of joining two clips fade and the rest.
  720. 49:58–50:03Fade and slash out, slow emerging slash, slow vanishing. Render slash export,
  721. 50:03–50:10Turing the edited video into the final file. SRT, the standard file format of subtitles.
  722. 50:11–50:14Burned in, subtitles fixed permanently into the video.
  723. 50:15–50:22Thumbnail, the face image of a video. Hallucination, an AI's confident error. Content ID,
  724. 50:22–50:28YouTube's copyright checking system. Deepfake, a deceptive imitation of someone's face or voice.
  725. 50:28–50:31Monetization, earning from videos.
  726. 50:31–50:36Appendix B, A Collection of Prompts
  727. 50:36–50:40The lower ten ready prompt frames, fill the brackets with your own details and shape the
  728. 50:40–50:43final prompt by the method of Chapter 5.
  729. 50:43–50:441.
  730. 50:44–50:45Village Dawn
  731. 50:45–50:50A village in North Bihar, Dawn, light mist over the fields, thatched houses, a bayon tree
  732. 50:50–50:56in the distance, your detail, camera panning slowly to the right, soft golden light, realistic
  733. 50:56–50:57documentary style.
  734. 50:57–50:582.
  735. 50:58–51:04scene, a pond covered with lotus and water lilies, seasoned on the bank tree slash cat.
  736. 51:04–51:10Gentle ripples on the water, subtle movement of bird slash fish, camera moving forward very slowly,
  737. 51:10–51:16time of day light, semantic. Free, courtyard scene, the courtyard of a mythila home,
  738. 51:16–51:23a Tulsi platform, action, say, food cooking on the hearth, folk ornament on the wall, warm intimate
  739. 51:23–51:25light, close-up, realistic.
  740. 51:25–51:264.
  741. 51:26–51:27Festival scene.
  742. 51:27–51:33Preparations for name of festival, central object, say, soup and dora, lamps, courtyard
  743. 51:33–51:39or get, close-ups of busy hands, festive joy, warm light, documentary style.
  744. 51:39–51:405.
  745. 51:40–51:41Nature scene.
  746. 51:41–51:46Green paddy fields waving in the wind, time, the far horizon, bird and flight, drone
  747. 51:46–51:50view rising slowly, natural light, ultra-fine detail.
  748. 51:50–51:526.
  749. 51:52–51:58studied dish slash craft, an extreme close-up of object, background, light steam or sheen,
  750. 51:58–52:04camera circling the object slowly, soft studio-like light, appetizing slash appealing style.
  751. 52:04–52:057.
  752. 52:05–52:07Doreen illustration animation.
  753. 52:07–52:09Character description doing action.
  754. 52:09–52:15Place, hand-drawn animation style, soft watercolor, like a children's book illustration,
  755. 52:15–52:16slow pleasant motion.
  756. 52:16–52:178.
  757. 52:17–52:24imagination, an imagined scene of period methila, action, colors like an old palm leaf painting,
  758. 52:24–52:30slow-sulme camera, museum documentary style, the AI disclosure is obligatory here. This is imagination.
  759. 52:31–52:379. Title card background, an abstract ornamental pattern drifting slowly, color scheme,
  760. 52:37–52:43say, from million red and yellow, no letters, no figures, come even motion, 10 second loop,
  761. 52:43–52:45Add the lettering yourself at the edit.
  762. 52:46–52:4910. Motion instruction for image video.
  763. 52:49–52:54In this image, gentle natural motion in element, say, water slash leaves slash cloth,
  764. 52:54–52:59camera static slash moving forward very slowly, nothing else changes.
  765. 52:59–53:02The image's original color and line remain intact.
  766. 53:03–53:07The Pendex C, a guide to the tools as of 2026.
  767. 53:08–53:12This list reflects the state of things at the time of writing 2026,
  768. 53:12–53:14Names and features keep changing.
  769. 53:14–53:16Learn the categories.
  770. 53:16–53:19Do not memorize the names before adopting any tool.
  771. 53:19–53:20Check three things.
  772. 53:20–53:22What the free plan gives,
  773. 53:22–53:24whether commercial use of the output is permitted,
  774. 53:24–53:26and where the watermark stands.
  775. 53:26–53:28Text-to-video slash image video,
  776. 53:28–53:30Google's Vivo,
  777. 53:30–53:31OpenAI's Sora,
  778. 53:31–53:32Kling,
  779. 53:32–53:33Runway,
  780. 53:33–53:34Pixverse,
  781. 53:34–53:35Seedence,
  782. 53:35–53:36Haleuo,
  783. 53:36–53:37Luma,
  784. 53:37–53:38Pika,
  785. 53:38–53:39Realistic Scenes,
  786. 53:39–53:40Cinematic Clips,
  787. 53:40–53:41Avatar Video,
  788. 53:41–53:45Synthesia, Digital Speakers, Lip Sync, Many Languages.
  789. 53:45–53:51Text-to-speech and voice cloning, 11 labs and others, services centered on Indian languages
  790. 53:51–53:56are also growing, a mathily option will appear there first, keep watching.
  791. 53:56–54:02AI Music, Suno, Udio and others, music and song from a description or from lyrics.
  792. 54:02–54:08Image generation, included in nearly every video tool, separately, Mijrani, and the
  793. 54:08–54:11The image features attach to the chatbots.
  794. 54:11–54:18Editing, cap cut, mobile and computer, simple auto-cactions included, DaVinci Resolve, computer,
  795. 54:18–54:24free professional grade, canva, thumbnails, banners, simple video.
  796. 54:24–54:30All in one plant video, tools of the in-video kind, which, given a topic, join script,
  797. 54:30–54:34visuals and voice on their own, good for quick work, but your own step is faint in
  798. 54:34–54:39them, so for learning, the step-by-step method of this book is better.
  799. 54:39–54:46Chatpot script, translation, SRT, brainstorming, Claude, chat GPT, Gemini.
  800. 54:46–54:50And last of all, this book too will age with time, but you will not.
  801. 54:50–54:54The teacher self-temper this book has built in you will apply to every new tool, when
  802. 54:54–54:59you seal a new tool, ask three questions, where is the prompt written, what is in the
  803. 54:59–55:04settings, where and how does the result arrive, and within five minutes you will
  804. 55:04–55:04be working it.

Plain text

Making videos with AI. Teach yourself series, written by Gage and Rathakar. Preface. This book is for everyone who wants to say something through video but has no camera, no studio, and no training in editing. Artificial intelligence AI has now made all three possible with an ordinary mobile phone or computer. What once required equipment worth hundreds of thousands and a team of dozens can now be done with a written instruction, called a prompt. but let one thing be clear at the very outset. AI is no magic wand, it is a tool, like an axe, like a pen, the axe cuts the wood, but which tree to fell and what house to build, the carpenter decides, in the same way, AI will make the visuals and generate the voice, but what to say, why to say it, and for whom, the answers to these three questions remain with you, this book will teach you to handle the tool, the craftsmanship will remain your own. This book is written in the teacher-self style, that is, no teacher or training institute is needed. Each chapter is one step, begin with the first chapter, move forward in order, and be sure to do the exercise give at the end of every chapter. Reading and doing must walk together, only then does learning happen. One more thing, all the illustrations in this book are original explanatory figures, not copies of any company's actual screens. There There are two reasons. First, the screens of AI tools change every few months, so a real screenshot would be outdated before the book left the press. Second, once you understand the principle, any new screen will feel familiar. If you memorize a screen, every change will leave you stranded. The figures in this book show the structure that applies to nearly every tool alike. Where the prompt box sits, what the settings contain, what a timeline looks like. Once you grasp this framework, you open any new tool and work it out on your own, and that is the true meaning of teach yourself. The absence of technical books in Mathaly has long been a sore point, and this book was first written in Mathaly. This English edition carries the same conviction, whatever you learn here, use it to put your own language, your own region, and your own stories on screen, the tales, psalms, paintings, history and festivals of Mithila, and of every homeland like it, are are all waiting for their videos. CHAPTER 1 WHAT IS AI VIDEO? Artificial Intelligence AI is the capacity of a computer to do human-like work, understanding language, recognizing images and now, creating images and video, when we write to an AI, show two children playing by a pond, and it produces a moving picture of that very scene on its own. This is called Generative AI. To generate means to bring forth. This AI does not fetch a video from somewhere, it composes a new video that never existed anywhere before. How is this possible? Understand it in simple terms, these AI models have been shown crores of images and videos, having seen so much, they have learned what an egret looks like, how water ripples, what color the morning light takes, and how a walking person's feet rise and fall, when you write an egret beside a pond at dawn, the model builds such a scene from its learned knowledge, just as a painter, who has seen thousands of ponds can now draw a pond from memory without one before their eyes. There are chiefly four routes to making video with AI. First, text to video, you only write, and the AI creates the entire scene. This is the most astonishing route, but it offers somewhat less control, the scene that arrives may not match your imagination exactly. Second, image a video, you supply one still image, a photograph or an AI generated picture, and the AI sets it in motion. Here the control is greater because you yourself chose the opening frame. Third, avatar video. A digital speaker avatar sits on screen and speaks your written text, like a news anchor, for lessons, announcements and lectures. This is extremely useful. Fourth, AI assisted editing. Here the footage is your own recording, but the cutting and joining, the subtitles, the background removal. All this the AI does. This book covers all four routes, because in practice a good video is usually a blend of them. Now understand what AI cannot do. AI does not yet produce long videos in one go, it typically makes short clips of 5 to 20 seconds, which are joined to build a longer video. AI may seem sometimes contain errors, a hand grows too many or too few fingers, written letters come out garbled, a character's face changes between one clip and the next, and the biggest point of all. AI knows nothing of Mathila, of Mathili, or of what lies in your heart. It will give only as much as you know how to ask. That is why half of the spoke is devoted to how to ask, that is, planning and prompts. Exercise. On paper, write down three subjects for videos you would like to make, for each, write one line, who will watch this video and what will they gain from watching it. Chapter 2. Getting Ready. What Do You Need? No costly machinery is needed to make video with AI, all that is required is this. First, a device, a smartphone is sufficient, a computer or laptop adds convenience, writing prompts, managing files and editing are easier on a large screen, no specially powerful computer is needed because the video is made not on your device but on the company servers, your device merely sends the instruction and downloads the finished video. Second, the internet, But the better the speed, the shorter the weight, video files are large, so downloading will consume data. Budget a few gigabytes a month. Third, an email account, nearly every AI tool requires an account, and most allow direct sign-in with a Google email Gmail, one suggestion. Keep a separate email for this work, so that the newsletters of AI tools do not flood your main inbox. Fourth, an understanding of credits. Most AI video tools run on a credit system. A credit is a kind of coupon, making one video deduct some credits, a free account receives a small allowance, daily and some tools monthly and others, once only in a few. Paid plans bring more credits, higher quality such as 1080p or 4k, videos without watermarks and longer clips. Adapt a practical policy here, which I call cheap first, then deep, while learning a new tool, run small experiments on the free credits, learn to write prompts, learn the tools temperament when it feels time for serious work, say, running a YouTube channel, then take a paid plan on one tool. Buying plans on every tool is wasteful, one or two suffice. 5. A system for files. It sounds a small matter, but the experience know how work drowns without order. Make a separate folder for each video project on your computer or phone. Inside it keep four subfolders, script scripts and prompts, clips raw AI made video, voice Peace music voiceover and music and final the finished video give every file a meaningful name. A file called clip a one-pond on will still be recognizable six months later. Video three final new two will not. And finally, patience, AI tools are sometimes busy, sometimes give strange results, sometimes fail entirely to understand your perfectly good prompt. All this is natural. Those who persist through three or four attempts learn, those who quit at the first odd result are left behind. Exercise, create the folder system described above on your device, make a new email account if you wish, and open the website of any one AI video tool simply to see what the free plan offers, do not make anything yet, only look. Chapter 3, Planning the video, from idea to script The biggest mistake beginners make is to open the tool straight away and start pressing a video made without a plan looks exactly like a house built without a drawing. That is why this chapter comes before the tools. Before every video, write the answers to three questions. 1. What is the purpose? To teach say how Matubati painting is made, to tell a story, to show scenes of a village or to an else. 1. Video, 1 purpose. Hold to this rule. A video that contains everything contains nothing. 2. Who is the audience? Children, students, expatriates far from home, curious outsiders, or the general viewer, the audience decides how simple the language should be, how long the video should run, and what the visuals should look like, bright colors and quick movement for children, stillness and gravity for adults. Free, How Long, In The Beginning, Make Videos Of 30 Seconds To 2 Minutes. A short video is easier to make, easier to fix, and better liked by today's viewer, remember, AI produces clips of 5 to 10 seconds, so a 1 minute video means 6 to 12 clips. Now the script, a script is nothing complicated or literary, it is simply a two column table, in the left column, what is seen the visual, in the right column, what is heard the voice or subtitle, for a 30 second video, 5 or 6 rows are enough. Here AI can help in a second role, as script assistant, give your subject to a chatbot such as Claude, ChatGPT, or Gemini and AskIt to draft the script, but the AskInG needs skill, merely saying write a video script on Madhubani will fetch a generic, lifeless text, ask like this. I am making a 60 second YouTube video on Madhubani painting, audience, young Indian viewers who have heard the name but know little more, voice, simple and warm, I need a script in two columns, on the left, a description of each visual which I will give to an AI video tool. On the right, the voice of her text. 6 scenes, 10 seconds each. Notice, subject, duration, audience, language, format, and use are all stated. The clearer the ask, the better the result. This same principle will apply later to video prompts. Do not accept the chatbot's script with your eyes closed, read it, test it against your own knowledge. Are the facts right? Does the language sound like your own? any line that falls flat, or have it rewritten make the third scene more tender, shorten the last line, the script is your signature, the AI is only a scribe. Exercise. Take one of the three subjects you chose in chapter 1, write the answers to the three questions purpose, audience, duration. Then, using the detailed style of asking shown above, have a chatbot draft a two-column script and make at least two corrections to it with your own hand. Chapter 4. The Storyboard. A Map of Scenes. The script is a map of words, the storyboard is a map of scenes, a storyboard lays out one sketch, a rough drawing or description, for every scene of the video in order. In filmmaking this method is a century old, and in the AI age its importance has not shrunk but grown. Why? Because AI needs a separate, clear instruction for every clip, and the storyboard is precisely that list of instructions. Do not worry, no drawing skill is required. Round faces, stick figures, and arrows are enough, if you would rather not sketch at all, write three lines for each scene, what is seen, what the camera does, and how many seconds. Learn a little of the camera's language, for these very words will serve you later in prompts. Wide shot, the whole scene from a distance, for establishing the place, like a full view of the village. Close up, from near, for feeling and fine detail, like the steam over a cup of teal on the hearth. Tracking shot, the camera moves along with the subject, like following a child on the way to school. Drone view, from above as a bird sees, like a sweeping view of pond and fields. Dallied in slash out, the camera slowly draws near or pulls away, for emotional weight. Static shot, the camera stays in one place, for calm, settled scenes. Keep one simple formula for the order of scenes, establish, develop, resolve. The The first scene tells where we are established, usually a wide shot. The middle scenes bring the subject close develop, close ups, tracking, the last scene gathers it all, a feeling, a message, or title card resolve. Vigor 2 shows the storyboard of a 30-second video called My Village. See how six scenes travel from dawn to dusk, morning mist establish, tea, school and pond develop, lamps and title resolve. Beside each scene its duration and camera note are written, each of these scenes will, further on, become one prompt. While making the storyboard, attend to continuity, if the first scene is morning, the second must not suddenly be night, if a character wears a red kurta, the red kurta must appear in every scene, and this must be written into every prompt, because AI does not remember the previous clip, write the characters' description on a separate sheet, a character card, and pasted word for word into every prompt. The small trick is the simplest way to keep the clips consistent. Exercise. Turn your script into a 16 storyboard. For each scene, write the visual description, the camera, the duration. If there is a character, make a character card dress, age, appearance, in three lines. Chapter 5. The Art of Prompt Writing. A prompt is the written instruction you give to the AI. It is the most valuable skill of the AI age, and the happy news is that it demands no technical knowledge, only clarity of language and language is our home ground. See the difference between a poor prompt and a good one. Poor village scene, this tells that AI nothing, a village of which country, which season, day or night, what is happening, it will invent something from its own mind, most likely some European or placeless village, now the good one, a village in North India, Early morning, light missed over the fields, thatched and tiled houses. A banyan tree in the distance, camera panning slowly to the right, soft gold and light, realistic documentary style. Now the AI holds the complete picture. Keep in mind the six-part formula shown in Figure 3. Subdecked, who or what is central, two egrets. Action, what are they doing? Catching fish. Scene. Where and when, upon full of lotus, at dawn. Camera. What does the camera do? Slowly moving in. Light. What kind of light? Golden morning light. Style. What look? Documentary, realistic. Let every prompt carry all six parts and keep roughly this order. Subject and action first, style last, some useful words of style, realistic, cinematic, animation, like a watercolor painting, like old film, documentary. The question of language, most AI video tools understand English prompts best, some also follow Hindi and other languages, the practical path is this, they can script in your own language and shape the final prompt in English, here a chatbot helps, give it your scene description in your language and say, turn this into an English video prompt with 6-part subject, action, scene, camera, light, style, thus the imagination stays yours, only the translation is mechanical. There is also the practice of the negative prompt, where you state what you do not want, for example, no blur, no distorted hands, no text on screen, no watermark, some tools give it a separate box, in others it is added to the main prompt. Now the most important principle of all, the improvements cycle figure 4, the first prompt rarely yields the video of your wishes, and this is no failure, it is the method itself. Look at the result, name the fault, change just that much in the prompt, and generate again. Did the egret come out too large? At a small egret, does the scene look garish? Change to soft, gentle light, usually within 3-5 cycles a usable clip arrives. Change only one or two things per cycle. Change everything at once and you will never know which change did the work. Keep saving your good prompts in one file, The you of six months hence will thank the you of today. Exercise, for the first scene of your storyboard, write one complete prompt using the six-part formula in your own language, then make its English form yourself or through a chatbot. Keep both in your script folder. Chapter 6. Text-to-video tools, the first clip. The moment has come to take the tool in hand. In this chapter we understand the common structure of text-to-video tools and make the first clip. First, an introduction to the tools, this field changes at great speed. Every few months a new model arrives and the old ones grow stronger. At the time of writing 2026 the leading names are, Google's BO, OpenAI's Sora, Kling, Oneway, Pixverse, Seedance, Haleuo, Luma, and Pika. Each has its own temperament, one excels at realistic scenes, another at stylized or artistic ones, another is faster and cheaper, the names will keep changing, but the structure, described below, remains nearly the same in all of them, so learn the structure, not the names. Look at figure 5, in almost every tool you will find these 5 things. 1. The prompt box, a large empty field where you write your instruction, this is the heart of the tool. 2. Model selection. A single company offers several models, new and powerful costly, Older or fast cheap, choose the cheap model while learning. The good model for final, publishable clips. 3. Aspect Ratio 16 colon 9 for YouTube and television wide, 916 for Reels, Shorts, and Status Upright, 1 colon 1 square, the side at the outset where the video will go, because changing the ratio later crops the scene. 4. Duration Usually options of 5, 8 or 10 seconds. A shorter clip costs fewer credits and carries fewer errors. 5. The generate button and the results area, press the button, wait from a few seconds to a few minutes and the clip appears in the results area. Download it from there. Now the method for the first clip, step by step. 1. Open the tools website and create an account with your email. 2. Check the free credits, how many, and when they renew. 3. Use the cheap slash fast model, 16 colon 9 ratio and the shortest duration. 4. Pace the prompt you built in chapter 5. 5. Press generate and wait. 6. Watch the finished clip in full, not once, but two or three times. Watch the hands, the faces, any lettering, the way things move. 7. Download it into your clip's folder under a meaningful name. Even if the clip is imperfect, it will serve for comparison. 8. the improvement cycle. Name the fault, refine the prompt, generate again. A few practical tricks. Generate two or three clips from the same prompt, AI gives a somewhat different result each time, and you pick the best. A clip whose main subject is right but whose edges carry faults can often be saved by cropping in the edit, do not discard it at once and set yourself a daily usage limit before the credits run dry or weeks credits will vanish in one enthusiastic evening. Exercise. Make the clip for the first scene of your storyboard, running at least three improvement cycles. Save the final clip together with the prompt that produced it. Chapter 7. From Enage to Video, The Road of Greater Control. Text to video has won in convenience. You have no hold over the opening scene. Whatever the AI makes, it makes. The remedy is image of video. First, prepare a still image that is exactly to your mind, then tell the AI to set this image in motion. Hear the look of the scene, the character's face, the clothing, all are fixed in advance. The AI only adds the movement. Where will the image come from? Free sources. First, your own photographs, your village, your festivals, nature, art, animating a photograph you took yourself is the most authentic or out. Remember, the photograph must be your own or used with the owner's permission, and before animating a photograph of a living person, be sure to take their consent. This is both courtesy and part of the ethics, described in chapter 15. Second. AI-generated images. There are separate tools for image generation and most video tools include an image-making feature. The image prompt follows the same six-part formula, only in place of camera movement, describe the composition. Images are cheap and quick to make. so run your improvement cycle on the image first. Generating 10 images and choosing the best is far cheaper than generating 10 videos. Third, scanning your own artwork. A work of Mathila painting if it is your own or you hold the rights can be scanned and given gentle motion. A fish stirring, the line ornament shimmering, exercise great restraint here, the dignity of traditional art lies in its stillness, so keep the motion extremely slight, right gentle, slow, subtle motion in the prompt. The method is simple. To use the image video option in the tool, upload your image and write the motion instructional on side. This prompt now carries less seen description and more motion description, what moves, in which direction, how fast, and what the camera does. For example, gentle ripples on the water, the egrets wings moving slowly, camera moving forward very slowly, nothing else changes, that last phrase, nothing else changes, matters, without it the AI will sometimes transform the whole scene. This is also the best remedy for the problem of character consistency. If the same character appears again and again in your video, first create or choose one excellent image of that character. And for every scene animate that same image with different motion instructions, some advanced tools offer a feature called character reference or consistent character where the character's image, given once, appears in every clip. Look for this feature in your tool. Exercise. Take any image or own photograph or AI generated and make three different motion versions of it. In one, only the camera moves. In the second, only some element of the scene moves. In the third, both, compare the three, which feels most natural. Chapter 8 Voice, Voice over and text to speech. The soul of a video lives not in the visuals but in the voice. A viewer will forgive a blurry scene, but will close the video at a bad voice, so read this chapter with care, and for speakers of mathily and other less-served languages, it holds some special advice. There are two ways to add voice, record your own, or have AI generate a text-to-speech, TTS for short. First, your own voice, because for mathily and languages like it, this remains the best route, the pure pronunciation, the natural cadence, the rise and fall of feeling. No machine yet renders these as well as a native speaker and no studio is needed. A smartphone microphone today is quite good enough, follow a few rules, record in a quiet room fan off. Windows shut. Night or early morning is best. Hold the phone about a hand span from your mouth, speak standing or sitting upright. The voice stays open. Keep the script before you, but speak as if telling, not reading, and record paragraph by paragraph rather than all in one take, when you slip, you re-speak only that much. AI can polish a recorded voice. Many tools, usually named enhanced voice, or studio sound in editing apps strip the background, noise and give the voice a studio finish, use it without fail, the difference between a plain recording and an enhanced one will astonish you. Now text to speech figure six, here you type the text, choose the language and the speaker female or male, young or mature, adjust pace and pitch, and download an MP3 file for Hindi, English, and other major languages this facility is very mature, for mathily the situation is improving. Some tools have begun to offer a mathily voice, and tools built for Indian languages are the most likely place to find one. Search in your tool. If mathily appears, first test it with a short passage to hear how pure the pronunciation is. If no mathily voice is available, two remedies, The first and best. Your own voice by the method above. The second. Making do with the hindi voice. Write the text phonetically, listen and adjust the spelling until it sounds right. Keep the pace a little slow. Use short sentences. The result will not carry a fully mathal cadence, but it will serve. Remember, this is a compromise, not an ideal. Wherever feeling at purity matter poetry, stories, children's material, give your own voice. A word on a newer facility, voice-cloning, some tools, from a few minutes of your recording, build a digital replica of your voice, which will then read any text in your own tones, for content in a less-served language this is attractive, teach the tool your voice wants, and the voice-overs of many videos can be made, but two iron rules, clone only your own voice, imitating anyone else's voice without written permission is absolutely forbidden, and where a cloned voice is used in a video, disclosing it is good practice. The joining of voice and visuals will happen at the edit Chapter 11, so keep the voice file separately in the voice music folder, make the voiceover first and the video clips after. This order is wise, because hearing the length of the voice tells you how many seconds each scene needs. Exercise. Make the voiceover of your script both ways, once recorded in your own voice, once through a TTS tool, listen to both with your eyes closed, which sounds more like you. Why? Chapter 9. Avatar Videos, The Digital Speaker. Imagine, every fortnight you must make an announcement video for a journal's new issue, or 50 lectures for a course. Camera, lighting, dress, recording, every single time. Impossible. The remedy is the Avatar video. A digital human sits on screen and speaks your written text, lips moving, eyes blinking, hands gesturing, like a news anchor. The well-known tools of this class are Hei-jen, Synthesia, and others, and the class itself is growing fast. The structure is nearly the same in all figure 7. 1. Choose the avatar. Tools carry hundreds of ready-made avatars, of different ages, dress, and bearing. Some tools also let you build your own avatar from a photograph or a short video. That is, you remain on screen without recording each time, if you make your own avatar. the same consent rule given for voice cloning, only your own likeness, never another's. 2. Give the script, it can take two forms, written text which the tool will speak through TTS, or your own recorded audio file which the avatar will lip sync, for mathily the second road is usually better, upload your own mathily voiceover, and the avatar speaks it, the lip sync you get is surprisingly good. 3. Choose the background and layout, a library, an office, a plain color or an image of your own. Choose the aspect ratio 16 colon 9 or 916, and generate. The beauty of the avatar video lies in its practicality, not in spectacle, news-style presentation, journal announcements, lesson explanations, introductions of an institution. Information videos, in all these it is excellent for the feeling-laden delivery of story and poetry, a human is still better. One tip on presentation. Do not keep the avatar on screen for the whole video without relief. In between, show related scenes, images or text cards called B-roll while the avatar's voice runs beneath. The video comes alive. This weaving happens at the edit, keep the avatar clip and the B-roll clips as separate files. And yes, when the avatar in a video looks human, the viewer has a right to know it is a digital speaker. Write one line in the description. stays intact and trust as a channel's real capital. Exercise, on the free plan of any avatar tool, make a 30-second introduction video, subject, an introduction to my village or an introduction to a favorite book, try both methods, type text and uploaded voice. Chapter 10 Music and sound effects. Visuals for the eye, voice for the ear and music, for the heart, the same scene feels lifeless without music and comes alive with the right score, but with music comes the greatest danger of all. Copyright, put someone's song in your video without permission and YouTube can block the video, others can claim its earnings, and the channel can be penalized, so rule one, never a famous film's song or commercial recording, unless you hold written permission. Then where will the music come from? Free lawful sources. First, AI generated music, there are now tools that compose music from a written description. Write slow, tender, flute-led, 60 seconds and the music is ready. Some tools even build a full song, voice included, from your lyrics. Music made this way for your own video is generally safe to use, but read each tool's license once, especially whether commercial use including YouTube monetization is permitted. Second, copyright-free music libraries. YouTube's own audio library inside YouTube studio is free and safe. Beyond it, many websites offer freely licensed music, some entirely free, some on the condition of attribution. If attribution is required, do not forget to write the musician's name in the video description. Third, your own recorded music, Mathilla has its own rich musical tradition, and so does every region, if you or someone you know, sings or plays, then a folk tune recorded by yourselves is the most authentic source of all, and it gives your video an identity no AI can. The tune of a folk's song is traditional, but a particular recording or arrangement belongs to its maker. Keep this distinction in mind. The craft of laying music, under a voiceover keep the music low, 20 to 30% of the main voice. Where there is no voiceover opening, close, scene changes the music may rise, let the mood of the music match the mood of the video, brightness for a morning scene, tenderness for a farewell, and at the end let the music sink away slowly fade out, music cut off abruptly jolts the ear. Sound effects are the small sounds, birdsong, the splash of water, the rustle of wind, they make a scene believable, some newer video models generate sound along with the scene native audio, if your tool has this, keep it on, if not take sounds from a free library and add them at the edit. Exercise. Gather background music of two different styles for your video, One AI generated, one from a free library. Play each behind the video in your mind's eye at least and consider which mood fits better. Chapter 11. Editing, turning clips into a video. Now you hold all the ingredients, video clips, voiceover, music, editing is the kitchen where these ingredients become the dish and the good news, editing skill, once learned, Serves in every video, it does not keep changing the way AI tools do. Shoes in editor, free editors exist for both mobile and computer. CapCite is at present the most popular and the simplest, on the computer. DaVinci Resolve is professional grade even in its free form. Shoes either, the structure figure 8 is the same in all. The media area, where you bring and import all your files. The preview, where the video plays as you work. The timeline, the most important of all. The line of time on which the clips are arranged in order, the timeline has several strips tracks, one for video, one for voice, one for music, one for subtitles, stacked one above another, all playing together. Keep the basic order of editing thus. 1. First lay the voice over on the timeline, this is the spine of the video, the visuals will be arranged upon it. 2. Listening to the voice. Place the video clips in order, let the scenes show what the words are saying, Cut the clips, keep the best portion of each, remove the rest. AI clips are often awkward at the very start and the very end. Trim both edges and the clip cleans up. Pre, add transitions, the manner of passing from one clip to the next, the rule, the fewer, the better, the plane cut is the purest, a light fade at a change of mood. Thinning, twirling, color transitions are the mark of the beginner. Four, lay the music on its track and to bring its level down chapter 10. 5. Color correction Most editors have a one-click filter or enhance. If your AI clips came from different tools, put the same filter on all, the colors fall into step, and the video feels stitched of one cloth. 6. Add a title card at the start 3-4 seconds and a closing card at the end the channel's name, a word of thanks. Export settings 1080p, MP4 format, 30 frames per second, the standard for YouTube. Keep the exported file in the final folder. One suggestion after the video is done, watch it once from beginning to end without stopping as a viewer, wherever your attention drifts, know that a cut is needed there, then show it to someone at home, the reaction of one first viewer teaches more than a hundred critics. Exercise. Join all your clips, voice, and music into your first complete video with title card and closing card, export it, show it to a family member and write down their first reaction. Chapter 12. Subtitles. A video that can be read. Most viewers today watch video without sound, on the bus, in the office, in bed at night, no subtitles, no viewers, and for content in a language like Mathalie the importance of subtitles is doubled. Subtitles in their original language build the habit of reading it, while Hindi or English subtitles bring in viewers who do not know the language at all, that is, your story travels the whole world. The standard format of subtitles is SRT, a plain text file in which three things repeat over and over Figure 9, a serial number, a timeline from which second to which second, and the text, this file can be made and corrected even in an ordinary text editor Notepad. But matching the timing by hand is laborious, and here AI helps again, two ways. First, automatic transcription, the auto captions feature in an editor such as cap cut listens to the video's voice and writes the subtitles itself. Timing included. In Hindi and English this is very accurate. A mathily voice it will usually hear as Hindi and write accordingly, then you correct the text it made, even so, correcting is far faster than writing from scratch, because the timing arrives ready made. Second, from the script, you already have the script written, give a chatbot your script and the video's total duration and say, divide this into SRT format, each subtitle at most two lines, at a comfortable reading pace, load the file in the editor and nudge the timings forward or back. The craft rules of subtitling, at most two lines at a time, roughly 32 to 40 characters per line. Each text stays on screen at least one second, a sentence breaks where the meaning allows I went slash to the market, no, I went to the market together, letters in white with a light dark shadow or strip behind, So they can be read even over a bright scene, place them at the lower middle of the screen, but not so low that the real format cuts them off. On YouTube, subtitles can be given in two ways, burned into the video joint at export from the editor or upload it as a separate SRT file which the viewer can switch on and off. The best method, give the original language subtitles as a separate file, and add Hindi and English as separate SRT files too. A chatbot will help with the translation the duty of checking it remains yours, thus one video reaches the viewers of three languages. Exercise, make the SRT of your video in its own language by either method, then make its Hindi or English translation file, run both with the video and check. Is the timing right? Is any line too long? Chapter 13. Publishing on YouTube. The video was made, now carry it to the viewer, YouTube remains the largest and the most Lasting platform, a video placed here keeps being watched year upon year, while on real-format platforms a video's life is a few days, so make YouTube the main house, let reels and short speed its windows. Making a channel is simple, sign into YouTube with your Google account and create one, choose the channel's name with thought, short, easy to say, and suggestive of the subject. The channel picture logo and banner too can be made with an AI image tool. And up o' time keep the checklist of figure 10 before you. A few points in detail. Title. The main matter in the first three or four words, with a title in your own language, adding Hindi or English in brackets helps the video surface in search, because seekers search in every language. Description. The first two lines are the most valuable. These appear in search results. Write the video's essence here. Key words included. Below them, the chapter list what comes at which minute, the list of sources, and channel introduction. Thumbnail. The viewer sees the thumbnail before the title, the rule, one image, one feeling, at most three words, and words large enough to be read on a small mobile screen make an attractive thumbnail with an AI image tool, but never a misleading one. What is in the thumbnail must be in the video where the viewer feels cheated and trust in the channel is gone. The AI disclosure, YouTube now expects that realistic looking AI generated or AI altered content be declared at upload in answer to the altered content question, this is not mere rule keeping, it is honesty with the viewer, for plainly imaginary styles such as animation the duty usually does not arise, but when in doubt, declaring is always the better course. Language setting Choose the video's language. Mathely is in YouTube's list, as are many others, this helps the video reach those searching in that language, and it strengthens the statistics of the language's content besides. After the upload, what then? Watch the response of the first hours and days, reply to comments, the early conversation waters the channel's roots, and keep regularity. One video a fortnight makes 24 in a year, and this bears more fruit than 100 videos at random. Viewers attach themselves to a program, not to scattered surprises. For reels and shorts, cut the most engaging 30-60 seconds of your main video into the 9-16 ratio with the editor's reframe or crop feature, and write at the end. Full video on the channel, this is the window that leads new viewers to the house. Exercise. Upload your video, completing all seven points of the checklist. Also cut a shorts version and upload it separately. After one week, compare the figures of the two views, watch time. After 14, special considerations for Mathili content. This chapter is the heart of this book. The tools are universal, but our purpose is particular, a Mathili for Mathila, and readers working in any less served language will find the same principles apply to their own. First, purity of language, AI tools, when writing Mathili usually let the shadow of Hindi fall across it, because they have learned far more Hindi, scripts, translations, Subtitles made by a chatbot. Check every one with your own eyes. The plain rule. The A.I.'s mathily is a draft, not an authority, where in doubt. Trust your ear. Speak the line aloud. Whatever grates on the ear is the shadow of Hindi. Second, the question of script. Mathily is written in Devanagari, and it also has its own ancient script, Turhuta mythilakshara. Keep the video subtitles and text cards in Devanagari. The most people will be able to read them. used Turhuta for beauty and identity, entitled cards, in the logo, in a watermark, thus the script stays before the eye, and curiosity awakens too, remember. AI image tools cannot yet write Devanagari or Turhuta letters correctly, always add written, text yourself at the edit, never have the AI write it. Third, authenticity of the visuals. Tell an AI an Indian village and it will produce a generalized North Indian scene, which is not Mithila, give the prompt Mithila's particular signs, the pond, the banyan, the mango orchard, the patty field, fish, pond, the thatched house, the tall-sea platform in the courtyard, wall ornament in the manner of Arapan. Even then, what comes will be Mathila-like, not Mathila, so wherever possible, blend in real photographs and footage by the method of Chapter 7. The mixture of AI scenes and real scenes gives the most authentic result of all. Fourth, the honor of Madhubani's slash Mathila painting, the AI can be told to generate in Madhubani style, and it will imitate the colors and the line. But pause here and think. Methila painting is a living tradition, the livelihood of thousands of artists rests on it and each of its manners Barney, Kachni, Godna, Gober carries its own lineage. An AI made Madhubani like image lifts the traditions appearance without its labor and its knowledge. My counsel, in your videos show real works by real artists with permission and credit. This honors the artist and strengthens your video at once. If you do use AI ornament in the Madhubani manner, say plainly that it is an AI mediation, not authentic Madhubani. 5. An inexhaustible store of subjects The field of mathily video is still nearly empty. Whoever makes makes first. Some directions. Children's material songs, tales, letters, the child audiences the fastest growing of all. Festival explainer C. H. Chathai, Samachakpa, Jhursital, Huat, Wai, Hau, recipes, folk tales and the verses of Vidya Patti presented with images, the vocabulary of village and home a visual dictionary of the words now slipping away, introductions to the places of Mathila. Each direction could be a channel in itself. And the last word, the patience of quality, in Methili the audience will be smaller than in Hindi. This is natural, but the loyalty of the Methili viewer is greater. They will comment, they will share, they will return again and again, look not at numbers but at relationships, a hundred devoted viewers are worth more than 10,000 indifferent ones. Exercise, choose one of the six directions above and plan three consecutive videos on its subject plus a one-line summary each. Three, because one video is an experiment, three are a direction. Chapter 15. Copyright, ethics, and responsibility. The powerful tool comes responsibility. This chapter is short, but bring its every line into practice. Copyright, the root principle, what you did not make, you do not use without permission, film songs, portions of others' videos, the text of books, other people's photographs, the rule covers them all. Everyone does it is no argument, YouTube's automatic system content ID catches it, and the channel bears the penalty. When freely licensed material to read the conditions, one says a tribution required, another no commercial use. Know also the question of rights over your own AI made material. In most tools terms, permission for commercial use of generated video and images comes with the paid plan, at a sometimes limited on the free one. For any tool you work with regularly, read its terms of use once, especially the ownership and commercial use sections. The dignity of persons, three iron prohibitions, no imitation of anyone's face or voice without written consent living or departed, both. Never show a person in a scene that lowers their honor, nor put in their mouth words they never spoke the deep ache, and take special care with the images and voices of children, before posting a video even of your own child. Consider that it is becoming public. Toward truth. A eye can make scenes that look real, and therefore it can deceive. The rule, let fiction be called fiction and news be news. When you make a video on a historical or cultural subject, check the facts against your own trusted sources. An AI chatbot speaks its errors with full confidence this is called hallucination. Wherever an AI scene looks real, declare it Chapter 13. A viewer's trust takes years to earn and one video to lose. And toward yourself. Let AI's convenience never turn into laziness. Take work from AI, not thought. The day you hand the script, the choices, and the final judgment over to AI, that day the video ceases to be yours, and viewers can smell the difference. Keep the tool in your hand, let not the hand become the tool. Exercise. Run a responsibility check on the first video you made, the music's license, permission for the image sources, the facts verified, the AI disclosure, keep the four answers in writing. Repeat this exercise with every video, within days it will become second nature. Chapter 16 The Practice Project, a first video from start to finish. Now all the threads in one place, this chapter is the model of one complete project, a 92nd video called CHathai, the Great Festival of Sun Worship, read it, then repeat the same sequence on a subject of your own. Day 1 Planning Chapters 3, 4, Purpose, to explain simply the spirit and the observance of Sirei Chathai, audience, Mathal families living away from home, the new generation especially, duration, 90 seconds, that is, 9 or 10 clips, a two-column script drafted by a chat pot, corrected by hand in three places, local detail added to the description of Nahikai, the feeling of the final line deepened, 10 scenes on the storyboard, the river-gat at dawn established, the courtyard where Thekua is being made, the supendora of the offering, the evening argia, the morning argia, the peran, and the closing card resolve. Day 2, Voice Chapter 8, the voiceover recorded in my own voice, at night, in a quiet room, paragraph by paragraph, polished with the editor's enhanced voice, total 82 seconds, which means the scene plan sits right. These three four, visuals chapters 5-7, a six-part prompt built for every scene thought out and mathily, translated into English. Methylas signs in each, a river ghat in north Bihar, banana stems, bamboo kania, fruit in the soup and aura, three clips came out well at the first attempt, four clips took two or three improvement cycles. For two scenes the tekua and the soup the image of video method was used, first the image perfected, then subtle motion given. And in one scene a real photograph of my own was used, the crowd at the gate, which no AI could ever have made. This mixture is the essence of Chapter 14. Day 5. Music and Assembly Chapters 10-11. AI Music. Slow, devotional, flute and emridang, 90 seconds, the license checked, the editing order, voice first, then clips, plain cuts, a fade in two places, one and the same light warm filter on every clip, on the title card CH-athai in Thirhuda and the full title in Devanagari. Day 6, subtitles and publication chapters 12-13, the Mathili SRT from the script, then Hindi and English translation files translated by Chatbot, checked by hand, the thumbnail, the scene of the evening argh, three words, hui chathai, ad upload, language Mathili, the AI disclosure switched on, music credit and sources in the description. Alongside, a 35-second shorts version. Day 7 Review Chapter 15, the responsibility check on all four points shown to the family, the suggestion came that the morning argia scene feels short, noted down for the next video. See, seven days, an hour or two a day, and from nothing to a published video, the first project will take longer than this, let it, by the third or fourth video the sequence becomes habit and the time falls by half. A last word. This book ends here, but the learning begins here. The tools will change. The models will change. The screens will change. But the framework you have learned plan, prompt, improvement cycle, assembly, responsibility will stand. Now go and tell the story of your Mathila in your own voice, in your own language. The Pentax A. Blossary. AI artificial intelligence, the capacity of a computer to learn and understand in a human-like way. Generative AI. AI that creates new material, text, images, video, voice. Prompt. The written instruction given to an AI. Negative prompt. The list of what is not wanted. Text to video. The method of making video from a written description. Image of video. Method of setting a still image in motion. TTS text-to-speech, the method of producing a spoken voice from written text. Avatar, a digital speaker who delivers text on screen. Lip sync, the matching of voice and lip movement. Voice cloning, making a digital replica of a voice. Credit, the usage currency of AI tools, each generation deducts some. Model, a particular version or engine of an AI. Clip, a short segment of video usually 5-20 seconds. Aspect Ratio, the ratio of a screen's width to its height, 16 colon 9 wide, 9-16 upright. Resolution, the fineness of the image, 1080p standard, 4K extra fine. Watermark, an identifying mark printed on content. Storyboard, the scene by scene sketch plan. B-roll, supporting footage beyond the main speaker. Timeline, the line of time in an editor on which clips are arranged. Track, a layer of the timeline, video, voice, music, and so on. Cut, trimming a clip, or passing directly from one clip to the next. Transition, the manner of joining two clips fade and the rest. Fade and slash out, slow emerging slash, slow vanishing. Render slash export, Turing the edited video into the final file. SRT, the standard file format of subtitles. Burned in, subtitles fixed permanently into the video. Thumbnail, the face image of a video. Hallucination, an AI's confident error. Content ID, YouTube's copyright checking system. Deepfake, a deceptive imitation of someone's face or voice. Monetization, earning from videos. Appendix B, A Collection of Prompts The lower ten ready prompt frames, fill the brackets with your own details and shape the final prompt by the method of Chapter 5. 1. Village Dawn A village in North Bihar, Dawn, light mist over the fields, thatched houses, a bayon tree in the distance, your detail, camera panning slowly to the right, soft golden light, realistic documentary style. 2. scene, a pond covered with lotus and water lilies, seasoned on the bank tree slash cat. Gentle ripples on the water, subtle movement of bird slash fish, camera moving forward very slowly, time of day light, semantic. Free, courtyard scene, the courtyard of a mythila home, a Tulsi platform, action, say, food cooking on the hearth, folk ornament on the wall, warm intimate light, close-up, realistic. 4. Festival scene. Preparations for name of festival, central object, say, soup and dora, lamps, courtyard or get, close-ups of busy hands, festive joy, warm light, documentary style. 5. Nature scene. Green paddy fields waving in the wind, time, the far horizon, bird and flight, drone view rising slowly, natural light, ultra-fine detail. 6. studied dish slash craft, an extreme close-up of object, background, light steam or sheen, camera circling the object slowly, soft studio-like light, appetizing slash appealing style. 7. Doreen illustration animation. Character description doing action. Place, hand-drawn animation style, soft watercolor, like a children's book illustration, slow pleasant motion. 8. imagination, an imagined scene of period methila, action, colors like an old palm leaf painting, slow-sulme camera, museum documentary style, the AI disclosure is obligatory here. This is imagination. 9. Title card background, an abstract ornamental pattern drifting slowly, color scheme, say, from million red and yellow, no letters, no figures, come even motion, 10 second loop, Add the lettering yourself at the edit. 10. Motion instruction for image video. In this image, gentle natural motion in element, say, water slash leaves slash cloth, camera static slash moving forward very slowly, nothing else changes. The image's original color and line remain intact. The Pendex C, a guide to the tools as of 2026. This list reflects the state of things at the time of writing 2026, Names and features keep changing. Learn the categories. Do not memorize the names before adopting any tool. Check three things. What the free plan gives, whether commercial use of the output is permitted, and where the watermark stands. Text-to-video slash image video, Google's Vivo, OpenAI's Sora, Kling, Runway, Pixverse, Seedence, Haleuo, Luma, Pika, Realistic Scenes, Cinematic Clips, Avatar Video, Synthesia, Digital Speakers, Lip Sync, Many Languages. Text-to-speech and voice cloning, 11 labs and others, services centered on Indian languages are also growing, a mathily option will appear there first, keep watching. AI Music, Suno, Udio and others, music and song from a description or from lyrics. Image generation, included in nearly every video tool, separately, Mijrani, and the The image features attach to the chatbots. Editing, cap cut, mobile and computer, simple auto-cactions included, DaVinci Resolve, computer, free professional grade, canva, thumbnails, banners, simple video. All in one plant video, tools of the in-video kind, which, given a topic, join script, visuals and voice on their own, good for quick work, but your own step is faint in them, so for learning, the step-by-step method of this book is better. Chatpot script, translation, SRT, brainstorming, Claude, chat GPT, Gemini. And last of all, this book too will age with time, but you will not. The teacher self-temper this book has built in you will apply to every new tool, when you seal a new tool, ask three questions, where is the prompt written, what is in the settings, where and how does the result arrive, and within five minutes you will be working it.

← Transcript index