विदेह First Maithili Fortnightly eJournal · ISSN 2229-547X
☰ विदेह मेनू
Videha · विदेह — offline browser utility

PDF · PPTX · DOCX · ZIP · चित्र औजारPDF Tools · Office Merge · ZIP Browser · Batch PNG/JPG · Speech/Text

✦ VERSION 32 · PDF · PPTX/DOCX merge · ZIP browser · thousands-of-images PNG↔JPG · Speech/Text ✦
browser-local processing — कोनो फाइल Videha सर्वर पर अपलोड नहि होइत · files stay on your device
🗜️
PDF फाइल एतऽ खसाउ वा क्लिक करू · Drop PDFs here or click
एकाधिक फाइल चुनि सकैत छी · select multiple files
Smart/Force मोड scan/image PDF छोट करबाक कोशिश करैत अछि। Searchable text बचाबय लेल ध्यान राखू: raster compression text केँ image बना सकैत अछि।
65% · संतुलित
120 DPI
रंगीन हटा कऽ आओर छोट करू · drop colour for smaller scans (good for text scans)
✅ Smart mode: scan/image PDF केँ छोट करबाक कोशिश करैत अछि। Output पैघ भेल तँ असली PDF राखैत अछि — फाइल size बढ़त नहि।
ध्यान: ई browser tool OCR नहि जोड़ैत। Searchable PDF पर compression करब त text image भऽ सकैत अछि।
🖼️
PDF फाइल एतऽ खसाउ वा क्लिक करू · Drop PDFs here or click
हर पन्ना एकटा JPG बनत · each page becomes a JPG
PNG बड़ नीक गुणवत्ता मुदा पैघ फाइल · PNG is lossless but larger
85%
150 DPI
खाली छोड़ू = सब पन्ना · blank = all pages
🔎
देवनागरी/तिरहुता/English PDF सँ पाठ निकालू · Extract Devanagari/Tirhuta/English text
केवल embedded/selectable text निकालैत अछि — scanned image PDF लेल OCR नहि
Cover/graphics पन्ना छोड़बाक लेल Pages मे 2- लिखू। ई OCR नहि, केवल embedded/selectable text सुधारि निकालैत अछि।
ई OCR नहि अछि। PDF मे text layer रहत तँ निकालत; केवल image-scan PDF मे खाली/अल्प text भेटि सकैत अछि।
खाली = सभ पन्ना; 2- = दोसर पन्ना सँ अन्त धरि; -5 = पहिल पाँच पन्ना
गलत लाइन-क्रम/column mixing हो तँ बदलू।
Embedded/selectable text केँ पन्नाक क्षेत्र अनुसार छाँटत। Cover/design पन्ना लेल left/right/top/bottom उपयोगी।
extra spaces/blank lines साफ करू · normalize spaces and blank lines
border/QR/barcode/decorative repeated lines हटाबय केर कोशिश · remove decorative junk lines
🔎 Parsing mode: searchable/selectable PDF सँ Unicode text निकालत। देवनागरी U+0900–097F, तिरहुता U+11480–114DF, आ English/Latin A–Z पहचानैत अछि।
Cover/design पन्ना मे embedded text रहितो क्रम बिगड़ि सकैत अछि; एहि लेल preset, layout, Pages=2-, Parse Crop/Zone, आ Junk हटाउ जोड़ल गेल अछि। Scan-only PDF केँ searchable बनाबय लेल OCR चाही।
👁️
स्कैन PDF/चित्र सँ OCR पाठ निकालू · OCR scanned PDF/images
देवनागरी · तिरहुता · English — local OCR engine/model रहला पर offline चलत
⚠️ OCR गुणवत्ता चेतावनी: cover, poster, photo-background, border, QR/barcode, decorative page पर OCR बहुत खराब हो सकैत अछि। साफ किताब-पन्ना/typed text page पर चलाउ। Cover/graphics लेल पहिले page range सँ ओ पन्ना छोड़ू वा नीचे “Sparse/cover text” preset चुनू।
पहिल विकल्प साफ पन्ना लेल अछि। Cover मे सजावट/रेखा/चित्र केँ अक्षर बुझि OCR बिगड़ैत अछि।
Mixed page पर केवल “hin” नहि चुनू; English + Devanagari चुनू। Tirhuta लेल genuine tirhuta.traineddata.gz चाही — tir Tigrinya अछि।
Cover छोड़य लेल 2- लिखू। खाली छोड़ू = सभ पन्ना।
150 DPI
साफ printed page पर threshold नीक; cover/graphics पर none/gray नीक।
Book page लेल 6/4; cover/poster लेल 11 कोशिश करू।
Cover/poster मे सजावट आ QR/barcode हटाबय लेल केवल असली text क्षेत्र OCR करू। Whole cover OCR करब बेसीतर खराब होयत।
OCR output मे blank lines/extra spaces साफ करू
👁️ Improved offline OCR: v20 मे reusable OCR worker, image-preprocessing, layout presets, आ Crop/Zone जोड़ल गेल अछि। Cover/poster पर पूरा पन्ना नहि, केवल text क्षेत्र OCR करू।
Expected local files: assets/ocr/tesseract.min.js, worker.min.js, tesseract-core.wasm.js, and assets/ocr/lang/eng.traineddata.gz / hin.traineddata.gz / Devanagari.traineddata.gz. For Tirhuta use custom assets/ocr/lang/tirhuta.traineddata.gz.
OCR लेल localhost सँ खोलू। Direct file:// खोलला पर worker/WASM fail भऽ सकैत अछि।
✂️
PDF केँ पन्ना वा size अनुसार बाँटू · Split PDF by pages or approximate size
No artificial file-size cap; practical capacity depends on browser RAM/device power.
Page split exact अछि। Size split browser-rendered PDF मे approximate अछि, कारण output size image quality/DPI पर निर्भर करैत अछि।
Semicolon अलग output file बनायत।
130 DPI
82%
⚠️ Split output एहि browser tool मे rendered/raster PDF होयत; searchable text layer सुरक्षित नहि रहि सकैत अछि। Unlimited capacity = code limit नहि, browser/device memory limit लागू।
🧩
PDF, PNG, JPG जोड़ि एक PDF बनाउ · Merge PDF/images into one PDF
Drag files, reorder, then merge. No artificial count/size limit; browser memory applies.
130 DPI
82%
PDF pages are rendered; image files are placed as pages.
⚠️ PDF merge एहि standalone version मे rasterizes PDF pages through canvas. Use lower DPI for very large files.
🖼️
PNG/JPG split, merge, PDF बनाउ · Image split/merge/convert
Multiple PNG/JPG files; output ZIP or one merged file/PDF.
rows columns
JPG/PNG for split or merged image; PDF makes one PDF from selected images.
88%
Large image merge/split uses browser canvas; very large panoramas may exceed browser canvas memory.
📊
PPTX फाइल सभ चुनू · Select PowerPoint .pptx files
Natural numeric order + manual ↑/↓ reorder · processing stays in the browser
Merge queue
GitHub/browser version safely targets modern .pptx. Legacy binary .ppt and macro-heavy .pptm are intentionally not accepted, because a browser cannot guarantee Microsoft Office-compatible conversion without PowerPoint.
Ready.
📝
DOCX फाइल सभ चुनू · Select Word .docx files
Natural numeric order + manual ↑/↓ reorder · images and common OOXML relationships preserved
Merge queue
Browser merge preserves the document body, ordinary images, hyperlinks, tables, numbering and common related parts. Conflicting custom style IDs use the first document's definition; very complex tracked changes, macros, protected forms, footnote/comment ecosystems or section-specific headers may need Word for perfect fidelity. Legacy .doc/.docm are not accepted.
Ready.
🗃️
सैकड़ो ZIP भीतरक कोनो फाइल देखू/निकालू · Browse any files inside hundreds of ZIPs
PDF · DOCX · PPTX · TXT · HTML · XML · MP3 · MP4 · PNG · JPG · JSON · CSV · fonts · any ordinary file
0 files 0 selected
ZIP path traversal such as ../, absolute paths and drive-letter paths are blocked. Preserved-folder output is grouped below each source ZIP name to avoid collisions. The table renders at most 1,000 matching rows at once for browser responsiveness, but filtering/extraction applies to the complete scanned set.
✓Source ZIPTypePath inside ZIPSize
Select ZIP files and press “ZIP scan”.
🔄
हजारो PNG/JPG केँ ZIP सँ एक बेरमे बदलू · Batch-convert thousands of images inside ZIPs
Multiple ZIPs · sequential/chunked canvas processing · one downloadable result ZIP + manifest
ZIP queue
90%
transparent pixels are filled with this background
Images are decoded and converted one at a time, then canvas/ImageBitmap resources are released. This greatly reduces crashes with thousands of files. Final ZIP creation still needs enough browser RAM for the resulting archive.
Ready.
▶️ Standalone browser HTML YouTube सँ audio directly download/convert नहि कऽ सकैत, कारण YouTube restrictions/CORS आ copyright/terms issues। ई panel केवल authorized workflow, URL checking, notes, आ external lawful conversion reminder लेल अछि।
Maithili speech recognition browser मे प्रायः native नहि; mai-IN fail भेल त hi-IN fallback कोशिश करू।
Chrome/Edge मे Web Speech API online service प्रयोग कऽ सकैत अछि; offline guarantee नहि।
MP3/WAV/M4A/OGG चुनू; फाइल local preview लेल रहत।
Ready. Browser support required. Audio upload preview added.
1.0× 1.0
🔊 Browser SpeechSynthesis voice playback direct MP3 export नहि दैत अछि। “Record & download” बटन microphone/system-audio recording workflow चलायत: browser audio/mpeg support करैत अछि तँ MP3, नहि तँ WebM/OGG audio download होयत। Unlimited = no artificial text-length cap; browser voice/memory limit लागू।
Ready. MP3 depends on MediaRecorder audio/mpeg support; fallback file may be WebM/OGG.