Advancing Myanmar Optical Character Recognition: Introducing the MyanmarOCR-ImageText Dataset
A technical review of Khant Sint Heinn's 41,664 image-to-text synthetic dataset for Burmese OCR and Vision-Language Models on Hugging Face.
Category
Deep dives and reviews on open and benchmark datasets for ML training.
A technical review of Khant Sint Heinn's 41,664 image-to-text synthetic dataset for Burmese OCR and Vision-Language Models on Hugging Face.
A technical breakdown of Khant Sint Heinn's 1,216 parallel Pali-Burmese dataset featuring Romanized phonetics and usage context on Hugging Face.
A technical breakdown of Khant Sint Heinn's 3,030 human-curated Burmese sentence dataset sourced from news portals and hosted on Hugging Face.
A technical breakdown of Khant Sint Heinn's 1,349 hand-written dialogue entries designed to fine-tune Large Language Models into a protective 'Boyfriend Roleplay' persona on Hugging Face.
A technical breakdown of Khant Sint Heinn's 2,503 human-verified English-Burmese parallel dataset extracted from Hugging Face Course video subtitles.
A technical breakdown of Khant Sint Heinn's 1,600 manually labeled Burmese sentence dataset designed for context-aware disambiguation between 'လဲ' and 'လည်း'.
A technical breakdown of Khant Sint Heinn's 12,021 human-verified Myanmar-English parallel dataset on Hugging Face.
A comprehensive review of Khant Sint Heinn's 102,600 row Myanmar spoken vs. written classification dataset on Hugging Face.