YTL AI Labs has launched Nemotron-Personas-Malaysia, an open dataset containing 1.35 million synthetic personas designed to help developers build AI models that better reflect Malaysia’s diverse population. The dataset was developed in collaboration with NVIDIA and is based on official Malaysian demographic and labour statistics.
The personas were generated from 150,000 base records derived from sources including the Department of Statistics Malaysia, OpenDOSM, MyCensus, eStatistik and national labour-force publications. YTL AI Labs said the dataset does not contain personal data that can be linked to any real individual.

Diverse Demographics
Each synthetic persona represents a hypothetical individual generated using probability tables derived from demographic statistics. The dataset covers demographic characteristics including age, gender, ethnicity, occupation and region, alongside personality traits based on the OCEAN psychological model.
It represents major ethnic groups in Malaysia, including Malay, Chinese, Indian, Kadazan-Dusun, Bajau, Murut, Iban, Bidayuh and Melanau. The dataset also spans different religions, occupations and regional backgrounds to provide a broader representation of the country’s population. According to YTL AI Labs, each persona contains 39 data fields, allowing developers and researchers to use the dataset for testing and training AI systems across a range of scenarios.
Notably, Nemotron-Personas-Malaysia is the first dataset in NVIDIA’s Nemotron-Personas collection to feature Bahasa Melayu as its primary language. The collection also includes datasets representing countries such as the US, Japan, India, Singapore, Brazil, France, South Korea, El Salvador, Vietnam and Belgium.

For AI Testing And Development
The company said the dataset can be used to evaluate how AI systems perform across different Malaysian demographics, languages, cultures and socioeconomic backgrounds. For example, a customer service AI could be tested against synthetic customers of different ages, states, occupations and language backgrounds.
Similarly, financial institutions could use the personas to assess whether their AI systems can understand different ways Malaysians express financial needs. Government agencies and researchers could also use the dataset to identify whether digital services or AI models perform differently across demographic groups.

YTL AI Labs CEO Foong Chee Mun said the development of AI should focus not only on the size of the underlying models, but also on how well they understand the local context and society they are intended to serve. He added that the goal is to allow Malaysians to build AI systems using global technology while retaining control over the local context and priorities that shape how the technology is used.
Nemotron-Personas-Malaysia forms part of YTL’s sovereign AI efforts, which also include its ILMU model series and YTL AI Cloud infrastructure. The dataset is publicly available through Hugging Face for developers, researchers and enterprises to access.
(Source: The Edge Malaysia)

