A command-line benchmark for testing model bias in a Malaysian cultural, linguistic, and legal context.
Most fairness benchmarks flatten away the details that matter in Malaysia. MY-FairnessBench keeps ethnicity, socio-economic history, Manglish, and code-switching in the test instead.
What I made
The CLI sends the same prompt through a set of international models and ILMU, Malaysia's first multimodal LLM. It then saves the responses to a timestamped Markdown report for comparison.
This is still a research prototype. The useful part is the framing: local bias cannot be measured well with examples written for a different social and legal setting.
How it works
5 stepsFrame a prompt around Malaysian context.
Send the same case through each configured provider.
Keep the raw responses together.
Look for differences in language and treatment.
Write a dated Markdown report.
Repository breakdown
This comes from GitHub's detected language breakdown. It measures source size, not the time or difficulty of the work.
Language mix
GitHub source bytesIndependent research build
Research prototype
Private GitHub repository