As 2022 comes to an end, one of the trends that the AI field is leaning towards exploring in 2023 is responsible AI. Unfortunately, biased algorithms have not only shown the extent to which machine learning can be abused, but also exposed unfair societal norms and discrimination.

Many tech giants and AI think tanks have begun focussing their resources on developing responsible algorithms. Today, as a part of the Google for India event, the tech giant announced a $1 million grant to IIT Madras for conducting research into the various aspects of bias in artificial intelligence. It also announced the launch of Project Bindi, an undertaking that aims to fight against fairness and social appearance issues in India. 

Biased AI, especially when applied to a country as diverse as India, presents an issue that needs solving before it becomes a problem. While there are already instances of algorithmic biases causing issues of discrimination in India-focused applications, such initiatives will not only shed light on the matter but also prevent the amplification of existing biases. 

Why IIT Madras? 

Google chose to reward  IIT Madras with the research grant mainly because of the institution’s presence and veracity in the Indian AI field. It has discrete research centres for data modelling, deployable AI, scientific research in sports and AI deployment in the medical field. 

The college also boasts an excellence centre for data science and AI backed by the Robert Bosch group. As a part of its AI4Bharat initiative, IIT-M also recently inaugurated the Nilekani centre, which conducts research into building Indian language technologies. The institute also conducts specific research into explainable AI and fair and social good in algorithms – a matter of importance for Google. 

Interestingly, the company has a track record of shutting down teams focusing on ethical AI projects, as well as a history of racist and sexist bots. However, CEO Sundar Pichai had pledged support for more ethical AI projects back in 2021, as seen by the recent spate of announcements in Google for India 2022.  

Google’s grant to IIT-M focuses particularly on training natural language processing models to mitigate gender bias. This will be done through the establishment of a multi-disciplinary centre for responsible AI. 

AI bias in India

AI has a history of showing the mirror to the ugly side of human biases. From bias caused by existing societal norms to bias derived from improper datasets, this problem has been plaguing the Western AI community. A society as diverse as India presents a unique challenge to those looking to tackle responsible algorithms in the country, as multiple biases need to be identified and worked against. 

Even though India’s biases don’t seem to be documented well in datasets used by AI, they are ever-present in daily life. An example of this is a study conducted by Google Research that found bias arising from data that polled only Internet users. Due to the fact that only 50% of India’s population uses the Internet, it was found that collected datasets under-represented Muslim and Dalit populations due to their lack of Internet use. 

Even in datasets commonly used by AI researchers, Indians are misrepresented. The ImageNet dataset, one of the most-widely used datasets for facial recognition, contains less than 3% of Indian and Chinese faces. These kinds of biases will not only skew algorithms towards under-representation of global diversity, but also create a loop of pushing stereotypes associated with marginalities. 

A bigger problem for training language models for India is the lack of availability of data for some of the lesser spoken languages. Google’s step towards adding Hinglish support is a welcome change, but India has a range of accents and variations in English itself. The vernacular diversity of India, combined with the difficulty in collecting accurate last-mile language data, will lead to the creation of biased datasets, which Google aims to remedy through research. 

A prime example of this is the Northeast states. Commonly referred to as the seven sisters, these states are linguistically diverse to a great extent. Moreover, they are also under-represented in many AI applications, as there is barely any data on the languages spoken in the region. Measures such as AI For All, AI4Bharat, and Google’s 100 language model should take these parts of the country into consideration regardless of the difficulties involved. 

Conclusion

While Western biases have been accurately described and quantified, studying discrimination in India is a long and arduous process. Although, steps are being taken to avoid gender bias in algorithms made for India, caste bias, discrimination on basis of skin tone, and religious bias are still yet to be quantified. 

It is almost impossible to completely eliminate bias in AI, but a good step forward is looking at how other fields handle the difference in perspectives worldwide. Normative data, such as those used in statistics, can be used to describe the various characteristics contained in any given population, thus providing a tool to dispel the closed perspectives that bias usually causes. 

Google Cloud chief Thomas Kurian also believes that fairness of the models is critical for AI adoption. He said as models get more and more sophisticated, people often forget about other important issues. 

Google seems to be going all out when it comes to responsible AI. For instance, this paper by Google Research, shows different research pathways to achieve algorithmic fairness. From problem formulation, to optimizations in data collection, to training bias-free models, the researchers have offered a clear picture of how we can reduce bias in algorithms made for India. 

The post Google Commits to Solving AI Biases in India appeared first on Analytics India Magazine.