Share it

Artificial Intelligence (AI) has become a state policy. The United States prioritizes the principles of capital, leaving all important decisions in private hands. The freedom (of making money) of the individual above all. China, on the other hand, controls, through the omnipresent state, the advances and uses of AI, following the Eastern ethic of prioritizing collective goals over personal ones (are you a good citizen?). And in the middle of both, Europe is looking for its place. A fit that allows us to combine the individual rights and social well-being that characterize the common European principles, with technological innovation and the valorization of data. For a decade, this search has seemed incapable, inefficient, and fruitless, with a couple of notable exceptions.

The The most relevant European contribution has been the GDPR. One ambitious and demanding legislation on the management of personal data that forces the two alternative superpowers to reconsider their positions as leaders. Setting the global data management agenda is an unprecedented milestone in the field of digital rights. Beyond the GDPR, it remains to be seen whether the guides for the ethical and trustworthy use of AI, promoted by the European Commission (HLEG-AI Guidelines) and, especially, the legislation being prepared for the European level (AI Act), will be successful or not.

But beyond these ethical and legislative contributions, what else can Europe contribute?
Until 2022, it seemed that the European research fabric, distributed, linked to teaching and competitive funds, was unable to keep up with the centralized and hyper-specialized entities of the US and China. The world, which is headed towards a few American and Chinese companies (e.g. e.g., OpenAI and DeepMind) would set the public and industrial agenda for AI, given the apparent indifference of Chinese secrecy. Competing with companies that have agile processes, lacking bureaucracy and ethical considerations, is complicated. Even more so when the value of contributions is measured in impact on public opinion. Spending millions of dollars training an AI model just to have the most expensive one of all? Oh, yeah!

Europe has not entered this game, rightly so, even though it has an adequate computational infrastructure. On the one hand, well-regulated public supercomputing centers cannot compete in flexibility in the marketing race that is 21st century AI. On the other hand, the limited scientific relevance of making larger AI models simply because we can does not justify the enormous environmental and economic expense associated with them.

That’s how things are., the most impressive AI seemed destined to be a technology controlled by a few, with a subscription-based business model, and where trained models are the most valuable business property.. Proprietary models such as GPT3 (for automatic text generation) or DALL-e (for automatic image generation) seemed unattainable. Until 2022. This year has changed everything. An initiative funded by French research institutions, with the participation of researchers from over sixty nationalities (BigScience), trained a natural language model the size of GPT3. In the spring, BLOOM was released, the largest natural language model for public and free use. Any individual or company can pass the Turing test with a desktop computer. But that’s not all. In August, Stability.AI, a European company, invested €600K in training an image-generating model called StableDiffusion, achieving results comparable to DALL-E2, and then release it, allowing public and commercial use of the model. The community exploded, and the tool’s popularity has already surpassed that of existing proprietary models, putting Stability.AI center stage. There is already talk of multi-billion dollar offers for the company, which will help overturn the dominant mentality that said “a private AI model is more valuable than a public one.”

BLOOM and StableDiffusion They are tools that will enable the creation of public and private initiatives in all sectors related to natural language and image. They will be the seed of many products and services that we will see appear in our lives during 2023, revolutionizing sectors such as audiovisual, art, advertising, personal care and entertainment. But that’s not all. BLOOM and StableDiffusion have influenced the prevailing philosophy of their competitors. In a reactive move, OpenAI released in September the code and model of Whisper, a voice-to-text transcriber, capable of interpreting convoluted accents in 52 languages ​​(where Catalan is the tenth best performing). The oil slick from the open models is starting to spread.

The European route It has been slow in coming, and 2022 could be considered the year of its birth. A path of open code and models, where innovation is democratized and accessible, overcoming old dogmas about the control of resources, and moving towards meritocratic competition with ethical principles. It remains to be seen to what extent the European bureaucracy will embrace and support this path.

Now, the most complicated part remains to be done. There are a host of ethical and legal aspects to consider, arising from training, releasing and using these models, which now require urgent attention. Europe is called to play an essential role, balancing fundamental rights, equity and social progress. Questions like, is it legitimate for a model to learn from copyrighted data? Is this compatible with the survival of sectors such as the artistic and cultural? Is it appropriate to use models trained on data that reflect the worst biases in our society, reproducing and perpetuating discrimination? The answers we find to these questions will shape the future of AI and our society. The availability of open models allows us to start working on solving these challenges today, but it also forces us not to leave it for tomorrow.

Dario Gasulla
Dario Garcia Gasulla

Polytechnic University of Catalonia (UPC); Barcelona Supercomputing Center (BSC)

Other articles

En aquest context de creixement i evolució dels models de llenguatge constant, és important  que considerem adoptar i adaptar petits models de llenguatge quan desenvolupem eines basades en models de llenguatge natural.

On December 13, 2023, a Workshop on AI Training in Catalonia, organized by CIDAI and AIR. The workshop brought together professionals from the academic, industrial […]

Marina Alonso Poal, Senior Data Scientist @ NTT DATA AI Center of Excellence Adil Moujahid, Technical Manager @ NTT DATA AI Center of Excellence   […]

CIDAI