We Don’t See the Same Thing
Artificial intelligence can recognize an image. Understanding what it means to the person looking at it is another story.
And I wonder why, standing in front of that same image, I want to cry and you simply keep walking.
I enter MALBA and come across a work by the American artist Flavin: tubes of light that transform the space. I keep going. I reach the Cuban artist Belkis Ayón. I stop in front of her work on the Abakuá Secret Society. I look at it and something happens to me: I want to cry. My foreign friend observes it for a few seconds and keeps walking.
In my head, I have an uninterrupted conversation with the world around me. Those of us who have devoted our professional and personal lives to art move through life in observation mode. It is a permanent state. The senses switched “on.” At the end of the day, as if you had somehow become part of every work you saw, you are no longer sure whether your body is skin, pixel, clay, or oil paint. Emotions accumulate in layers.
What do we feel when we look at a work of art? Are the sadness, beauty, fear, or calm we recognize in an image universal? Or do we also learn how to feel through our culture, our history, and the images that surround us?
A newly published study has taken these questions into the territory of artificial intelligence.
The experiment has just been published as a preprint and cannot yet be considered established research. But the question it raises is too interesting for me to ignore.
It is called ArtECulture: Benchmarking Culture-Conditioned Visual Emotion Understanding in Multimodal Large Language Models and was presented in August by a team of researchers primarily affiliated with the National University of Singapore (NUS). A team specializing in technology decided to ask a question that touches directly on the territory of culture.
The experiment
ArtECulture brings together 6,792 works of art and 92,062 emotional annotations, built around three cultural contexts: Chinese, Arab, and Anglophone. The dataset seeks to include Western and non-Western content in balanced proportions, precisely to avoid one of the problems that many AI systems carry with them: having learned about the world primarily through data produced from a Western perspective.
The researchers tested 16 multimodal language models, both open and proprietary, in a zero-shot setting: that is, without giving them specific training beforehand to solve the test. The task went far beyond recognizing objects or describing an image. The AI had to try to predict what emotion a work might provoke in people belonging to a particular culture and explain the reasons behind that interpretation.
In other words, it had to change perspective. It was not enough to say: “there is a person crying.” The question was closer to: “how might someone from this cultural context interpret this scene emotionally?” And that is where the problems began.
The best model achieved less than 50% accuracy. The study also found that the systems struggled more when they had to move away from the cultural frameworks most familiar from their training data. The authors describe this phenomenon as a form of Anglocentric affective bias. The machine can recognize an image. What it still struggles with is understanding the world that exists around that image.
A work of art never comes alone
The research brings to mind an idea developed almost a century ago by art historian Erwin Panofsky. For Panofsky, looking at a work of art is not simply a matter of identifying forms. A first level of looking recognizes objects and gestures, a second identifies themes and symbols, and a deeper interpretation tries to reconstruct the cultural context that gives them meaning.
The difference may seem small, but it is enormous. A figure with raised arms may be celebrating. It may be praying. It may be protesting. It may be saying goodbye. The body is the same; the meaning changes.
An AI can recognize the gesture. Understanding why that gesture means something different depending on who is looking is another story. Culture is also inside the image, even when we cannot see it.
A museum guide for machines
The authors of ArtECulture explore the problem, but they also test a possible solution: they build a knowledge base containing concepts and emotional associations linked to different cultures and allow the models to consult that information before producing their answers.
This procedure, known as retrieval-augmented generation, improves both emotional prediction and the explanation provided by the AI. It is almost like giving it a museum guide. But a particular kind of guide. It does not merely explain what a work represents. It tells the machine that the same image can activate different associations in different communities.
How much of what we feel do we learn?
A few years ago, research projects such as ArtEmis began exploring the relationship between artistic images and human emotions. That project brought together tens of thousands of works from WikiArt and hundreds of thousands of human responses to study not only what emotion an image produced, but also how people explained that reaction. ArtECulture adds a decisive variable: culture.
For a long time, we talked about artificial intelligence as though it could learn some kind of universal language of images. As though, after observing millions of photographs, paintings, and videos, it could discover what everything means. That universal dictionary does not exist. Red can mean danger, passion, celebration, or mourning. A mask can mean carnival, ritual, protection, or threat. A landscape can represent nature, territory, memory, or identity. It depends on who is looking.
The next challenge is understanding
The conversation around AI and art has focused in recent years on one question: can a machine create? We already know that it can produce astonishing images. Music, texts, videos, and visual experiences as well. But ArtECulture proposes shifting the discussion toward understanding why a work matters. As we try to teach a machine to interpret our emotions, we discover something about ourselves: looking is never a purely visual operation.
We look with our individual and collective memory. With our language. With our education. With the images of our childhood. With the symbols we learned to recognize. With the stories of the place we come from. With our layers.
We look from somewhere. AI is still learning where that somewhere is. It can already recognize a tear. Now it has to learn that not all tears mean the same thing.
A work of art never ends with what we see. It begins there.