News

DAGOBAH: tables, AI will understand

June 22, 2020 - Big Data & AI

Human activities produce a massive amount of raw data in tabular form. EURECOM and Orange are developing the DAGOBAH semantic annotation platform. The aim is to deploy a generic solution to optimize AI applications such as personal assistants, but also to facilitate the management of complex datasets in any enterprise.

In everyday life, a keyword search on the Internet is often enough to fill in our thousands of memory gaps, clarify our doubts or satisfy our curiosity. The results even anticipate our needs by offering more information than we ask for: a singer's biography, a few of his songs, the dates of his next concerts... But have you ever wondered how the search engine always provides the answer to your questions? In order to display the most relevant results, computer programs need to understand the meaning and nuances of the data (often in tabular form) used to answer users' queries. This is one of the major challenges of the DAGOBAH platform, the result of a partnership between EURECOM and Orange research teams that began in 2019.

DAGOBAH's goal: to automatically understand tabular data produced by humans. The absence of an explicit context for this type of data, compared with text, means that understanding it depends on the knowledge of the reader. "A human knows how to detect the orientation of a table, the presence of headers or row merges, the relationship between columns, and so on. We want to teach the machine this natural interpretation", describes Raphaël Troncy, a data science researcher at Eurecom.

The art of harnessing encyclopedic knowledge

Having identified the shape of a table, DAGOBAH sets out to understand its content. Let's take the example of two columns. The first lists directors' names, the second film titles. How does DAGOBAH go about interpreting this dataset, the nature and content of which it knows nothing about? It performs a semantic annotation, i.e. it places a sort of label on each element of the table. To do this, it must determine the nature of the content of a column (names of directors, etc.) and determine the relationship between the two columns. Here: director - directed - film. Except that one element can mean several things. For example, "Lincoln" refers to a family name, a British or American city, the title of a Steven Spielberg film, etc. In short, the platform needs to remove any ambiguity about the content of a cell from its overall context.

To achieve its goals, DAGOBAH interrogates existing generalist encyclopedic knowledge bases (Wikidata, DBpedia). In these databases, knowledge is often formalized and associated with attributes: "Wes Anderson" is associated with "director". In order to process a new table, DAGOBAH compares each element with its database and proposes attribute candidates: "film title", "city", etc. But only one element can remain. But only one candidate remains. So, for each column, the candidates are grouped together and put to a majority vote. The nature sought is then deduced with a greater or lesser probability.

However, this method has its limits when it comes to complex tables. For example, industrial data may not only be suitable for the general public, but may also contain statistics relating to business knowledge or specific scientific data that is difficult to identify.

Neural networks to the rescue

To reduce the risk of ambiguity, DAGOBAH uses neural networks and the lexical embedding technique. The principle: represent the content of a cell as a vector in a multi-dimensional space. Within this space, the vectors of two semantically related words are grouped together geometrically in the same place. Visually speaking, directors group together and so do film titles. The application of this principle to DAGOBAH is based on the assumption that elements in the same column must be sufficiently similar to form a coherent whole. "To remove ambiguities between candidates, we group candidate families in vector space. The problem then comes down to selecting the most relevant group in the context of the table under consideration", explains Thomas Labbé, data scientist at Orange. This method becomes more efficient than a simple majority vote search when table context information is scarce.

However, one of the drawbacks of using deep learning is the lack of visibility over what is happening within the neural network. "We modify hyperparameters that we turn like kitchen knobs to obtain the best results. The approach is very empirical and time-consuming, as we repeat the experiment many times," explains Raphaël Troncy. The approach is particularly time-consuming. The teams are also working on scaling it up. In this respect, Orange's infrastructures dedicated to Big Data are a major asset. Ultimately, the researchers aim to deploy an end-to-end, all-purpose approach that is sufficiently generic to meet the needs of a wide variety of applications.

Towards industrial applications

Semantic interpretation of tables: a goal, not an end. " Working with Eurecom gives us almost real-time knowledge of the latest academic advances, as well as an informed opinion on the technical avenues we plan to pursue," says Yoan Chabot, artificial intelligence researcher at Orange. DAGOBAH's use of encyclopedic data enables it to optimize natural language question/answer engines such as voice assistants. But the Holy Grail will be to offer an automatic processing solution for specific business knowledge in an industrial environment. " Our solution would enable us to address the market of private players, and not just public ones, with a view to internal use in companies that produce massive amounts of tabular data," adds Yoan Chabot.

The challenge will then be considerable, as industry has no knowledge graphs to which DAGOBAH can refer. The next step will therefore be to semantically annotate datasets from embryonic knowledge bases. In order to achieve their objectives, the academic and industrial partners have committed themselves for the second year running to an international challenge on semantic annotation, a theme that is very much in vogue in the scientific community. Over a four-month period, they will have the opportunity to test their approach in real-life conditions, before comparing their results with the rest of the international community next November.

Further information DAGOBAH: A picture speaks only to those who know how to annotate it

Latest news

[A GREAT STORY] State-of-the-art computing infrastructure supporting research in artificial intelligence

For the past ten years, Télécom Paris has had a large-scale computing infrastructure, established in part thanks to the support of the Carnot TSN Institute. This world-class platform allows the school’s research teams and students to train and test their AI models free of charge.
,

[BELLE HISTOIRE] AI to optimize robot-assisted knee osteoarthritis surgery

As part of a thesis conducted with Ganymed Robotics and LaTIM (a laboratory under the joint supervision of IMT Atlantique, a component school of the Carnot TSN institute), Anna Gounot is developing AI models capable of predicting the state of knee cartilage from scanner images, in order to improve the precision of prosthesis fitting to treat osteoarthritis.

[VIDEO] Hadaptic Evident: an experimental platform at the heart of digital health

At the crossroads of digital technologies, healthcare and applications, the Hadaptic Evident platform at Télécom SudParis, a component school of the Carnot TSN institute, is a unique experimentation and co-innovation facility. Awarded the label of the Carnot institute Télécom & Société numérique, it is a concrete illustration of the ability of academic research to respond to the major societal challenges linked to autonomy, ageing and well-being.

Need more information?

News
© 2022 Carnot Télécom & Société Numérique | Legal Notice