A framework to operationalize and automate the data integration lifecycle
(English) Data plays a key role in today’s world. Many organizations collect and store massive amounts of data from many different data sources. As a result, these data collections show a diversity in structure and semantics that grows as the data sources expand and evolve. These factors challenge t...
| Autor: | |
|---|---|
| Formato: | tesis doctoral |
| Fecha de publicación: | 2025 |
| País: | España |
| Recursos: | Universitat Politècnica de Catalunya (UPC) |
| Repositorio: | UPCommons. Portal del coneixement obert de la UPC |
| Idioma: | inglés |
| OAI Identifier: | oai:upcommons.upc.edu:2117/442278 |
| Acesso em linha: | https://hdl.handle.net/2117/442278 https://dx.doi.org/10.5821/dissertation-2117-442278 |
| Access Level: | acceso abierto |
| Palavra-chave: | Data Integration Data Discovery Knowledge Graphs Data Wrangling 004 - Informàtica Àrees temàtiques de la UPC::Informàtica |
| id |
ES_ae270027a2c4f47a7dde7b21af0eccd0 |
|---|---|
| oai_identifier_str |
oai:upcommons.upc.edu:2117/442278 |
| network_acronym_str |
ES |
| network_name_str |
España |
| repository_id_str |
|
| dc.title.none.fl_str_mv |
A framework to operationalize and automate the data integration lifecycle |
| title |
A framework to operationalize and automate the data integration lifecycle |
| spellingShingle |
A framework to operationalize and automate the data integration lifecycle Flores Herrera, Javier de Jesús|||0000-0002-2998-9962 Data Integration Data Discovery Knowledge Graphs Data Wrangling 004 - Informàtica Àrees temàtiques de la UPC::Informàtica |
| title_short |
A framework to operationalize and automate the data integration lifecycle |
| title_full |
A framework to operationalize and automate the data integration lifecycle |
| title_fullStr |
A framework to operationalize and automate the data integration lifecycle |
| title_full_unstemmed |
A framework to operationalize and automate the data integration lifecycle |
| title_sort |
A framework to operationalize and automate the data integration lifecycle |
| dc.creator.none.fl_str_mv |
Flores Herrera, Javier de Jesús|||0000-0002-2998-9962 |
| author |
Flores Herrera, Javier de Jesús|||0000-0002-2998-9962 |
| author_facet |
Flores Herrera, Javier de Jesús|||0000-0002-2998-9962 |
| author_role |
author |
| dc.subject.none.fl_str_mv |
Data Integration Data Discovery Knowledge Graphs Data Wrangling 004 - Informàtica Àrees temàtiques de la UPC::Informàtica |
| topic |
Data Integration Data Discovery Knowledge Graphs Data Wrangling 004 - Informàtica Àrees temàtiques de la UPC::Informàtica |
| description |
(English) Data plays a key role in today’s world. Many organizations collect and store massive amounts of data from many different data sources. As a result, these data collections show a diversity in structure and semantics that grows as the data sources expand and evolve. These factors challenge traditional data management methods, which depend on fixed structures and stable conditions. There is a mismatch between old assumptions and new realities, where it is not enough to just collect data and run conventional tools. Instead, we must rethink how we integrate data to support high variety, handle large-scale collections, and accommodate new available data. This PhD thesis proposes innovative and advanced techniques to support and automate the data integration lifecycle. First, we describe how to represent and standardize data sources using graph-based schemas. These schemas provide a solid foundation for all steps of the data integration lifecycle. Next, we introduce an integration method that leverages graph-based schemas to add new data incrementally without disrupting existing integration structures. This approach ensures that data integration remains flexible and scalable as organizations grow. We also help users find the right datasets to integrate. By focusing on data discovery, we reduce the time spent exploring irrelevant data sources and suggest relevant ones for integration. To this end, we focus first on facilitating the discovery of joinable attributes among datasets. We propose a new qualitative metric and use data profiles and learning models to decide which attributes are worth joining. To further enhance data discovery, we introduce contextual pre-filtering. Using data profiles and graph-based schemas, we can focus on promising datasets before applying data discovery tools. This pre-filtering step not only boosts the accuracy of existing data discovery tools but also optimizes their performance by narrowing the search space. In summary, this thesis helps bridge the gap between conventional data methods and modern, diverse data ecosystems. The results contribute to the field of data integration by offering scalable and automated solutions that match the changing needs of data integration today. |
| publishDate |
2025 |
| dc.date.none.fl_str_mv |
2025 2025-06-16 2025 2025-09-23 |
| dc.type.none.fl_str_mv |
doctoral thesis http://purl.org/coar/resource_type/c_db06 VoR http://purl.org/coar/version/c_970fb48d4fbd8a85 |
| dc.type.openaire.fl_str_mv |
info:eu-repo/semantics/doctoralThesis |
| format |
doctoralThesis |
| dc.identifier.none.fl_str_mv |
https://hdl.handle.net/2117/442278 https://dx.doi.org/10.5821/dissertation-2117-442278 |
| url |
https://hdl.handle.net/2117/442278 https://dx.doi.org/10.5821/dissertation-2117-442278 |
| dc.language.none.fl_str_mv |
Inglés eng |
| language_invalid_str_mv |
Inglés |
| language |
eng |
| dc.rights.none.fl_str_mv |
open access http://purl.org/coar/access_right/c_abf2 |
| dc.rights.openaire.fl_str_mv |
info:eu-repo/semantics/openAccess |
| rights_invalid_str_mv |
open access http://purl.org/coar/access_right/c_abf2 |
| eu_rights_str_mv |
openAccess |
| dc.format.none.fl_str_mv |
application/pdf |
| dc.publisher.none.fl_str_mv |
Universitat Politècnica de Catalunya |
| publisher.none.fl_str_mv |
Universitat Politècnica de Catalunya |
| dc.source.none.fl_str_mv |
reponame:UPCommons. Portal del coneixement obert de la UPC instname:Universitat Politècnica de Catalunya (UPC) |
| instname_str |
Universitat Politècnica de Catalunya (UPC) |
| reponame_str |
UPCommons. Portal del coneixement obert de la UPC |
| collection |
UPCommons. Portal del coneixement obert de la UPC |
| repository.name.fl_str_mv |
|
| repository.mail.fl_str_mv |
|
| _version_ |
1869416537517981696 |
| spelling |
A framework to operationalize and automate the data integration lifecycleFlores Herrera, Javier de Jesús|||0000-0002-2998-9962Data IntegrationData DiscoveryKnowledge GraphsData Wrangling004 - InformàticaÀrees temàtiques de la UPC::Informàtica(English) Data plays a key role in today’s world. Many organizations collect and store massive amounts of data from many different data sources. As a result, these data collections show a diversity in structure and semantics that grows as the data sources expand and evolve. These factors challenge traditional data management methods, which depend on fixed structures and stable conditions. There is a mismatch between old assumptions and new realities, where it is not enough to just collect data and run conventional tools. Instead, we must rethink how we integrate data to support high variety, handle large-scale collections, and accommodate new available data. This PhD thesis proposes innovative and advanced techniques to support and automate the data integration lifecycle. First, we describe how to represent and standardize data sources using graph-based schemas. These schemas provide a solid foundation for all steps of the data integration lifecycle. Next, we introduce an integration method that leverages graph-based schemas to add new data incrementally without disrupting existing integration structures. This approach ensures that data integration remains flexible and scalable as organizations grow. We also help users find the right datasets to integrate. By focusing on data discovery, we reduce the time spent exploring irrelevant data sources and suggest relevant ones for integration. To this end, we focus first on facilitating the discovery of joinable attributes among datasets. We propose a new qualitative metric and use data profiles and learning models to decide which attributes are worth joining. To further enhance data discovery, we introduce contextual pre-filtering. Using data profiles and graph-based schemas, we can focus on promising datasets before applying data discovery tools. This pre-filtering step not only boosts the accuracy of existing data discovery tools but also optimizes their performance by narrowing the search space. In summary, this thesis helps bridge the gap between conventional data methods and modern, diverse data ecosystems. The results contribute to the field of data integration by offering scalable and automated solutions that match the changing needs of data integration today.(Català) Les dades tenen un paper fonamental en el món actual. Moltes organitzacions recopilen i emmagatzemen grans volums de dades procedents de diverses fonts. Aquestes fonts poden variar tant en l’estructura com en la modelització de conceptes i van creixent i evolucionant a mesura que s’hi afegeixen noves fonts de dades. Això posa a prova els mètodes clàssics de gestió de dades, que depenen d’estructures fixes i condicions estables. Avui, ja no n’hi ha prou de reunir dades i emprar eines convencionals. Cal replantejar la manera d’integrar les dades per gestionar-ne la gran varietat, tractar grans volums i incorporar noves fonts a mesura que s’integren. Aquesta tesi proposa tècniques per automatitzar el cicle de vida de la integració de dades. Primer, mostrem com representar i estandarditzar les fonts mitjançant esquemes basats en graf. Aquests esquemes serveixen de fonament sòlid per a cada pas de la integració. Tot seguit, presentem un mètode que aprofita aquests esquemes per afegir noves fonts de manera incremental sense alterar les estructures existents, tot mantenint flexibilitat i escalabilitat a mesura que les organitzacions creixen. També fem més àgil la cerca de conjunts de dades que valgui la pena integrar. En centrar-nos en el descobriment de dades, reduïm el temps destinat a explorar fonts irrellevants i proposem les més adequades. Per fer-ho, introduïm una mètrica qualitativa i fem servir perfils de dades i models d’aprenentatge per decidir quins atributs cal unir. A més, incorporem un prefiltrat contextual que detecta els conjunts de dades més prometedors abans d’aplicar eines de descobriment, cosa que millora la precisió i redueix la càrrega computacional. En resum, aquesta tesi escurça la distància entre els mètodes tradicionals i els entorns moderns de dades. Ofereix solucions escalables i automatitzades que s’adapten a les necessitats canviants de la integració de dades.(Español) Los datos desempeñan un papel fundamental en el mundo actual. Muchas organizaciones recopilan y almacenan grandes volúmenes de datos desde diversas fuentes. Estas fuentes pueden variar en estructura y modelado de conceptos que van creciendo y evolucionando a medida que más fuentes de datos son integradas. Esto pone a prueba los métodos clásicos de gestión de datos, que dependen de estructuras fijas y condiciones estables. Hoy en día, no basta con reunir datos y usar herramientas convencionales. En su lugar, debemos replantearnos cómo integrar datos para manejar una alta variedad, gestionar grandes volúmenes y acomodar nuevas fuentes a medida que se integran. Esta tesis propone técnicas para automatizar el ciclo de vida de la integración de datos. Primero, mostramos cómo representar y estandarizar las fuentes con esquemas basados en grafos. Estos esquemas sirven de base sólida para cada paso de la integración. Luego, presentamos un método que emplea dichos esquemas para añadir nuevas fuentes de forma incremental sin afectar las estructuras existentes, manteniendo flexibilidad y escalabilidad a medida que las organizaciones crecen. También agilizamos la búsqueda de conjuntos de datos que valga la pena integrar. Al centrarnos en el descubrimiento de datos, reducimos el tiempo dedicado a explorar fuentes irrelevantes y sugerimos las más adecuadas. Para ello, introducimos una métrica cualitativa y usamos perfiles de datos y modelos de aprendizaje para decidir qué atributos se deben unir. Además, aportamos un prefiltrado contextual que identifica los conjuntos de datos más prometedores antes de aplicar herramientas de descubrimiento, lo que mejora la precisión y reduce la carga computacional. En resumen, esta tesis acorta la brecha entre los métodos tradicionales y los entornos modernos de datos. Ofrece soluciones escalables y automatizadas que se adaptan a las cambiantes necesidades de la integración de datos.Universitat Politècnica de Catalunya20252025-06-1620252025-09-23doctoral thesishttp://purl.org/coar/resource_type/c_db06VoRhttp://purl.org/coar/version/c_970fb48d4fbd8a85info:eu-repo/semantics/doctoralThesisapplication/pdfhttps://hdl.handle.net/2117/442278https://dx.doi.org/10.5821/dissertation-2117-442278reponame:UPCommons. Portal del coneixement obert de la UPCinstname:Universitat Politècnica de Catalunya (UPC)Inglésengopen accesshttp://purl.org/coar/access_right/c_abf2info:eu-repo/semantics/openAccessoai:upcommons.upc.edu:2117/4422782026-05-27T15:37:01Z |
| score |
15,812455 |