Reinforcement Learning for Value Alignment

[eng] As autonomous agents become increasingly sophisticated and we allow them to perform more complex tasks, it is of utmost importance to guarantee that they will act in alignment with human values. This problem has received in the AI literature the name of the value alignment problem. Current app...

ver descrição completa

Detalhes bibliográficos
Autor: Rodríguez Soto, Manel
Formato: tesis doctoral
Estado:Versión publicada
Fecha de publicación:2023
País:España
Recursos:Universidad de Barcelona
Repositorio:Dipòsit Digital de la UB
OAI Identifier:oai:diposit.ub.edu:2445/202126
Acesso em linha:https://hdl.handle.net/2445/202126
http://hdl.handle.net/10803/688998
Access Level:acceso abierto
Palavra-chave:Intel·ligència artificial
Aprenentatge automàtic
Aprenentatge per reforç (Intel·ligència artificial)
Sistemes multiagent
Valors (Filosofia)
Artificial intelligence
Machine learning
Reinforcement learning
Multiagent systems
Values
id ES_31b059ca5e37af4edaa544be43e2cb3a
oai_identifier_str oai:diposit.ub.edu:2445/202126
network_acronym_str ES
network_name_str España
repository_id_str
dc.title.none.fl_str_mv Reinforcement Learning for Value Alignment
title Reinforcement Learning for Value Alignment
spellingShingle Reinforcement Learning for Value Alignment
Rodríguez Soto, Manel
Intel·ligència artificial
Aprenentatge automàtic
Aprenentatge per reforç (Intel·ligència artificial)
Sistemes multiagent
Valors (Filosofia)
Artificial intelligence
Machine learning
Reinforcement learning
Multiagent systems
Values
title_short Reinforcement Learning for Value Alignment
title_full Reinforcement Learning for Value Alignment
title_fullStr Reinforcement Learning for Value Alignment
title_full_unstemmed Reinforcement Learning for Value Alignment
title_sort Reinforcement Learning for Value Alignment
dc.creator.none.fl_str_mv Rodríguez Soto, Manel
author Rodríguez Soto, Manel
author_facet Rodríguez Soto, Manel
author_role author
dc.contributor.none.fl_str_mv López Sánchez, Maite
Rodríguez-Aguilar, Juan A. (Juan Antonio)
Universitat de Barcelona. Facultat de Matemàtiques
dc.subject.none.fl_str_mv Intel·ligència artificial
Aprenentatge automàtic
Aprenentatge per reforç (Intel·ligència artificial)
Sistemes multiagent
Valors (Filosofia)
Artificial intelligence
Machine learning
Reinforcement learning
Multiagent systems
Values
topic Intel·ligència artificial
Aprenentatge automàtic
Aprenentatge per reforç (Intel·ligència artificial)
Sistemes multiagent
Valors (Filosofia)
Artificial intelligence
Machine learning
Reinforcement learning
Multiagent systems
Values
description [eng] As autonomous agents become increasingly sophisticated and we allow them to perform more complex tasks, it is of utmost importance to guarantee that they will act in alignment with human values. This problem has received in the AI literature the name of the value alignment problem. Current approaches apply reinforcement learning to align agents with values due to its recent successes at solving complex sequential decision-making problems. However, they follow an agent-centric approach by expecting that the agent applies the reinforcement learning algorithm correctly to learn an ethical behaviour, without formal guarantees that the learnt ethical behaviour will be ethical. This thesis proposes a novel environment-designer approach for solving the value alignment problem with theoretical guarantees. Our proposed environment-designer approach advances the state of the art with a process for designing ethical environments wherein it is in the agent's best interest to learn ethical behaviours. Our process specifies the ethical knowledge of a moral value in terms that can be used in a reinforcement learning context. Next, our process embeds this knowledge in the agent's learning environment to design an ethical learning environment. The resulting ethical environment incentivises the agent to learn an ethical behaviour while pursuing its own objective. We further contribute to the state of the art by providing a novel algorithm that, following our ethical environment design process, is formally guaranteed to create ethical environments. In other words, this algorithm guarantees that it is in the agent's best interest to learn value- aligned behaviours. We illustrate our algorithm by applying it in a case study environment wherein the agent is expected to learn to behave in alignment with the moral value of respect. In it, a conversational agent is in charge of doing surveys, and we expect it to ask the users questions respectfully while trying to get as much information as possible. In the designed ethical environment, results confirm our theoretical results: the agent learns an ethical behaviour while pursuing its individual objective.
publishDate 2023
dc.date.none.fl_str_mv 2023
dc.type.none.fl_str_mv info:eu-repo/semantics/doctoralThesis
info:eu-repo/semantics/publishedVersion
format doctoralThesis
status_str publishedVersion
dc.identifier.none.fl_str_mv https://hdl.handle.net/2445/202126
http://hdl.handle.net/10803/688998
url https://hdl.handle.net/2445/202126
http://hdl.handle.net/10803/688998
dc.language.none.fl_str_mv Inglés
language_invalid_str_mv Inglés
dc.rights.none.fl_str_mv cc by (c) Rodríguez Soto, Manel, 2023
http://creativecommons.org/licenses/by/3.0/es/
info:eu-repo/semantics/openAccess
rights_invalid_str_mv cc by (c) Rodríguez Soto, Manel, 2023
http://creativecommons.org/licenses/by/3.0/es/
eu_rights_str_mv openAccess
dc.format.none.fl_str_mv application/pdf
dc.publisher.none.fl_str_mv Universitat de Barcelona
publisher.none.fl_str_mv Universitat de Barcelona
dc.source.none.fl_str_mv Tesis Doctorals - Facultat - Matemàtiques
reponame:Dipòsit Digital de la UB
instname:Universidad de Barcelona
instname_str Universidad de Barcelona
reponame_str Dipòsit Digital de la UB
collection Dipòsit Digital de la UB
repository.name.fl_str_mv
repository.mail.fl_str_mv
_version_ 1869405627317485568
spelling Reinforcement Learning for Value AlignmentRodríguez Soto, ManelIntel·ligència artificialAprenentatge automàticAprenentatge per reforç (Intel·ligència artificial)Sistemes multiagentValors (Filosofia)Artificial intelligenceMachine learningReinforcement learningMultiagent systemsValues[eng] As autonomous agents become increasingly sophisticated and we allow them to perform more complex tasks, it is of utmost importance to guarantee that they will act in alignment with human values. This problem has received in the AI literature the name of the value alignment problem. Current approaches apply reinforcement learning to align agents with values due to its recent successes at solving complex sequential decision-making problems. However, they follow an agent-centric approach by expecting that the agent applies the reinforcement learning algorithm correctly to learn an ethical behaviour, without formal guarantees that the learnt ethical behaviour will be ethical. This thesis proposes a novel environment-designer approach for solving the value alignment problem with theoretical guarantees. Our proposed environment-designer approach advances the state of the art with a process for designing ethical environments wherein it is in the agent's best interest to learn ethical behaviours. Our process specifies the ethical knowledge of a moral value in terms that can be used in a reinforcement learning context. Next, our process embeds this knowledge in the agent's learning environment to design an ethical learning environment. The resulting ethical environment incentivises the agent to learn an ethical behaviour while pursuing its own objective. We further contribute to the state of the art by providing a novel algorithm that, following our ethical environment design process, is formally guaranteed to create ethical environments. In other words, this algorithm guarantees that it is in the agent's best interest to learn value- aligned behaviours. We illustrate our algorithm by applying it in a case study environment wherein the agent is expected to learn to behave in alignment with the moral value of respect. In it, a conversational agent is in charge of doing surveys, and we expect it to ask the users questions respectfully while trying to get as much information as possible. In the designed ethical environment, results confirm our theoretical results: the agent learns an ethical behaviour while pursuing its individual objective.[cat] A mesura que els agents autònoms es tornen cada cop més sofisticats i els permetem realitzar tasques més complexes, és de la màxima importància garantir que actuaran d'acord amb els valors humans. Aquest problema ha rebut a la literatura d'IA el nom del problema d'alineació de valors. Els enfocaments actuals apliquen aprenentatge per reforç per alinear els agents amb els valors a causa dels seus èxits recents a l'hora de resoldre problemes complexos de presa de decisions seqüencials. Tanmateix, segueixen un enfocament centrat en l'agent en esperar que l'agent apliqui correctament l'algorisme d'aprenentatge de reforç per aprendre un comportament ètic, sense garanties formals que el comportament ètic après serà ètic. Aquesta tesi proposa un nou enfocament de dissenyador d'entorn per resoldre el problema d'alineació de valors amb garanties teòriques. El nostre enfocament de disseny d'entorns proposat avança l'estat de l'art amb un procés per dissenyar entorns ètics en què és del millor interès de l'agent aprendre comportaments ètics. El nostre procés especifica el coneixement ètic d'un valor moral en termes que es poden utilitzar en un context d'aprenentatge de reforç. A continuació, el nostre procés incorpora aquest coneixement a l'entorn d'aprenentatge de l'agent per dissenyar un entorn d'aprenentatge ètic. L'entorn ètic resultant incentiva l'agent a aprendre un comportament ètic mentre persegueix el seu propi objectiu. A més, contribuïm a l'estat de l'art proporcionant un algorisme nou que, seguint el nostre procés de disseny d'entorns ètics, està garantit formalment per crear entorns ètics. En altres paraules, aquest algorisme garanteix que és del millor interès de l'agent aprendre comportaments alineats amb valors. Il·lustrem el nostre algorisme aplicant-lo en un estudi de cas on s'espera que l'agent aprengui a comportar-se d'acord amb el valor moral del respecte. En ell, un agent de conversa s'encarrega de fer enquestes, i esperem que faci preguntes als usuaris amb respecte tot intentant obtenir la màxima informació possible. En l'entorn ètic dissenyat, els resultats confirmen els nostres resultats teòrics: l'agent aprèn un comportament ètic mentre persegueix el seu objectiu individual.Universitat de BarcelonaLópez Sánchez, MaiteRodríguez-Aguilar, Juan A. (Juan Antonio)Universitat de Barcelona. Facultat de Matemàtiques2023info:eu-repo/semantics/doctoralThesisinfo:eu-repo/semantics/publishedVersionapplication/pdfhttps://hdl.handle.net/2445/202126http://hdl.handle.net/10803/688998Tesis Doctorals - Facultat - Matemàtiquesreponame:Dipòsit Digital de la UBinstname:Universidad de BarcelonaIngléscc by (c) Rodríguez Soto, Manel, 2023http://creativecommons.org/licenses/by/3.0/es/info:eu-repo/semantics/openAccessoai:diposit.ub.edu:2445/2021262026-05-27T06:46:51Z
score 15,301629