Multi-task reinforcement learning in reproducing Kernel Hilbert spaces via cross-learning

Cerviño, Juan - Bazerque, Juan Andrés - Calvo-Fullana, Miguel - Ribeiro, Alejandro

Resumen:

Reinforcement learning is a framework to optimize an agent’s policy using rewards that are revealed by the system as a response to an action. In its standard form, reinforcement learning involves a single agent that uses its policy to accomplish a specific task. These methods require large amounts of reward samples to achieve good performance, and may not generalize well when the task is modified, even if the new task is related. In this paper we are interested in a collaborative scheme in which multiple policies are optimized jointly. To this end, we we introduce cross-learning , in which policies are trained for related tasks in separate environments, and they are constrained to be close to one another. Two properties make our new approach attractive: (i) it produces a multi-task central policy that can be used as a starting point to adapt quickly to one of the tasks trained for, and (ii) as in meta-learning , it adapts to environments related but different to those seen during training. We focus on policies belonging to reproducing kernel Hilbert spaces for which we bound the distance between the task-specific policies and the cross-learned policy. To solve the resulting optimization problem, we resort to a projected policy gradient algorithm and prove that it converges to a near-optimal solution with high probability. We evaluate our methodology with a navigation example in which an agent moves through environments with obstacles of multiple shapes and avoids obstacles not trained for.

Detalles Bibliográficos
2021
ARL DCIST CRA W911NF-17-2-0181
Intel Science and Technology Center for Wireless Autonomous Systems
Task analysis
Training data
Navigation
Convergence
Reinforcement learning
Optimization
Multi-task learning
Meta-learning
Inglés
Universidad de la República
COLIBRI
https://hdl.handle.net/20.500.12008/51580
Acceso abierto
Licencia Creative Commons Atribución (CC - By 4.0)
_version_ 1875692335161933824
author Cerviño, Juan
author2 Bazerque, Juan Andrés
Calvo-Fullana, Miguel
Ribeiro, Alejandro
author2_role author
author
author
author_facet Cerviño, Juan
Bazerque, Juan Andrés
Calvo-Fullana, Miguel
Ribeiro, Alejandro
author_role author
bitstream.checksum.fl_str_mv 6429389a7df7277b72b7924fdc7d47a9
a0ebbeafb9d2ec7cbb19d7137ebc392c
aa4ded3991caf203ada54d801dbdfafc
58cb336ce230a47d2f88ad02838a665f
4ddc3993ffc983074c0d51c82009744b
bitstream.checksumAlgorithm.fl_str_mv MD5
MD5
MD5
MD5
MD5
bitstream.url.fl_str_mv http://localhost:8080/xmlui/bitstream/20.500.12008/51580/5/license.txt
http://localhost:8080/xmlui/bitstream/20.500.12008/51580/2/license_url
http://localhost:8080/xmlui/bitstream/20.500.12008/51580/3/license_text
http://localhost:8080/xmlui/bitstream/20.500.12008/51580/4/license_rdf
http://localhost:8080/xmlui/bitstream/20.500.12008/51580/1/CBCR21.pdf
collection COLIBRI
dc.contributor.filiacion.none.fl_str_mv Cerviño Juan, University of Pennsylvania, Philadelphia, USA
Bazerque Juan Andrés, Universidad de la República (Uruguay). Facultad de Ingeniería.
Calvo-Fullana Miguel, Massachusetts Institute of Technology, Boston, USA
Ribeiro Alejandro, University of Pennsylvania, Philadelphia, USA
dc.creator.none.fl_str_mv Cerviño, Juan
Bazerque, Juan Andrés
Calvo-Fullana, Miguel
Ribeiro, Alejandro
dc.date.accessioned.none.fl_str_mv 2025-09-11T18:01:35Z
dc.date.available.none.fl_str_mv 2025-09-11T18:01:35Z
dc.date.issued.none.fl_str_mv 2021
dc.description.abstract.none.fl_txt_mv Reinforcement learning is a framework to optimize an agent’s policy using rewards that are revealed by the system as a response to an action. In its standard form, reinforcement learning involves a single agent that uses its policy to accomplish a specific task. These methods require large amounts of reward samples to achieve good performance, and may not generalize well when the task is modified, even if the new task is related. In this paper we are interested in a collaborative scheme in which multiple policies are optimized jointly. To this end, we we introduce cross-learning , in which policies are trained for related tasks in separate environments, and they are constrained to be close to one another. Two properties make our new approach attractive: (i) it produces a multi-task central policy that can be used as a starting point to adapt quickly to one of the tasks trained for, and (ii) as in meta-learning , it adapts to environments related but different to those seen during training. We focus on policies belonging to reproducing kernel Hilbert spaces for which we bound the distance between the task-specific policies and the cross-learned policy. To solve the resulting optimization problem, we resort to a projected policy gradient algorithm and prove that it converges to a near-optimal solution with high probability. We evaluate our methodology with a navigation example in which an agent moves through environments with obstacles of multiple shapes and avoids obstacles not trained for.
dc.description.sponsorship.none.fl_txt_mv ARL DCIST CRA W911NF-17-2-0181
Intel Science and Technology Center for Wireless Autonomous Systems
dc.format.extent.es.fl_str_mv 16 p.
dc.format.mimetype.es.fl_str_mv application/pdf
dc.identifier.citation.es.fl_str_mv Cerviño, J., Bazerque, J., Calvo-Fullana, M. y otros. Multi-task reinforcement learning in reproducing Kernel Hilbert spaces via cross-learning [Preprint]. Publicado en: IEEE Transactions on Signal Processing, vol. 69, oct. 2021, pp. 5947-5962. DOI: 10.1109/TSP.2021.3122303.
dc.identifier.uri.none.fl_str_mv https://hdl.handle.net/20.500.12008/51580
dc.language.iso.none.fl_str_mv en
eng
dc.rights.license.none.fl_str_mv Licencia Creative Commons Atribución (CC - By 4.0)
dc.rights.none.fl_str_mv info:eu-repo/semantics/openAccess
dc.source.none.fl_str_mv reponame:COLIBRI
instname:Universidad de la República
instacron:Universidad de la República
dc.subject.es.fl_str_mv Task analysis
Training data
Navigation
Convergence
Reinforcement learning
Optimization
Multi-task learning
Meta-learning
dc.title.none.fl_str_mv Multi-task reinforcement learning in reproducing Kernel Hilbert spaces via cross-learning
dc.type.es.fl_str_mv Preprint
dc.type.none.fl_str_mv info:eu-repo/semantics/preprint
dc.type.version.none.fl_str_mv info:eu-repo/semantics/submittedVersion
description Reinforcement learning is a framework to optimize an agent’s policy using rewards that are revealed by the system as a response to an action. In its standard form, reinforcement learning involves a single agent that uses its policy to accomplish a specific task. These methods require large amounts of reward samples to achieve good performance, and may not generalize well when the task is modified, even if the new task is related. In this paper we are interested in a collaborative scheme in which multiple policies are optimized jointly. To this end, we we introduce cross-learning , in which policies are trained for related tasks in separate environments, and they are constrained to be close to one another. Two properties make our new approach attractive: (i) it produces a multi-task central policy that can be used as a starting point to adapt quickly to one of the tasks trained for, and (ii) as in meta-learning , it adapts to environments related but different to those seen during training. We focus on policies belonging to reproducing kernel Hilbert spaces for which we bound the distance between the task-specific policies and the cross-learned policy. To solve the resulting optimization problem, we resort to a projected policy gradient algorithm and prove that it converges to a near-optimal solution with high probability. We evaluate our methodology with a navigation example in which an agent moves through environments with obstacles of multiple shapes and avoids obstacles not trained for.
eu_rights_str_mv openAccess
format preprint
id COLIBRI_e2a0d0f3249d806d20b2ba43ad245e3e
identifier_str_mv Cerviño, J., Bazerque, J., Calvo-Fullana, M. y otros. Multi-task reinforcement learning in reproducing Kernel Hilbert spaces via cross-learning [Preprint]. Publicado en: IEEE Transactions on Signal Processing, vol. 69, oct. 2021, pp. 5947-5962. DOI: 10.1109/TSP.2021.3122303.
instacron_str Universidad de la República
institution Universidad de la República
instname_str Universidad de la República
language eng
language_invalid_str_mv en
network_acronym_str COLIBRI
network_name_str COLIBRI
oai_identifier_str oai:colibri.udelar.edu.uy:20.500.12008/51580
publishDate 2021
reponame_str COLIBRI
repository.mail.fl_str_mv karina.camps@seciu.edu.uy
repository.name.fl_str_mv COLIBRI - Universidad de la República
repository_id_str 4771
rights_invalid_str_mv Licencia Creative Commons Atribución (CC - By 4.0)
spelling Cerviño Juan, University of Pennsylvania, Philadelphia, USABazerque Juan Andrés, Universidad de la República (Uruguay). Facultad de Ingeniería.Calvo-Fullana Miguel, Massachusetts Institute of Technology, Boston, USARibeiro Alejandro, University of Pennsylvania, Philadelphia, USA2025-09-11T18:01:35Z2025-09-11T18:01:35Z2021Cerviño, J., Bazerque, J., Calvo-Fullana, M. y otros. Multi-task reinforcement learning in reproducing Kernel Hilbert spaces via cross-learning [Preprint]. Publicado en: IEEE Transactions on Signal Processing, vol. 69, oct. 2021, pp. 5947-5962. DOI: 10.1109/TSP.2021.3122303.https://hdl.handle.net/20.500.12008/51580Reinforcement learning is a framework to optimize an agent’s policy using rewards that are revealed by the system as a response to an action. In its standard form, reinforcement learning involves a single agent that uses its policy to accomplish a specific task. These methods require large amounts of reward samples to achieve good performance, and may not generalize well when the task is modified, even if the new task is related. In this paper we are interested in a collaborative scheme in which multiple policies are optimized jointly. To this end, we we introduce cross-learning , in which policies are trained for related tasks in separate environments, and they are constrained to be close to one another. Two properties make our new approach attractive: (i) it produces a multi-task central policy that can be used as a starting point to adapt quickly to one of the tasks trained for, and (ii) as in meta-learning , it adapts to environments related but different to those seen during training. We focus on policies belonging to reproducing kernel Hilbert spaces for which we bound the distance between the task-specific policies and the cross-learned policy. To solve the resulting optimization problem, we resort to a projected policy gradient algorithm and prove that it converges to a near-optimal solution with high probability. We evaluate our methodology with a navigation example in which an agent moves through environments with obstacles of multiple shapes and avoids obstacles not trained for.Submitted by Ribeiro Jorge (jribeiro@fing.edu.uy) on 2025-09-10T18:05:06Z No. of bitstreams: 2 license_rdf: 24942 bytes, checksum: 58cb336ce230a47d2f88ad02838a665f (MD5) CBCR21.pdf: 897474 bytes, checksum: 4ddc3993ffc983074c0d51c82009744b (MD5)Approved for entry into archive by Machado Jimena (jmachado@fing.edu.uy) on 2025-09-11T17:43:38Z (GMT) No. of bitstreams: 2 license_rdf: 24942 bytes, checksum: 58cb336ce230a47d2f88ad02838a665f (MD5) CBCR21.pdf: 897474 bytes, checksum: 4ddc3993ffc983074c0d51c82009744b (MD5)Made available in DSpace by Camps Karina (karina.camps@seciu.edu.uy) on 2025-09-11T18:01:35Z (GMT). No. of bitstreams: 2 license_rdf: 24942 bytes, checksum: 58cb336ce230a47d2f88ad02838a665f (MD5) CBCR21.pdf: 897474 bytes, checksum: 4ddc3993ffc983074c0d51c82009744b (MD5) Previous issue date: 2021ARL DCIST CRA W911NF-17-2-0181Intel Science and Technology Center for Wireless Autonomous Systems16 p.application/pdfenengLas obras depositadas en el Repositorio se rigen por la Ordenanza de los Derechos de la Propiedad Intelectual de la Universidad de la República.(Res. Nº 91 de C.D.C. de 8/III/1994 – D.O. 7/IV/1994) y por la Ordenanza del Repositorio Abierto de la Universidad de la República (Res. Nº 16 de C.D.C. de 07/10/2014)info:eu-repo/semantics/openAccessLicencia Creative Commons Atribución (CC - By 4.0)Task analysisTraining dataNavigationConvergenceReinforcement learningOptimizationMulti-task learningMeta-learningMulti-task reinforcement learning in reproducing Kernel Hilbert spaces via cross-learningPreprintinfo:eu-repo/semantics/preprintinfo:eu-repo/semantics/submittedVersionreponame:COLIBRIinstname:Universidad de la Repúblicainstacron:Universidad de la RepúblicaCerviño, JuanBazerque, Juan AndrésCalvo-Fullana, MiguelRibeiro, AlejandroSistemas y ControlLICENSElicense.txtlicense.txttext/plain; charset=utf-84267http://localhost:8080/xmlui/bitstream/20.500.12008/51580/5/license.txt6429389a7df7277b72b7924fdc7d47a9MD55CC-LICENSElicense_urllicense_urltext/plain; charset=utf-844http://localhost:8080/xmlui/bitstream/20.500.12008/51580/2/license_urla0ebbeafb9d2ec7cbb19d7137ebc392cMD52license_textlicense_texttext/html; charset=utf-826818http://localhost:8080/xmlui/bitstream/20.500.12008/51580/3/license_textaa4ded3991caf203ada54d801dbdfafcMD53license_rdflicense_rdfapplication/rdf+xml; charset=utf-824942http://localhost:8080/xmlui/bitstream/20.500.12008/51580/4/license_rdf58cb336ce230a47d2f88ad02838a665fMD54ORIGINALCBCR21.pdfCBCR21.pdfapplication/pdf897474http://localhost:8080/xmlui/bitstream/20.500.12008/51580/1/CBCR21.pdf4ddc3993ffc983074c0d51c82009744bMD5120.500.12008/515802025-09-11 15:01:35.959oai:colibri.udelar.edu.uy:20.500.12008/51580VGVybWlub3MgeSBjb25kaWNpb25lcyByZWxhdGl2YXMgYWwgZGVwb3NpdG8gZGUgb2JyYXMKCgpMYXMgb2JyYXMgZGVwb3NpdGFkYXMgZW4gZWwgUmVwb3NpdG9yaW8gc2UgcmlnZW4gcG9yIGxhIE9yZGVuYW56YSBkZSBsb3MgRGVyZWNob3MgZGUgbGEgUHJvcGllZGFkIEludGVsZWN0dWFsICBkZSBsYSBVbml2ZXJzaWRhZCBEZSBMYSBSZXDDumJsaWNhLiAoUmVzLiBOwrogOTEgZGUgQy5ELkMuIGRlIDgvSUlJLzE5OTQg4oCTIEQuTy4gNy9JVi8xOTk0KSB5ICBwb3IgbGEgT3JkZW5hbnphIGRlbCBSZXBvc2l0b3JpbyBBYmllcnRvIGRlIGxhIFVuaXZlcnNpZGFkIGRlIGxhIFJlcMO6YmxpY2EgKFJlcy4gTsK6IDE2IGRlIEMuRC5DLiBkZSAwNy8xMC8yMDE0KS4gCgpBY2VwdGFuZG8gZWwgYXV0b3IgZXN0b3MgdMOpcm1pbm9zIHkgY29uZGljaW9uZXMgZGUgZGVww7NzaXRvIGVuIENPTElCUkksIGxhIFVuaXZlcnNpZGFkIGRlIFJlcMO6YmxpY2EgcHJvY2VkZXLDoSBhOiAgCgphKSBhcmNoaXZhciBtw6FzIGRlIHVuYSBjb3BpYSBkZSBsYSBvYnJhIGVuIGxvcyBzZXJ2aWRvcmVzIGRlIGxhIFVuaXZlcnNpZGFkIGEgbG9zIGVmZWN0b3MgZGUgZ2FyYW50aXphciBhY2Nlc28sIHNlZ3VyaWRhZCB5IHByZXNlcnZhY2nDs24KYikgY29udmVydGlyIGxhIG9icmEgYSBvdHJvcyBmb3JtYXRvcyBzaSBmdWVyYSBuZWNlc2FyaW8gIHBhcmEgZmFjaWxpdGFyIHN1IHByZXNlcnZhY2nDs24geSBhY2Nlc2liaWxpZGFkIHNpbiBhbHRlcmFyIHN1IGNvbnRlbmlkby4KYykgcmVhbGl6YXIgbGEgY29tdW5pY2FjacOzbiBww7pibGljYSB5IGRpc3BvbmVyIGVsIGFjY2VzbyBsaWJyZSB5IGdyYXR1aXRvIGEgdHJhdsOpcyBkZSBJbnRlcm5ldCBtZWRpYW50ZSBsYSBwdWJsaWNhY2nDs24gZGUgbGEgb2JyYSBiYWpvIGxhIGxpY2VuY2lhIENyZWF0aXZlIENvbW1vbnMgc2VsZWNjaW9uYWRhIHBvciBlbCBwcm9waW8gYXV0b3IuCgoKRW4gY2FzbyBxdWUgZWwgYXV0b3IgaGF5YSBkaWZ1bmRpZG8geSBkYWRvIGEgcHVibGljaWRhZCBhIGxhIG9icmEgZW4gZm9ybWEgcHJldmlhLCAgcG9kcsOhIHNvbGljaXRhciB1biBwZXLDrW9kbyBkZSBlbWJhcmdvIHNvYnJlIGxhIGRpc3BvbmliaWxpZGFkIHDDumJsaWNhIGRlIGxhIG1pc21hLCBlbCBjdWFsIGNvbWVuemFyw6EgYSBwYXJ0aXIgZGUgbGEgYWNlcHRhY2nDs24gZGUgZXN0ZSBkb2N1bWVudG8geSBoYXN0YSBsYSBmZWNoYSBxdWUgaW5kaXF1ZSAuCgpFbCBhdXRvciBhc2VndXJhIHF1ZSBsYSBvYnJhIG5vIGluZnJpZ2UgbmluZ8O6biBkZXJlY2hvIHNvYnJlIHRlcmNlcm9zLCB5YSBzZWEgZGUgcHJvcGllZGFkIGludGVsZWN0dWFsIG8gY3VhbHF1aWVyIG90cm8uCgpFbCBhdXRvciBnYXJhbnRpemEgcXVlIHNpIGVsIGRvY3VtZW50byBjb250aWVuZSBtYXRlcmlhbGVzIGRlIGxvcyBjdWFsZXMgbm8gdGllbmUgbG9zIGRlcmVjaG9zIGRlIGF1dG9yLCAgaGEgb2J0ZW5pZG8gZWwgcGVybWlzbyBkZWwgcHJvcGlldGFyaW8gZGUgbG9zIGRlcmVjaG9zIGRlIGF1dG9yLCB5IHF1ZSBlc2UgbWF0ZXJpYWwgY3V5b3MgZGVyZWNob3Mgc29uIGRlIHRlcmNlcm9zIGVzdMOhIGNsYXJhbWVudGUgaWRlbnRpZmljYWRvIHkgcmVjb25vY2lkbyBlbiBlbCB0ZXh0byBvIGNvbnRlbmlkbyBkZWwgZG9jdW1lbnRvIGRlcG9zaXRhZG8gZW4gZWwgUmVwb3NpdG9yaW8uCgpFbiBvYnJhcyBkZSBhdXRvcsOtYSBtw7psdGlwbGUgL3NlIHByZXN1bWUvIHF1ZSBlbCBhdXRvciBkZXBvc2l0YW50ZSBkZWNsYXJhIHF1ZSBoYSByZWNhYmFkbyBlbCBjb25zZW50aW1pZW50byBkZSB0b2RvcyBsb3MgYXV0b3JlcyBwYXJhIHB1YmxpY2FybGEgZW4gZWwgUmVwb3NpdG9yaW8sIHNpZW5kbyDDqXN0ZSBlbCDDum5pY28gcmVzcG9uc2FibGUgZnJlbnRlIGEgY3VhbHF1aWVyIHRpcG8gZGUgcmVjbGFtYWNpw7NuIGRlIGxvcyBvdHJvcyBjb2F1dG9yZXMuCgpFbCBhdXRvciBzZXLDoSByZXNwb25zYWJsZSBkZWwgY29udGVuaWRvIGRlIGxvcyBkb2N1bWVudG9zIHF1ZSBkZXBvc2l0YS4gTGEgVURFTEFSIG5vIHNlcsOhIHJlc3BvbnNhYmxlIHBvciBsYXMgZXZlbnR1YWxlcyB2aW9sYWNpb25lcyBhbCBkZXJlY2hvIGRlIHByb3BpZWRhZCBpbnRlbGVjdHVhbCBlbiBxdWUgcHVlZGEgaW5jdXJyaXIgZWwgYXV0b3IuCgpBbnRlIGN1YWxxdWllciBkZW51bmNpYSBkZSB2aW9sYWNpw7NuIGRlIGRlcmVjaG9zIGRlIHByb3BpZWRhZCBpbnRlbGVjdHVhbCwgbGEgVURFTEFSICBhZG9wdGFyw6EgdG9kYXMgbGFzIG1lZGlkYXMgbmVjZXNhcmlhcyBwYXJhIGV2aXRhciBsYSBjb250aW51YWNpw7NuIGRlIGRpY2hhIGluZnJhY2Npw7NuLCBsYXMgcXVlIHBvZHLDoW4gaW5jbHVpciBlbCByZXRpcm8gZGVsIGFjY2VzbyBhIGxvcyBjb250ZW5pZG9zIHkvbyBtZXRhZGF0b3MgZGVsIGRvY3VtZW50byByZXNwZWN0aXZvLgoKTGEgb2JyYSBzZSBwb25kcsOhIGEgZGlzcG9zaWNpw7NuIGRlbCBww7pibGljbyBhIHRyYXbDqXMgZGUgbGFzIGxpY2VuY2lhcyBDcmVhdGl2ZSBDb21tb25zLCBlbCBhdXRvciBwb2Ryw6Egc2VsZWNjaW9uYXIgdW5hIGRlIGxhcyA2IGxpY2VuY2lhcyBkaXNwb25pYmxlczoKCgpBdHJpYnVjacOzbiAoQ0MgLSBCeSk6IFBlcm1pdGUgdXNhciBsYSBvYnJhIHkgZ2VuZXJhciBvYnJhcyBkZXJpdmFkYXMsIGluY2x1c28gY29uIGZpbmVzIGNvbWVyY2lhbGVzLCBzaWVtcHJlIHF1ZSBzZSByZWNvbm96Y2EgYWwgYXV0b3IuCgpBdHJpYnVjacOzbiDigJMgQ29tcGFydGlyIElndWFsIChDQyAtIEJ5LVNBKTogUGVybWl0ZSB1c2FyIGxhIG9icmEgeSBnZW5lcmFyIG9icmFzIGRlcml2YWRhcywgaW5jbHVzbyBjb24gZmluZXMgY29tZXJjaWFsZXMsIHBlcm8gbGEgZGlzdHJpYnVjacOzbiBkZSBsYXMgb2JyYXMgZGVyaXZhZGFzIGRlYmUgaGFjZXJzZSBtZWRpYW50ZSB1bmEgbGljZW5jaWEgaWTDqW50aWNhIGEgbGEgZGUgbGEgb2JyYSBvcmlnaW5hbCwgcmVjb25vY2llbmRvIGEgbG9zIGF1dG9yZXMuCgpBdHJpYnVjacOzbiDigJMgTm8gQ29tZXJjaWFsIChDQyAtIEJ5LU5DKTogUGVybWl0ZSB1c2FyIGxhIG9icmEgeSBnZW5lcmFyIG9icmFzIGRlcml2YWRhcywgc2llbXByZSB5IGN1YW5kbyBlc29zIHVzb3Mgbm8gdGVuZ2FuIGZpbmVzIGNvbWVyY2lhbGVzLCByZWNvbm9jaWVuZG8gYWwgYXV0b3IuCgpBdHJpYnVjacOzbiDigJMgU2luIERlcml2YWRhcyAoQ0MgLSBCeS1ORCk6IFBlcm1pdGUgZWwgdXNvIGRlIGxhIG9icmEsIGluY2x1c28gY29uIGZpbmVzIGNvbWVyY2lhbGVzLCBwZXJvIG5vIHNlIHBlcm1pdGUgZ2VuZXJhciBvYnJhcyBkZXJpdmFkYXMsIGRlYmllbmRvIHJlY29ub2NlciBhbCBhdXRvci4KCkF0cmlidWNpw7NuIOKAkyBObyBDb21lcmNpYWwg4oCTIENvbXBhcnRpciBJZ3VhbCAoQ0Mg4oCTIEJ5LU5DLVNBKTogUGVybWl0ZSB1c2FyIGxhIG9icmEgeSBnZW5lcmFyIG9icmFzIGRlcml2YWRhcywgc2llbXByZSB5IGN1YW5kbyBlc29zIHVzb3Mgbm8gdGVuZ2FuIGZpbmVzIGNvbWVyY2lhbGVzIHkgbGEgZGlzdHJpYnVjacOzbiBkZSBsYXMgb2JyYXMgZGVyaXZhZGFzIHNlIGhhZ2EgbWVkaWFudGUgbGljZW5jaWEgaWTDqW50aWNhIGEgbGEgZGUgbGEgb2JyYSBvcmlnaW5hbCwgcmVjb25vY2llbmRvIGEgbG9zIGF1dG9yZXMuCgpBdHJpYnVjacOzbiDigJMgTm8gQ29tZXJjaWFsIOKAkyBTaW4gRGVyaXZhZGFzIChDQyAtIEJ5LU5DLU5EKTogUGVybWl0ZSB1c2FyIGxhIG9icmEsIHBlcm8gbm8gc2UgcGVybWl0ZSBnZW5lcmFyIG9icmFzIGRlcml2YWRhcyB5IG5vIHNlIHBlcm1pdGUgdXNvIGNvbiBmaW5lcyBjb21lcmNpYWxlcywgZGViaWVuZG8gcmVjb25vY2VyIGFsIGF1dG9yLgoKTG9zIHVzb3MgcHJldmlzdG9zIGVuIGxhcyBsaWNlbmNpYXMgaW5jbHV5ZW4gbGEgZW5hamVuYWNpw7NuLCByZXByb2R1Y2Npw7NuLCBjb211bmljYWNpw7NuLCBwdWJsaWNhY2nDs24sIGRpc3RyaWJ1Y2nDs24geSBwdWVzdGEgYSBkaXNwb3NpY2nDs24gZGVsIHDDumJsaWNvLiBMYSBjcmVhY2nDs24gZGUgb2JyYXMgZGVyaXZhZGFzIGluY2x1eWUgbGEgYWRhcHRhY2nDs24sIHRyYWR1Y2Npw7NuIHkgZWwgcmVtaXguCgpDdWFuZG8gc2Ugc2VsZWNjaW9uZSB1bmEgbGljZW5jaWEgcXVlIGhhYmlsaXRlIHVzb3MgY29tZXJjaWFsZXMsIGVsIGRlcMOzc2l0byBkZWJlcsOhIHNlciBhY29tcGHDsWFkbyBkZWwgYXZhbCBkZWwgamVyYXJjYSBtw6F4aW1vIGRlbCBTZXJ2aWNpbyBjb3JyZXNwb25kaWVudGUuCg==Institucionalhttps://www.colibri.udelar.edu.uyUniversidadhttps://udelar.edu.uy/https://www.colibri.udelar.edu.uy/oai/requestkarina.camps@seciu.edu.uyUruguayopendoar:47712025-09-11T18:01:35COLIBRI - Universidad de la Repúblicafalse
spellingShingle Multi-task reinforcement learning in reproducing Kernel Hilbert spaces via cross-learning
Cerviño, Juan
Task analysis
Training data
Navigation
Convergence
Reinforcement learning
Optimization
Multi-task learning
Meta-learning
status_str submittedVersion
title Multi-task reinforcement learning in reproducing Kernel Hilbert spaces via cross-learning
title_full Multi-task reinforcement learning in reproducing Kernel Hilbert spaces via cross-learning
title_fullStr Multi-task reinforcement learning in reproducing Kernel Hilbert spaces via cross-learning
title_full_unstemmed Multi-task reinforcement learning in reproducing Kernel Hilbert spaces via cross-learning
title_short Multi-task reinforcement learning in reproducing Kernel Hilbert spaces via cross-learning
title_sort Multi-task reinforcement learning in reproducing Kernel Hilbert spaces via cross-learning
topic Task analysis
Training data
Navigation
Convergence
Reinforcement learning
Optimization
Multi-task learning
Meta-learning
url https://hdl.handle.net/20.500.12008/51580