Multi-task reinforcement learning in reproducing Kernel Hilbert spaces via cross-learning
Resumen:
Reinforcement learning is a framework to optimize an agent’s policy using rewards that are revealed by the system as a response to an action. In its standard form, reinforcement learning involves a single agent that uses its policy to accomplish a specific task. These methods require large amounts of reward samples to achieve good performance, and may not generalize well when the task is modified, even if the new task is related. In this paper we are interested in a collaborative scheme in which multiple policies are optimized jointly. To this end, we we introduce cross-learning , in which policies are trained for related tasks in separate environments, and they are constrained to be close to one another. Two properties make our new approach attractive: (i) it produces a multi-task central policy that can be used as a starting point to adapt quickly to one of the tasks trained for, and (ii) as in meta-learning , it adapts to environments related but different to those seen during training. We focus on policies belonging to reproducing kernel Hilbert spaces for which we bound the distance between the task-specific policies and the cross-learned policy. To solve the resulting optimization problem, we resort to a projected policy gradient algorithm and prove that it converges to a near-optimal solution with high probability. We evaluate our methodology with a navigation example in which an agent moves through environments with obstacles of multiple shapes and avoids obstacles not trained for.
| 2021 | |
|
ARL DCIST CRA W911NF-17-2-0181 Intel Science and Technology Center for Wireless Autonomous Systems |
|
|
Task analysis Training data Navigation Convergence Reinforcement learning Optimization Multi-task learning Meta-learning |
|
| Inglés | |
| Universidad de la República | |
| COLIBRI | |
| https://hdl.handle.net/20.500.12008/51580 | |
| Acceso abierto | |
| Licencia Creative Commons Atribución (CC - By 4.0) |
| _version_ | 1875692335161933824 |
|---|---|
| author | Cerviño, Juan |
| author2 | Bazerque, Juan Andrés Calvo-Fullana, Miguel Ribeiro, Alejandro |
| author2_role | author author author |
| author_facet | Cerviño, Juan Bazerque, Juan Andrés Calvo-Fullana, Miguel Ribeiro, Alejandro |
| author_role | author |
| bitstream.checksum.fl_str_mv | 6429389a7df7277b72b7924fdc7d47a9 a0ebbeafb9d2ec7cbb19d7137ebc392c aa4ded3991caf203ada54d801dbdfafc 58cb336ce230a47d2f88ad02838a665f 4ddc3993ffc983074c0d51c82009744b |
| bitstream.checksumAlgorithm.fl_str_mv | MD5 MD5 MD5 MD5 MD5 |
| bitstream.url.fl_str_mv | http://localhost:8080/xmlui/bitstream/20.500.12008/51580/5/license.txt http://localhost:8080/xmlui/bitstream/20.500.12008/51580/2/license_url http://localhost:8080/xmlui/bitstream/20.500.12008/51580/3/license_text http://localhost:8080/xmlui/bitstream/20.500.12008/51580/4/license_rdf http://localhost:8080/xmlui/bitstream/20.500.12008/51580/1/CBCR21.pdf |
| collection | COLIBRI |
| dc.contributor.filiacion.none.fl_str_mv | Cerviño Juan, University of Pennsylvania, Philadelphia, USA Bazerque Juan Andrés, Universidad de la República (Uruguay). Facultad de Ingeniería. Calvo-Fullana Miguel, Massachusetts Institute of Technology, Boston, USA Ribeiro Alejandro, University of Pennsylvania, Philadelphia, USA |
| dc.creator.none.fl_str_mv | Cerviño, Juan Bazerque, Juan Andrés Calvo-Fullana, Miguel Ribeiro, Alejandro |
| dc.date.accessioned.none.fl_str_mv | 2025-09-11T18:01:35Z |
| dc.date.available.none.fl_str_mv | 2025-09-11T18:01:35Z |
| dc.date.issued.none.fl_str_mv | 2021 |
| dc.description.abstract.none.fl_txt_mv | Reinforcement learning is a framework to optimize an agent’s policy using rewards that are revealed by the system as a response to an action. In its standard form, reinforcement learning involves a single agent that uses its policy to accomplish a specific task. These methods require large amounts of reward samples to achieve good performance, and may not generalize well when the task is modified, even if the new task is related. In this paper we are interested in a collaborative scheme in which multiple policies are optimized jointly. To this end, we we introduce cross-learning , in which policies are trained for related tasks in separate environments, and they are constrained to be close to one another. Two properties make our new approach attractive: (i) it produces a multi-task central policy that can be used as a starting point to adapt quickly to one of the tasks trained for, and (ii) as in meta-learning , it adapts to environments related but different to those seen during training. We focus on policies belonging to reproducing kernel Hilbert spaces for which we bound the distance between the task-specific policies and the cross-learned policy. To solve the resulting optimization problem, we resort to a projected policy gradient algorithm and prove that it converges to a near-optimal solution with high probability. We evaluate our methodology with a navigation example in which an agent moves through environments with obstacles of multiple shapes and avoids obstacles not trained for. |
| dc.description.sponsorship.none.fl_txt_mv | ARL DCIST CRA W911NF-17-2-0181 Intel Science and Technology Center for Wireless Autonomous Systems |
| dc.format.extent.es.fl_str_mv | 16 p. |
| dc.format.mimetype.es.fl_str_mv | application/pdf |
| dc.identifier.citation.es.fl_str_mv | Cerviño, J., Bazerque, J., Calvo-Fullana, M. y otros. Multi-task reinforcement learning in reproducing Kernel Hilbert spaces via cross-learning [Preprint]. Publicado en: IEEE Transactions on Signal Processing, vol. 69, oct. 2021, pp. 5947-5962. DOI: 10.1109/TSP.2021.3122303. |
| dc.identifier.uri.none.fl_str_mv | https://hdl.handle.net/20.500.12008/51580 |
| dc.language.iso.none.fl_str_mv | en eng |
| dc.rights.license.none.fl_str_mv | Licencia Creative Commons Atribución (CC - By 4.0) |
| dc.rights.none.fl_str_mv | info:eu-repo/semantics/openAccess |
| dc.source.none.fl_str_mv | reponame:COLIBRI instname:Universidad de la República instacron:Universidad de la República |
| dc.subject.es.fl_str_mv | Task analysis Training data Navigation Convergence Reinforcement learning Optimization Multi-task learning Meta-learning |
| dc.title.none.fl_str_mv | Multi-task reinforcement learning in reproducing Kernel Hilbert spaces via cross-learning |
| dc.type.es.fl_str_mv | Preprint |
| dc.type.none.fl_str_mv | info:eu-repo/semantics/preprint |
| dc.type.version.none.fl_str_mv | info:eu-repo/semantics/submittedVersion |
| description | Reinforcement learning is a framework to optimize an agent’s policy using rewards that are revealed by the system as a response to an action. In its standard form, reinforcement learning involves a single agent that uses its policy to accomplish a specific task. These methods require large amounts of reward samples to achieve good performance, and may not generalize well when the task is modified, even if the new task is related. In this paper we are interested in a collaborative scheme in which multiple policies are optimized jointly. To this end, we we introduce cross-learning , in which policies are trained for related tasks in separate environments, and they are constrained to be close to one another. Two properties make our new approach attractive: (i) it produces a multi-task central policy that can be used as a starting point to adapt quickly to one of the tasks trained for, and (ii) as in meta-learning , it adapts to environments related but different to those seen during training. We focus on policies belonging to reproducing kernel Hilbert spaces for which we bound the distance between the task-specific policies and the cross-learned policy. To solve the resulting optimization problem, we resort to a projected policy gradient algorithm and prove that it converges to a near-optimal solution with high probability. We evaluate our methodology with a navigation example in which an agent moves through environments with obstacles of multiple shapes and avoids obstacles not trained for. |
| eu_rights_str_mv | openAccess |
| format | preprint |
| id | COLIBRI_e2a0d0f3249d806d20b2ba43ad245e3e |
| identifier_str_mv | Cerviño, J., Bazerque, J., Calvo-Fullana, M. y otros. Multi-task reinforcement learning in reproducing Kernel Hilbert spaces via cross-learning [Preprint]. Publicado en: IEEE Transactions on Signal Processing, vol. 69, oct. 2021, pp. 5947-5962. DOI: 10.1109/TSP.2021.3122303. |
| instacron_str | Universidad de la República |
| institution | Universidad de la República |
| instname_str | Universidad de la República |
| language | eng |
| language_invalid_str_mv | en |
| network_acronym_str | COLIBRI |
| network_name_str | COLIBRI |
| oai_identifier_str | oai:colibri.udelar.edu.uy:20.500.12008/51580 |
| publishDate | 2021 |
| reponame_str | COLIBRI |
| repository.mail.fl_str_mv | karina.camps@seciu.edu.uy |
| repository.name.fl_str_mv | COLIBRI - Universidad de la República |
| repository_id_str | 4771 |
| rights_invalid_str_mv | Licencia Creative Commons Atribución (CC - By 4.0) |
| spelling | Cerviño Juan, University of Pennsylvania, Philadelphia, USABazerque Juan Andrés, Universidad de la República (Uruguay). Facultad de Ingeniería.Calvo-Fullana Miguel, Massachusetts Institute of Technology, Boston, USARibeiro Alejandro, University of Pennsylvania, Philadelphia, USA2025-09-11T18:01:35Z2025-09-11T18:01:35Z2021Cerviño, J., Bazerque, J., Calvo-Fullana, M. y otros. Multi-task reinforcement learning in reproducing Kernel Hilbert spaces via cross-learning [Preprint]. Publicado en: IEEE Transactions on Signal Processing, vol. 69, oct. 2021, pp. 5947-5962. DOI: 10.1109/TSP.2021.3122303.https://hdl.handle.net/20.500.12008/51580Reinforcement learning is a framework to optimize an agent’s policy using rewards that are revealed by the system as a response to an action. In its standard form, reinforcement learning involves a single agent that uses its policy to accomplish a specific task. These methods require large amounts of reward samples to achieve good performance, and may not generalize well when the task is modified, even if the new task is related. In this paper we are interested in a collaborative scheme in which multiple policies are optimized jointly. To this end, we we introduce cross-learning , in which policies are trained for related tasks in separate environments, and they are constrained to be close to one another. Two properties make our new approach attractive: (i) it produces a multi-task central policy that can be used as a starting point to adapt quickly to one of the tasks trained for, and (ii) as in meta-learning , it adapts to environments related but different to those seen during training. We focus on policies belonging to reproducing kernel Hilbert spaces for which we bound the distance between the task-specific policies and the cross-learned policy. To solve the resulting optimization problem, we resort to a projected policy gradient algorithm and prove that it converges to a near-optimal solution with high probability. We evaluate our methodology with a navigation example in which an agent moves through environments with obstacles of multiple shapes and avoids obstacles not trained for.Submitted by Ribeiro Jorge (jribeiro@fing.edu.uy) on 2025-09-10T18:05:06Z No. of bitstreams: 2 license_rdf: 24942 bytes, checksum: 58cb336ce230a47d2f88ad02838a665f (MD5) CBCR21.pdf: 897474 bytes, checksum: 4ddc3993ffc983074c0d51c82009744b (MD5)Approved for entry into archive by Machado Jimena (jmachado@fing.edu.uy) on 2025-09-11T17:43:38Z (GMT) No. of bitstreams: 2 license_rdf: 24942 bytes, checksum: 58cb336ce230a47d2f88ad02838a665f (MD5) CBCR21.pdf: 897474 bytes, checksum: 4ddc3993ffc983074c0d51c82009744b (MD5)Made available in DSpace by Camps Karina (karina.camps@seciu.edu.uy) on 2025-09-11T18:01:35Z (GMT). No. of bitstreams: 2 license_rdf: 24942 bytes, checksum: 58cb336ce230a47d2f88ad02838a665f (MD5) CBCR21.pdf: 897474 bytes, checksum: 4ddc3993ffc983074c0d51c82009744b (MD5) Previous issue date: 2021ARL DCIST CRA W911NF-17-2-0181Intel Science and Technology Center for Wireless Autonomous Systems16 p.application/pdfenengLas obras depositadas en el Repositorio se rigen por la Ordenanza de los Derechos de la Propiedad Intelectual de la Universidad de la República.(Res. Nº 91 de C.D.C. de 8/III/1994 – D.O. 7/IV/1994) y por la Ordenanza del Repositorio Abierto de la Universidad de la República (Res. Nº 16 de C.D.C. de 07/10/2014)info:eu-repo/semantics/openAccessLicencia Creative Commons Atribución (CC - By 4.0)Task analysisTraining dataNavigationConvergenceReinforcement learningOptimizationMulti-task learningMeta-learningMulti-task reinforcement learning in reproducing Kernel Hilbert spaces via cross-learningPreprintinfo:eu-repo/semantics/preprintinfo:eu-repo/semantics/submittedVersionreponame:COLIBRIinstname:Universidad de la Repúblicainstacron:Universidad de la RepúblicaCerviño, JuanBazerque, Juan AndrésCalvo-Fullana, MiguelRibeiro, AlejandroSistemas y ControlLICENSElicense.txtlicense.txttext/plain; charset=utf-84267http://localhost:8080/xmlui/bitstream/20.500.12008/51580/5/license.txt6429389a7df7277b72b7924fdc7d47a9MD55CC-LICENSElicense_urllicense_urltext/plain; charset=utf-844http://localhost:8080/xmlui/bitstream/20.500.12008/51580/2/license_urla0ebbeafb9d2ec7cbb19d7137ebc392cMD52license_textlicense_texttext/html; charset=utf-826818http://localhost:8080/xmlui/bitstream/20.500.12008/51580/3/license_textaa4ded3991caf203ada54d801dbdfafcMD53license_rdflicense_rdfapplication/rdf+xml; charset=utf-824942http://localhost:8080/xmlui/bitstream/20.500.12008/51580/4/license_rdf58cb336ce230a47d2f88ad02838a665fMD54ORIGINALCBCR21.pdfCBCR21.pdfapplication/pdf897474http://localhost:8080/xmlui/bitstream/20.500.12008/51580/1/CBCR21.pdf4ddc3993ffc983074c0d51c82009744bMD5120.500.12008/515802025-09-11 15:01:35.959oai:colibri.udelar.edu.uy:20.500.12008/51580VGVybWlub3MgeSBjb25kaWNpb25lcyByZWxhdGl2YXMgYWwgZGVwb3NpdG8gZGUgb2JyYXMKCgpMYXMgb2JyYXMgZGVwb3NpdGFkYXMgZW4gZWwgUmVwb3NpdG9yaW8gc2UgcmlnZW4gcG9yIGxhIE9yZGVuYW56YSBkZSBsb3MgRGVyZWNob3MgZGUgbGEgUHJvcGllZGFkIEludGVsZWN0dWFsICBkZSBsYSBVbml2ZXJzaWRhZCBEZSBMYSBSZXDDumJsaWNhLiAoUmVzLiBOwrogOTEgZGUgQy5ELkMuIGRlIDgvSUlJLzE5OTQg4oCTIEQuTy4gNy9JVi8xOTk0KSB5ICBwb3IgbGEgT3JkZW5hbnphIGRlbCBSZXBvc2l0b3JpbyBBYmllcnRvIGRlIGxhIFVuaXZlcnNpZGFkIGRlIGxhIFJlcMO6YmxpY2EgKFJlcy4gTsK6IDE2IGRlIEMuRC5DLiBkZSAwNy8xMC8yMDE0KS4gCgpBY2VwdGFuZG8gZWwgYXV0b3IgZXN0b3MgdMOpcm1pbm9zIHkgY29uZGljaW9uZXMgZGUgZGVww7NzaXRvIGVuIENPTElCUkksIGxhIFVuaXZlcnNpZGFkIGRlIFJlcMO6YmxpY2EgcHJvY2VkZXLDoSBhOiAgCgphKSBhcmNoaXZhciBtw6FzIGRlIHVuYSBjb3BpYSBkZSBsYSBvYnJhIGVuIGxvcyBzZXJ2aWRvcmVzIGRlIGxhIFVuaXZlcnNpZGFkIGEgbG9zIGVmZWN0b3MgZGUgZ2FyYW50aXphciBhY2Nlc28sIHNlZ3VyaWRhZCB5IHByZXNlcnZhY2nDs24KYikgY29udmVydGlyIGxhIG9icmEgYSBvdHJvcyBmb3JtYXRvcyBzaSBmdWVyYSBuZWNlc2FyaW8gIHBhcmEgZmFjaWxpdGFyIHN1IHByZXNlcnZhY2nDs24geSBhY2Nlc2liaWxpZGFkIHNpbiBhbHRlcmFyIHN1IGNvbnRlbmlkby4KYykgcmVhbGl6YXIgbGEgY29tdW5pY2FjacOzbiBww7pibGljYSB5IGRpc3BvbmVyIGVsIGFjY2VzbyBsaWJyZSB5IGdyYXR1aXRvIGEgdHJhdsOpcyBkZSBJbnRlcm5ldCBtZWRpYW50ZSBsYSBwdWJsaWNhY2nDs24gZGUgbGEgb2JyYSBiYWpvIGxhIGxpY2VuY2lhIENyZWF0aXZlIENvbW1vbnMgc2VsZWNjaW9uYWRhIHBvciBlbCBwcm9waW8gYXV0b3IuCgoKRW4gY2FzbyBxdWUgZWwgYXV0b3IgaGF5YSBkaWZ1bmRpZG8geSBkYWRvIGEgcHVibGljaWRhZCBhIGxhIG9icmEgZW4gZm9ybWEgcHJldmlhLCAgcG9kcsOhIHNvbGljaXRhciB1biBwZXLDrW9kbyBkZSBlbWJhcmdvIHNvYnJlIGxhIGRpc3BvbmliaWxpZGFkIHDDumJsaWNhIGRlIGxhIG1pc21hLCBlbCBjdWFsIGNvbWVuemFyw6EgYSBwYXJ0aXIgZGUgbGEgYWNlcHRhY2nDs24gZGUgZXN0ZSBkb2N1bWVudG8geSBoYXN0YSBsYSBmZWNoYSBxdWUgaW5kaXF1ZSAuCgpFbCBhdXRvciBhc2VndXJhIHF1ZSBsYSBvYnJhIG5vIGluZnJpZ2UgbmluZ8O6biBkZXJlY2hvIHNvYnJlIHRlcmNlcm9zLCB5YSBzZWEgZGUgcHJvcGllZGFkIGludGVsZWN0dWFsIG8gY3VhbHF1aWVyIG90cm8uCgpFbCBhdXRvciBnYXJhbnRpemEgcXVlIHNpIGVsIGRvY3VtZW50byBjb250aWVuZSBtYXRlcmlhbGVzIGRlIGxvcyBjdWFsZXMgbm8gdGllbmUgbG9zIGRlcmVjaG9zIGRlIGF1dG9yLCAgaGEgb2J0ZW5pZG8gZWwgcGVybWlzbyBkZWwgcHJvcGlldGFyaW8gZGUgbG9zIGRlcmVjaG9zIGRlIGF1dG9yLCB5IHF1ZSBlc2UgbWF0ZXJpYWwgY3V5b3MgZGVyZWNob3Mgc29uIGRlIHRlcmNlcm9zIGVzdMOhIGNsYXJhbWVudGUgaWRlbnRpZmljYWRvIHkgcmVjb25vY2lkbyBlbiBlbCB0ZXh0byBvIGNvbnRlbmlkbyBkZWwgZG9jdW1lbnRvIGRlcG9zaXRhZG8gZW4gZWwgUmVwb3NpdG9yaW8uCgpFbiBvYnJhcyBkZSBhdXRvcsOtYSBtw7psdGlwbGUgL3NlIHByZXN1bWUvIHF1ZSBlbCBhdXRvciBkZXBvc2l0YW50ZSBkZWNsYXJhIHF1ZSBoYSByZWNhYmFkbyBlbCBjb25zZW50aW1pZW50byBkZSB0b2RvcyBsb3MgYXV0b3JlcyBwYXJhIHB1YmxpY2FybGEgZW4gZWwgUmVwb3NpdG9yaW8sIHNpZW5kbyDDqXN0ZSBlbCDDum5pY28gcmVzcG9uc2FibGUgZnJlbnRlIGEgY3VhbHF1aWVyIHRpcG8gZGUgcmVjbGFtYWNpw7NuIGRlIGxvcyBvdHJvcyBjb2F1dG9yZXMuCgpFbCBhdXRvciBzZXLDoSByZXNwb25zYWJsZSBkZWwgY29udGVuaWRvIGRlIGxvcyBkb2N1bWVudG9zIHF1ZSBkZXBvc2l0YS4gTGEgVURFTEFSIG5vIHNlcsOhIHJlc3BvbnNhYmxlIHBvciBsYXMgZXZlbnR1YWxlcyB2aW9sYWNpb25lcyBhbCBkZXJlY2hvIGRlIHByb3BpZWRhZCBpbnRlbGVjdHVhbCBlbiBxdWUgcHVlZGEgaW5jdXJyaXIgZWwgYXV0b3IuCgpBbnRlIGN1YWxxdWllciBkZW51bmNpYSBkZSB2aW9sYWNpw7NuIGRlIGRlcmVjaG9zIGRlIHByb3BpZWRhZCBpbnRlbGVjdHVhbCwgbGEgVURFTEFSICBhZG9wdGFyw6EgdG9kYXMgbGFzIG1lZGlkYXMgbmVjZXNhcmlhcyBwYXJhIGV2aXRhciBsYSBjb250aW51YWNpw7NuIGRlIGRpY2hhIGluZnJhY2Npw7NuLCBsYXMgcXVlIHBvZHLDoW4gaW5jbHVpciBlbCByZXRpcm8gZGVsIGFjY2VzbyBhIGxvcyBjb250ZW5pZG9zIHkvbyBtZXRhZGF0b3MgZGVsIGRvY3VtZW50byByZXNwZWN0aXZvLgoKTGEgb2JyYSBzZSBwb25kcsOhIGEgZGlzcG9zaWNpw7NuIGRlbCBww7pibGljbyBhIHRyYXbDqXMgZGUgbGFzIGxpY2VuY2lhcyBDcmVhdGl2ZSBDb21tb25zLCBlbCBhdXRvciBwb2Ryw6Egc2VsZWNjaW9uYXIgdW5hIGRlIGxhcyA2IGxpY2VuY2lhcyBkaXNwb25pYmxlczoKCgpBdHJpYnVjacOzbiAoQ0MgLSBCeSk6IFBlcm1pdGUgdXNhciBsYSBvYnJhIHkgZ2VuZXJhciBvYnJhcyBkZXJpdmFkYXMsIGluY2x1c28gY29uIGZpbmVzIGNvbWVyY2lhbGVzLCBzaWVtcHJlIHF1ZSBzZSByZWNvbm96Y2EgYWwgYXV0b3IuCgpBdHJpYnVjacOzbiDigJMgQ29tcGFydGlyIElndWFsIChDQyAtIEJ5LVNBKTogUGVybWl0ZSB1c2FyIGxhIG9icmEgeSBnZW5lcmFyIG9icmFzIGRlcml2YWRhcywgaW5jbHVzbyBjb24gZmluZXMgY29tZXJjaWFsZXMsIHBlcm8gbGEgZGlzdHJpYnVjacOzbiBkZSBsYXMgb2JyYXMgZGVyaXZhZGFzIGRlYmUgaGFjZXJzZSBtZWRpYW50ZSB1bmEgbGljZW5jaWEgaWTDqW50aWNhIGEgbGEgZGUgbGEgb2JyYSBvcmlnaW5hbCwgcmVjb25vY2llbmRvIGEgbG9zIGF1dG9yZXMuCgpBdHJpYnVjacOzbiDigJMgTm8gQ29tZXJjaWFsIChDQyAtIEJ5LU5DKTogUGVybWl0ZSB1c2FyIGxhIG9icmEgeSBnZW5lcmFyIG9icmFzIGRlcml2YWRhcywgc2llbXByZSB5IGN1YW5kbyBlc29zIHVzb3Mgbm8gdGVuZ2FuIGZpbmVzIGNvbWVyY2lhbGVzLCByZWNvbm9jaWVuZG8gYWwgYXV0b3IuCgpBdHJpYnVjacOzbiDigJMgU2luIERlcml2YWRhcyAoQ0MgLSBCeS1ORCk6IFBlcm1pdGUgZWwgdXNvIGRlIGxhIG9icmEsIGluY2x1c28gY29uIGZpbmVzIGNvbWVyY2lhbGVzLCBwZXJvIG5vIHNlIHBlcm1pdGUgZ2VuZXJhciBvYnJhcyBkZXJpdmFkYXMsIGRlYmllbmRvIHJlY29ub2NlciBhbCBhdXRvci4KCkF0cmlidWNpw7NuIOKAkyBObyBDb21lcmNpYWwg4oCTIENvbXBhcnRpciBJZ3VhbCAoQ0Mg4oCTIEJ5LU5DLVNBKTogUGVybWl0ZSB1c2FyIGxhIG9icmEgeSBnZW5lcmFyIG9icmFzIGRlcml2YWRhcywgc2llbXByZSB5IGN1YW5kbyBlc29zIHVzb3Mgbm8gdGVuZ2FuIGZpbmVzIGNvbWVyY2lhbGVzIHkgbGEgZGlzdHJpYnVjacOzbiBkZSBsYXMgb2JyYXMgZGVyaXZhZGFzIHNlIGhhZ2EgbWVkaWFudGUgbGljZW5jaWEgaWTDqW50aWNhIGEgbGEgZGUgbGEgb2JyYSBvcmlnaW5hbCwgcmVjb25vY2llbmRvIGEgbG9zIGF1dG9yZXMuCgpBdHJpYnVjacOzbiDigJMgTm8gQ29tZXJjaWFsIOKAkyBTaW4gRGVyaXZhZGFzIChDQyAtIEJ5LU5DLU5EKTogUGVybWl0ZSB1c2FyIGxhIG9icmEsIHBlcm8gbm8gc2UgcGVybWl0ZSBnZW5lcmFyIG9icmFzIGRlcml2YWRhcyB5IG5vIHNlIHBlcm1pdGUgdXNvIGNvbiBmaW5lcyBjb21lcmNpYWxlcywgZGViaWVuZG8gcmVjb25vY2VyIGFsIGF1dG9yLgoKTG9zIHVzb3MgcHJldmlzdG9zIGVuIGxhcyBsaWNlbmNpYXMgaW5jbHV5ZW4gbGEgZW5hamVuYWNpw7NuLCByZXByb2R1Y2Npw7NuLCBjb211bmljYWNpw7NuLCBwdWJsaWNhY2nDs24sIGRpc3RyaWJ1Y2nDs24geSBwdWVzdGEgYSBkaXNwb3NpY2nDs24gZGVsIHDDumJsaWNvLiBMYSBjcmVhY2nDs24gZGUgb2JyYXMgZGVyaXZhZGFzIGluY2x1eWUgbGEgYWRhcHRhY2nDs24sIHRyYWR1Y2Npw7NuIHkgZWwgcmVtaXguCgpDdWFuZG8gc2Ugc2VsZWNjaW9uZSB1bmEgbGljZW5jaWEgcXVlIGhhYmlsaXRlIHVzb3MgY29tZXJjaWFsZXMsIGVsIGRlcMOzc2l0byBkZWJlcsOhIHNlciBhY29tcGHDsWFkbyBkZWwgYXZhbCBkZWwgamVyYXJjYSBtw6F4aW1vIGRlbCBTZXJ2aWNpbyBjb3JyZXNwb25kaWVudGUuCg==Institucionalhttps://www.colibri.udelar.edu.uyUniversidadhttps://udelar.edu.uy/https://www.colibri.udelar.edu.uy/oai/requestkarina.camps@seciu.edu.uyUruguayopendoar:47712025-09-11T18:01:35COLIBRI - Universidad de la Repúblicafalse |
| spellingShingle | Multi-task reinforcement learning in reproducing Kernel Hilbert spaces via cross-learning Cerviño, Juan Task analysis Training data Navigation Convergence Reinforcement learning Optimization Multi-task learning Meta-learning |
| status_str | submittedVersion |
| title | Multi-task reinforcement learning in reproducing Kernel Hilbert spaces via cross-learning |
| title_full | Multi-task reinforcement learning in reproducing Kernel Hilbert spaces via cross-learning |
| title_fullStr | Multi-task reinforcement learning in reproducing Kernel Hilbert spaces via cross-learning |
| title_full_unstemmed | Multi-task reinforcement learning in reproducing Kernel Hilbert spaces via cross-learning |
| title_short | Multi-task reinforcement learning in reproducing Kernel Hilbert spaces via cross-learning |
| title_sort | Multi-task reinforcement learning in reproducing Kernel Hilbert spaces via cross-learning |
| topic | Task analysis Training data Navigation Convergence Reinforcement learning Optimization Multi-task learning Meta-learning |
| url | https://hdl.handle.net/20.500.12008/51580 |