Cognitive inference as the main predictor of AI reliability in automated behavioral coding of parent–child interactions

Amorocho, José - Balsa, Ana - Giraldo-Huertas, Juan José - Bloomfield, Juanita - Patrone, Paula - Cid, Alejandro

Resumen:

Observational coding of parent–child interactions is a gold standard in developmental science but remains unscalable. Multimodal generative AI could help, yet its reliability and failure modes are not well characterized. We benchmarked a multi-agent pipeline (GABRIEL) against a multi-rater expert consensus when scoring 22 PICCOLO items on 156 ten-minute free-play interactions from Uruguay. Agreement was summarized with Percent Agreement (PA) and Cohen’s κ, and disagreement with a unit-free normalized mean squared error (nMSE = MSE/Var(Yitem)). A priori item classes indexed cognitive inference (Low/Medium/High). Final calibration yielded modest agreement (PA = 50.7%, κ = .216). Disagreement was chiefly structured by inference (Kruskal–Wallis H = 308.70, p<.001), a pattern that persisted in late iterations. The model also overused the middle category (1) and underused “2.” No systematic differences in nMSE emerged by sex, age quartile, or maternal education. We conclude that generative AI is promising for scalable detection of concrete, low-inference behaviors, whereas high-inference judgments still require expert adjudication. A human-in-the-loop, co-intelligence workflow aligns current strengths with ethical oversight and supports equitable deployment at scale.

Detalles Bibliográficos
2026
Generative AI
Parent-Child Interaction
Observational Coding
Inter-Rater Reliability
PICCOLO
Multimodal AI
Inglés
Universidad de Montevideo
REDUM
https://hdl.handle.net/20.500.12806/2795
Acceso abierto
Attribution-NonCommercial-NoDerivatives 4.0 Internacional
_version_ 1875304570634108928
author Amorocho, José
author2 Balsa, Ana
Giraldo-Huertas, Juan José
Bloomfield, Juanita
Patrone, Paula
Cid, Alejandro
author2_role author
author
author
author
author
author_facet Amorocho, José
Balsa, Ana
Giraldo-Huertas, Juan José
Bloomfield, Juanita
Patrone, Paula
Cid, Alejandro
author_role author
bitstream.checksum.fl_str_mv d2ac6d3274f0b3928c2cfbf4ac6f0b6d
4460e5956bc1d1639be9ae6146a50347
691ed290c8bf8671811a9242b7fc04b6
d61271a755255cce727e8a46bedaef5e
7a9220d896ec1792ba4acd19a3b915be
bitstream.checksumAlgorithm.fl_str_mv MD5
MD5
MD5
MD5
MD5
bitstream.url.fl_str_mv http://redum.um.edu.uy/bitstream/20.500.12806/2795/1/Amorocho%2c%20Balsa%2c%20Giraldo-Huertas%2c%20Bloomfield%2c%20Patrone%2c%20Cid.pdf
http://redum.um.edu.uy/bitstream/20.500.12806/2795/2/license_rdf
http://redum.um.edu.uy/bitstream/20.500.12806/2795/3/license.txt
http://redum.um.edu.uy/bitstream/20.500.12806/2795/4/Amorocho%2c%20Balsa%2c%20Giraldo-Huertas%2c%20Bloomfield%2c%20Patrone%2c%20Cid.pdf.txt
http://redum.um.edu.uy/bitstream/20.500.12806/2795/5/Amorocho%2c%20Balsa%2c%20Giraldo-Huertas%2c%20Bloomfield%2c%20Patrone%2c%20Cid.pdf.jpg
collection REDUM
dc.creator.none.fl_str_mv Amorocho, José
Balsa, Ana
Giraldo-Huertas, Juan José
Bloomfield, Juanita
Patrone, Paula
Cid, Alejandro
dc.date.accessioned.none.fl_str_mv 2026-02-27T14:00:24Z
dc.date.available.none.fl_str_mv 2026-02-27T14:00:24Z
dc.date.issued.es.fl_str_mv 2026
dc.description.abstract.none.fl_txt_mv Observational coding of parent–child interactions is a gold standard in developmental science but remains unscalable. Multimodal generative AI could help, yet its reliability and failure modes are not well characterized. We benchmarked a multi-agent pipeline (GABRIEL) against a multi-rater expert consensus when scoring 22 PICCOLO items on 156 ten-minute free-play interactions from Uruguay. Agreement was summarized with Percent Agreement (PA) and Cohen’s κ, and disagreement with a unit-free normalized mean squared error (nMSE = MSE/Var(Yitem)). A priori item classes indexed cognitive inference (Low/Medium/High). Final calibration yielded modest agreement (PA = 50.7%, κ = .216). Disagreement was chiefly structured by inference (Kruskal–Wallis H = 308.70, p<.001), a pattern that persisted in late iterations. The model also overused the middle category (1) and underused “2.” No systematic differences in nMSE emerged by sex, age quartile, or maternal education. We conclude that generative AI is promising for scalable detection of concrete, low-inference behaviors, whereas high-inference judgments still require expert adjudication. A human-in-the-loop, co-intelligence workflow aligns current strengths with ethical oversight and supports equitable deployment at scale.
dc.format.extent.es.fl_str_mv 31 p.
dc.format.mimetype.es.fl_str_mv text/plain
dc.identifier.uri.none.fl_str_mv https://hdl.handle.net/20.500.12806/2795
dc.language.iso.none.fl_str_mv eng
dc.rights.es.fl_str_mv Abierto
dc.rights.license.none.fl_str_mv Attribution-NonCommercial-NoDerivatives 4.0 Internacional
dc.rights.none.fl_str_mv info:eu-repo/semantics/openAccess
dc.rights.uri.*.fl_str_mv http://creativecommons.org/licenses/by-nc-nd/4.0/
dc.source.none.fl_str_mv reponame:REDUM
instname:Universidad de Montevideo
instacron:Universidad de Montevideo
dc.subject.keyword.es.fl_str_mv Generative AI
Parent-Child Interaction
Observational Coding
Inter-Rater Reliability
PICCOLO
Multimodal AI
dc.title.none.fl_str_mv Cognitive inference as the main predictor of AI reliability in automated behavioral coding of parent–child interactions
dc.type.es.fl_str_mv Preprint
dc.type.none.fl_str_mv info:eu-repo/semantics/preprint
dc.type.version.es.fl_str_mv Aceptada
dc.type.version.none.fl_str_mv info:eu-repo/semantics/acceptedVersion
description Observational coding of parent–child interactions is a gold standard in developmental science but remains unscalable. Multimodal generative AI could help, yet its reliability and failure modes are not well characterized. We benchmarked a multi-agent pipeline (GABRIEL) against a multi-rater expert consensus when scoring 22 PICCOLO items on 156 ten-minute free-play interactions from Uruguay. Agreement was summarized with Percent Agreement (PA) and Cohen’s κ, and disagreement with a unit-free normalized mean squared error (nMSE = MSE/Var(Yitem)). A priori item classes indexed cognitive inference (Low/Medium/High). Final calibration yielded modest agreement (PA = 50.7%, κ = .216). Disagreement was chiefly structured by inference (Kruskal–Wallis H = 308.70, p<.001), a pattern that persisted in late iterations. The model also overused the middle category (1) and underused “2.” No systematic differences in nMSE emerged by sex, age quartile, or maternal education. We conclude that generative AI is promising for scalable detection of concrete, low-inference behaviors, whereas high-inference judgments still require expert adjudication. A human-in-the-loop, co-intelligence workflow aligns current strengths with ethical oversight and supports equitable deployment at scale.
eu_rights_str_mv openAccess
format preprint
id REDUM_5ada4b332199a7eabf341081a59a2bce
instacron_str Universidad de Montevideo
institution Universidad de Montevideo
instname_str Universidad de Montevideo
language eng
network_acronym_str REDUM
network_name_str REDUM
oai_identifier_str oai:redum.um.edu.uy:20.500.12806/2795
publishDate 2026
reponame_str REDUM
repository.mail.fl_str_mv nolascoaga@um.edu.uy
repository.name.fl_str_mv REDUM - Universidad de Montevideo
repository_id_str 10501
rights_invalid_str_mv Attribution-NonCommercial-NoDerivatives 4.0 Internacional
Abierto
http://creativecommons.org/licenses/by-nc-nd/4.0/
spelling Attribution-NonCommercial-NoDerivatives 4.0 InternacionalAbiertohttp://creativecommons.org/licenses/by-nc-nd/4.0/info:eu-repo/semantics/openAccess5966f393-0189-4978-88b6-86027c49c71c16b83a11-57bb-434e-8c04-df7fbe5ff1e829c64f3d-4249-456f-a848-ff321aa7b9e71636f80f-9b18-4eb3-aa6a-8d948681450f382fb732-32e3-4c28-b23e-fc1af67739276383a2bd-e52f-4ece-8cf6-e88ad065d71f2026-02-27T14:00:24Z2026-02-27T14:00:24Z2026https://hdl.handle.net/20.500.12806/279531 p.text/plainengCognitive inference as the main predictor of AI reliability in automated behavioral coding of parent–child interactionsPreprintAceptadainfo:eu-repo/semantics/acceptedVersioninfo:eu-repo/semantics/preprintObservational coding of parent–child interactions is a gold standard in developmental science but remains unscalable. Multimodal generative AI could help, yet its reliability and failure modes are not well characterized. We benchmarked a multi-agent pipeline (GABRIEL) against a multi-rater expert consensus when scoring 22 PICCOLO items on 156 ten-minute free-play interactions from Uruguay. Agreement was summarized with Percent Agreement (PA) and Cohen’s κ, and disagreement with a unit-free normalized mean squared error (nMSE = MSE/Var(Yitem)). A priori item classes indexed cognitive inference (Low/Medium/High). Final calibration yielded modest agreement (PA = 50.7%, κ = .216). Disagreement was chiefly structured by inference (Kruskal–Wallis H = 308.70, p<.001), a pattern that persisted in late iterations. The model also overused the middle category (1) and underused “2.” No systematic differences in nMSE emerged by sex, age quartile, or maternal education. We conclude that generative AI is promising for scalable detection of concrete, low-inference behaviors, whereas high-inference judgments still require expert adjudication. A human-in-the-loop, co-intelligence workflow aligns current strengths with ethical oversight and supports equitable deployment at scale.Generative AIParent-Child InteractionObservational CodingInter-Rater ReliabilityPICCOLOMultimodal AIreponame:REDUMinstname:Universidad de Montevideoinstacron:Universidad de MontevideoAmorocho, JoséBalsa, AnaGiraldo-Huertas, Juan JoséBloomfield, JuanitaPatrone, PaulaCid, AlejandroORIGINALAmorocho, Balsa, Giraldo-Huertas, Bloomfield, Patrone, Cid.pdfAmorocho, Balsa, Giraldo-Huertas, Bloomfield, Patrone, Cid.pdfapplication/pdf1276343http://redum.um.edu.uy/bitstream/20.500.12806/2795/1/Amorocho%2c%20Balsa%2c%20Giraldo-Huertas%2c%20Bloomfield%2c%20Patrone%2c%20Cid.pdfd2ac6d3274f0b3928c2cfbf4ac6f0b6dMD51CC-LICENSElicense_rdflicense_rdfapplication/rdf+xml; charset=utf-8805http://redum.um.edu.uy/bitstream/20.500.12806/2795/2/license_rdf4460e5956bc1d1639be9ae6146a50347MD52LICENSElicense.txtlicense.txttext/plain; charset=utf-82117http://redum.um.edu.uy/bitstream/20.500.12806/2795/3/license.txt691ed290c8bf8671811a9242b7fc04b6MD53TEXTAmorocho, Balsa, Giraldo-Huertas, Bloomfield, Patrone, Cid.pdf.txtAmorocho, Balsa, Giraldo-Huertas, Bloomfield, Patrone, Cid.pdf.txtExtracted texttext/plain81229http://redum.um.edu.uy/bitstream/20.500.12806/2795/4/Amorocho%2c%20Balsa%2c%20Giraldo-Huertas%2c%20Bloomfield%2c%20Patrone%2c%20Cid.pdf.txtd61271a755255cce727e8a46bedaef5eMD54THUMBNAILAmorocho, Balsa, Giraldo-Huertas, Bloomfield, Patrone, Cid.pdf.jpgAmorocho, Balsa, Giraldo-Huertas, Bloomfield, Patrone, Cid.pdf.jpgGenerated Thumbnailimage/jpeg1390http://redum.um.edu.uy/bitstream/20.500.12806/2795/5/Amorocho%2c%20Balsa%2c%20Giraldo-Huertas%2c%20Bloomfield%2c%20Patrone%2c%20Cid.pdf.jpg7a9220d896ec1792ba4acd19a3b915beMD5520.500.12806/27952026-07-08 14:35:37.658oai:redum.um.edu.uy:20.500.12806/2795TGljZW5jaWEgZGUgRGlzdHJpYnVjacOzbiBObyBFeGNsdXNpdmEgCkF1dG9yaXphY2nDs24gcGFyYSBsYSBwdWJsaWNhY2nDs24gZW4gZWwgUmVwb3NpdG9yaW8gRGlnaXRhbCBVbml2ZXJzaWRhZCBkZSBNb250ZXZpZGVvIChSRURVTSkKClBhcmEgcXVlIGVsIFJFRFVNIGFsbWFjZW5lLCByZXByb2R1emNhIHkgZGlmdW5kYSBww7pibGljYW1lbnRlIGxhIG9icmEgcXVlIHNlIGRlcG9zaXRhLCBlcyBuZWNlc2FyaW8gcXVlIGFjZXB0ZSBsb3Mgc2lndWllbnRlcyB0w6lybWlub3M6CgoxLglBdXRvcml6byBhIGxhIFVuaXZlcnNpZGFkIGRlIE1vbnRldmlkZW8gZWwgZGVyZWNobyBubyBleGNsdXNpdm8gZGUgYWxtYWNlbmFyLCByZXByb2R1Y2lyLCBjb211bmljYXIgeS9vIGRpc3RyaWJ1aXIgZ3JhdHVpdGEgeSBww7pibGljYW1lbnRlIGVzdGEgb2JyYSBiYWpvIGZvcm1hdG8gZWxlY3Ryw7NuaWNvIGVuIGVsIFJFRFVNLiAKMi4JRXN0b3kgZGUgYWN1ZXJkbyBlbiBxdWUgbGEgVW5pdmVyc2lkYWQgZGUgTW9udGV2aWRlbyBwdWVkYSBjb25zZXJ2YXIgbcOhcyBkZSB1bmEgY29waWEgZGUgZXN0YSBvYnJhIHksIHNpbiBhbHRlcmFyIHN1IGNvbnRlbmlkbywgY29udmVydGlybG8gYSBjdWFscXVpZXIgZm9ybWF0byBkZSBhcmNoaXZvLCBtZWRpbyBvIHNvcG9ydGUsIHBhcmEgcHJvcMOzc2l0b3MgZGUgc2VndXJpZGFkLCBwcmVzZXJ2YWNpw7NuIHkgYWNjZXNvLiAKMy4JQXV0b3Jpem8gbGEgcmVwcm9kdWNjacOzbiB0b3RhbCBvIHBhcmNpYWwgZGUgbGEgb2JyYSwgc2llbXByZSBhc29jaWFkYSBhIG1pIG5vbWJyZSBlbiBjYWxpZGFkIGRlIGF1dG9yLgo0LglFbiBuaW5ndW5hIGNpcmN1bnN0YW5jaWEgYXV0b3Jpem8gbGEgYWRhcHRhY2nDs24sIHRyYW5zZm9ybWFjacOzbiwgdHJhZHVjY2nDs24geSBlbiBnZW5lcmFsIGN1YWxxdWllciB0aXBvIGRlIG1vZGlmaWNhY2nDs24gYSBtaSBvYnJhLgo1LglEZWNsYXJvIHF1ZSBlc3RhIG9icmEgZXMgdW4gdHJhYmFqbyBvcmlnaW5hbCB5IHF1ZSBzb3kgYXV0b3IgZGUgbGEgbWlzbWEgeSBxdWUgbm8gaGUgb3RvcmdhZG8gZXNvcyBkZXJlY2hvcyBhIHRlcmNlcm9zIHF1ZSBwdWVkYW4gbGltaXRhciBhIGxhIFVNIHBhcmEgZWplcmNlcmxvcy4gCjYuCUVuIGVsIGNhc28gZGUgcXVlIGVzdGEgb2JyYSBjb250ZW5nYSBtYXRlcmlhbCBwYXJhIGVsIHF1ZSBubyBwb3NlbyBkZXJlY2hvcyBkZSBhdXRvciwgZGVjbGFybyBxdWUgaGUgb2J0ZW5pZG8gZWwgcGVybWlzbyBzaW4gcmVzdHJpY2Npb25lcyBkZWwKcHJvcGlldGFyaW8gZGUgbG9zIGRlcmVjaG9zIGRlIGF1dG9yIHBhcmEgb3RvcmdhciBhIGxhIFVuaXZlcnNpZGFkIGRlIE1vbnRldmlkZW8gbG9zIGRlcmVjaG9zIHJlcXVlcmlkb3MgcG9yIGVzdGEgbGljZW5jaWEsIHkgcXVlIGRpY2hvIG1hdGVyaWFsIGRlIHRlcmNlcm9zIGVzdMOhIGNsYXJhbWVudGUgaWRlbnRpZmljYWRvIHkgcmVjb25vY2lkbyBkZW50cm8gZGVsIHRleHRvIG8gY29udGVuaWRvIGRlIGxhIHByZXNlbnRhY2nDs24uIAo3LglEZWNsYXJvIHF1ZSBsYSBVbml2ZXJzaWRhZCBkZSBNb250ZXZpZGVvIHF1ZWRhIGV4Y2x1aWRhIGRlIHRvZGEgcmVzcG9uc2FiaWxpZGFkIGVuIGNhc28gZGUgZXZlbnR1YWxlcyByZWNsYW1hY2lvbmVzIGRlIHRlcmNlcm9zIHBvciBpbmZyaW5naXIgbG9zIGRlcmVjaG9zIGRlIGF1dG9yLiAKOC4JRGVjbGFybyBxdWUgbGFzIG9waW5pb25lcyBleHByZXNhZGFzIGVuIGVzdGUgZG9jdW1lbnRvIG5vIHNvbiBuZWNlc2FyaWFtZW50ZSBjb21wYXJ0aWRhcyBwb3IgbGEgVW5pdmVyc2lkYWQgZGUgTW9udGV2aWRlby4gCjkuCUF1dG9yaXpvIGEgbG9zIHVzdWFyaW9zIGRlbCBSRURVTSBwYXJhIHV0aWxpemFyLCByZXByb2R1Y2lyIHkgZGlzdHJpYnVpciBlc3RlIGRvY3VtZW50byBiYWpvIGxpY2VuY2lhIENyZWF0aXZlIENvbW1vbnMgNC4wLiAoQXRyaWJ1Y2nDs24tTm9Db21lcmNpYWwtU2luRGVyaXZhZGFzIDQuMCBJbnRlcm5hY2lvbmFsKS4KClNpIHRpZW5lIGFsZ3VuYSBkdWRhIHNvYnJlIGxvcyB0w6lybWlub3MgZGUgZXN0YSBhdXRvcml6YWNpw7NuLCBwb3IgZmF2b3IgZW52w61lbGEgYSBiaWJsaW90ZWNhQHVtLmVkdS51eQoKCgo=Institucionalhttps://redum.um.edu.uy/Universidadhttps://um.edu.uy/https://redum.um.edu.uy/oai/requestnolascoaga@um.edu.uyUruguayopendoar:105012026-07-08T17:35:37REDUM - Universidad de Montevideofalse
spellingShingle Cognitive inference as the main predictor of AI reliability in automated behavioral coding of parent–child interactions
Amorocho, José
Generative AI
Parent-Child Interaction
Observational Coding
Inter-Rater Reliability
PICCOLO
Multimodal AI
status_str acceptedVersion
title Cognitive inference as the main predictor of AI reliability in automated behavioral coding of parent–child interactions
title_full Cognitive inference as the main predictor of AI reliability in automated behavioral coding of parent–child interactions
title_fullStr Cognitive inference as the main predictor of AI reliability in automated behavioral coding of parent–child interactions
title_full_unstemmed Cognitive inference as the main predictor of AI reliability in automated behavioral coding of parent–child interactions
title_short Cognitive inference as the main predictor of AI reliability in automated behavioral coding of parent–child interactions
title_sort Cognitive inference as the main predictor of AI reliability in automated behavioral coding of parent–child interactions
topic Generative AI
Parent-Child Interaction
Observational Coding
Inter-Rater Reliability
PICCOLO
Multimodal AI
url https://hdl.handle.net/20.500.12806/2795