scieee AI-readable full text Open interactive document viewer

Evaluación forense de la credibilidad del testimonio y sintomatología internalizante en delitos cometidos en la esfera privada

Amado, Bárbara G.

Abstract

El análisis de la realidad de la declaración y la cuantificación del daño psicológico consecuencia de la victimización de un delito, contribuyen a la suficiencia de la prueba. Se puso a prueba la validez del Criteria-Based Content Analysis como técnica de análisis de contenido de la declaración y la Hipótesis Undeutsch en la cual se sustenta; así como su admisibilidad en la justicia a partir del cumplimiento de los Daubert Standards y criterios jurisprudenciales, y su validez como sistema categorial metódico. Asimismo, se cuantificó la huella (sintomatología internalizante) consecuencia de la victimización de abusos sexuales en la infancia y adolescencia y estimado sus repercusiones en la adultez. Se discuten las implicaciones de los resultados en la práctica de la pericial psicológica (estudio de la credibilidad y huella psicológica) y la admisibilidad del CBCA en la justicia.

Full text

TESIS DOCTORAL MODALIDAD DE COMPENDIO DE ARTÍCULOS EVALUACIÓN FORENSE DE LA CREDIBILIDAD DEL TESTIMONIO Y SINTOMATOLOGÍA INTERNALIZANTE EN DELITOS COMETIDOS EN LA ESFERA PRIVADA Bárbara González Amado PROGRAMA DE DOUTORAMENTO EN PSICOLOXÍA DO TRABALLO E AS ORGANIZACIÓNS, XURÍDICA-FORENSE E DO CONSUMIDOR E USUARIO (RD 99/2011) FACULTADE DE PSICOLOXÍA SANTIAGO DE COMPOSTELA 2017 TESIS DOCTORAL MODALIDAD DE COMPENDIO DE ARTÍCULOS EVALUACIÓN FORENSE DE LA CREDIBILIDAD DEL TESTIMONIO Y SINTOMATOLOGÍA INTERNALIZANTE EN DELITOS COMETIDOS EN LA ESFERA PRIVADA Fdo.:…………………………. Bárbara González Amado PROGRAMA DE DOUTORAMENTO EN PSICOLOXÍA DO TRABALLO E AS ORGANIZACIÓNS, XURÍDICA-FORENSE E DO CONSUMIDOR E USUARIO (RD 99/2011) FACULTADE DE PSICOLOXÍA SANTIAGO DE COMPOSTELA 2017 AUTORIZACIÓN DO DIRECTOR/TITOR E DIRECTORA DA TESE: AVALIACIÓN FORENSE DA CREDIBILIDADE DAS TESTEMUÑAS E SINTOMATOLOXÍA INTERNALIZANTE EN DELITOS COMETIDOS NA ESFERA PRIVADA D. Ramón Arce Fernández, Catedrático de Psicoloxía Xurídica e Forense da Universidade de Santiago de Compostela e Dna. Francisca Fariña Rivera, Catedrática de Psicoloxía Básica e Psicoloxía Xurídica do Menor da Universidade de Vigo, INFORMAN Que a presente tese, correspóndese co traballo realizado por Dna. Bárbara González Amado, baixo a nosa dirección, e autorizamos á presentación da tese indicada, considerando que reúne os requisitos esixidos no artigo 33 do regulamento de Estudos de Doutoramento da USC, e que como director e directora desta non incurrimos nas causas de abstención establecidas na lei 40/2015. Asinado. Ramón Arce Fernández Asinado. Francisca Fariña Rivera Prof. Dr. D. Ramón Arce Fernández e a Profa. Dra. Dna. Francisca Fariña Rivera, como director/titor e directora da tese titulada: AVALIACIÓN FORENSE DA CREDIBILIDADE DAS TESTEMUÑAS E SINTOMATOLOXÍA INTERNALIZANTE EN DELITOS COMETIDOS NA ESFERA PRIVADA. Pola presente DECLARAMOS: Que a tese presentada por Dona Bárbara González Amado é idónea para ser presentada, de acordo co artigo 41 do Regulamento de Estudos de Doutoramento, pola modalidade de compendio de ARTIGOS, nos que o doutorando tivo participación no peso da investigación e a súa contribución foi decisiva para levar a cabo este traballo. E que está en coñecemento dos coautores, tanto doutores como non doutores, participantes nos artigos, que ningún dos traballos reunidos nesta tese serán presentados por ningún deles noutra tese de Doutoramento, o que asino baixo a miña responsabilidade. Santiago de Compostela, a 9 de Maio de 2017 Asinado. Ramón Arce Fernández Asinado. Francisca Fariña Rivera AGRADECEMENTOS Agradecer, en primeiro lugar, ó meu tutor e director, Ramón Arce, e directora de tese, Francisca Fariña, pola confianza depositada en min durante estes anos e por todas as horas adicadas a transmitirme os seus coñecementos. A Mercedes Novo e Loli Seijo, polos seus consellos e pola súa inesgotable capacidade para facerme crer en min mesma. A todos eles darlles as gracias porque aquí comeza unha traxectoria profesional, cunha mochila que porto ás miñas costas cargada de aprendizaxes, e que non sería posible sen a súa labor. Non quero esquecerme do profesor Jesús Salgado, a quen teño que agradecer a súa inestimable axuda durante os momentos de dúbida e, por suposto, por ser quén de contaxiarme a complexa beleza do mundo do meta-análise. Agradezo ó profesor Rui Abrunhosa a súa hospitalidade, profesionalidade e a atención que me brindou sempre que lle foi requerida. As miñas compañeiras e compañeiros da Unidade de Psicoloxía Forense, agradecerlles as xornadas de choiva de ideas e de reflexión. Parte da mochila que levo tamén é labor vosa. Ó meu pai e á miña nai, ós que lles debo todo o que son, agradezo o seu apoio incondicional e as súas verbas sempre oportunas que me mantiveron en pé. Ó seu espírito de loita é o que me empurra a seguir cara adiante. A mochila está cargada de todas as vosas ensinanzas. A ti, meu compañeiro de vida, por servirme de reflexo e de guía. ÍNDICE 1. INTRODUCCIÓN ............................................................................................................ 21 1.1. LA PERICIAL PSICOLÓGICA Y EL ANÁLISIS DE LA CREDIBILIDAD DEL TESTIMONIO ........ 21 1.2. INSTRUMENTOS DE EVALUACIÓN DE LA REALIDAD DEL TESTIMONIO ........................... 24 1.3. STATEMENT VALIDITY ASSESSMENT (SVA) ................................................................ 25 1.3.1. Origen y estructura ............................................................................................. 25 1.3.2. Admisibilidad del SVA en los distintos países ................................................... 27 1.4. CRITERIA-BASED CONTENT ANALYSIS (CBCA) .......................................................... 28 1.4.1. Paradigmas de investigación. Estudios de campo y experimentales .................. 31 1.4.2. Limitaciones de la técnica .................................................................................. 32 1.5. ABUSO SEXUAL EN LA INFANCIA Y ADOLESCENCIA. SECUELAS INTERNALIZANTES DE LA VICTIMIZACIÓN ................................................................................................... 34 1.6. EL PROCEDIMIENTO META-ANALÍTICO ......................................................................... 37 1.6.1. Tipos de meta-análisis ........................................................................................ 40 2. OBJETIVOS ...................................................................................................................... 43 3. MÉTODO .......................................................................................................................... 45 4. RESULTADOS ................................................................................................................. 49 5. DISCUSIÓN ...................................................................................................................... 55 6. DISCUSSION .................................................................................................................... 65 7. REFERENCIAS ................................................................................................................ 75 8. ARTÍCULOS DE INVESTIGACIÓN ............................................................................ 83 8.1. UNDEUTSCH HYPOTHESIS AND CRITERIA BASED CONTENT ANALYSIS: A METAANALYTIC REVIEW ........................................................................................................ 85 8.2. CRITERIA-BASED CONTENT ANALYSIS (CBCA) REALITY CRITERIA IN ADULTS: A META-ANALYTIC REVIEW ........................................................................................... 107 8.3. PSYCHOLOGICAL INJURY IN VICTIMS OF CHILD SEXUAL ABUSE: A METAANALYTIC REVIEW ...................................................................................................... 123 8.4. ANÁLISIS DE CONTENIDO EN DECLARACIONES DE AGRESORES. UNA REVISIÓN META-ANALÍTICA ....................................................................................................... 153 9. ÍNDICE DE TABLAS ..................................................................................................... 161 10. APÉNDICE ...................................................................................................................... 173 10.1. RESUMEN ................................................................................................................... 175 10.2. ARTÍCULOS................................................................................................................. 185 21 1. INTRODUCCIÓN 1.1. LA PERICIAL PSICOLÓGICA Y EL ANÁLISIS DE LA CREDIBILIDAD DEL TESTIMONIO En la actualidad, la Ley 1/2000 de Enjuiciamiento Civil (LEC), en sus artículos 335 a 352, regula la intervención de los peritos en el estado español. Asimismo, la Ley de Enjuiciamiento Criminal (LECrim) en su art. 456, hace mención a que el Juez podrá solicitar un informe pericial, cuando no disponga de los conocimientos científicos, técnicos o artísticos necesarios para apreciar determinadas circunstancias del caso. Sin embargo, dicho informe no será vinculante para el dictamen del juez o tribunal, sino que su interpretación, basada en la sana crítica, servirá de apoyo junto con otras pruebas practicadas para valorar libremente el dictamen pericial (De Luca, Navarro, & Cameriere, 2013). Así, la función del psicólogo forense como perito judicial es la de auxiliar a la justicia, asesorando a jueces y tribunales a partir de los conocimientos propios de la disciplina. La jurisprudencia estadounidense de los años 80 no ha sido clara en el establecimiento del papel del psicólogo forense como asistente en tribunales de justicia en casos de abusos sexuales a menores. Se acepta, en determinados estados como Hawai, Montana o Minnesota, entre otros, el testimonio experto sobre la valoración de la credibilidad de la víctima en general (reputation evidence of character), aunque con ciertas limitaciones. Sin embargo, se restringe el papel del perito a determinadas casuísticas: emisión de un juicio sobre la credibilidad en relación con un patrón de características comportamentales de los niños abusados sexualmente; o cuando el testimonio experto mejora la credibilidad de la víctima. En el polo opuesto, ciertos estados rechazan de plano la admisibilidad del testimonio experto por entender que la credibilidad del testimonio es competencia exclusiva del juez y, por tanto, los peritos no deben concluir sobre la culpabilidad o inocencia del acusado; o cuando el testimonio no ayuda de ningún modo al jurado o bien cuando supone un perjuicio. No obstante, ante el creciente número de casos de abuso sexual infantil en la época y el conocimiento puramente profano de jueces y tribunales sobre esta problemática, el foco de atención se dirigió a la búsqueda de formas más eficaces para la verificación de los testimonios de las víctimas (Baker, 1990). Esta necesidad detectada llevó a modificar las reglas tradicionales (Federal Rules of Evidence), con el objetivo de admitir el testimonio experto en casos de abuso sexual infantil, tal y como se recoge en la LECrim española, es decir, cuando el perito experto posea unos conocimientos, habilidades, dilatada experiencia o educación de la que carecen los jueces y tribunales para emitir una opinión sobre determinados comportamientos en relación con el padecimiento de sintomatología clínica (Smith vs. State), e.g. aportar pruebas al jurado sobre el diagnóstico del Trastorno por Estrés Postraumático en víctimas de abuso sexual infantil. Hasta ese momento, sólo se admitía el testimonio experto cuando las pruebas fueran aceptadas generalmente por la comunidad científica, un precepto que, aunque vago, fue incluido posteriormente en los Daubert Standards (se hablará más detalladamente de ellos en el epígrafe siguiente). Empero, los criterios Daubert dejaron en un segundo plano este requerimiento de las reglas tradicionales BÁRBARA GONZÁLEZ AMADO 22 subrayando, como criterio de mayor relevancia, el uso de instrumentos validados empíricamente por la comunidad científica. Por su parte, en el contexto alemán, a comienzos del siglo XX el psicólogo William Stern, fue el primero en acentuar la relevancia de la evaluación de la veracidad de un testimonio, centrándose en las diferencias individuales de las personas declarantes, especialmente en aquellos casos en los que se carecía de evidencias físicas para la incriminación. Sin embargo, la fiabilidad del testimonio recaía, en muchas ocasiones, en elementos indirectos que ponían en duda la credibilidad del individuo previamente al enjuiciamiento. Por ejemplo, un historial de promiscuidad sexual (Undeutsch, 1989) o la minoría de edad (Baker, 1990) a la cual se le presupone una alta propensión a la fantasía (falso testimonio), restaban credibilidad a la declaración de mujeres y menores víctimas de abusos sexuales, respectivamente. Posteriormente, la defensa de Udo Undeutsch de una pericial psicológica en un tribunal de menores marcó un punto de inflexión en la Psicología del Testimonio en general, y en la evaluación de la credibilidad en particular. Con base en esta intervención, el tribunal concluyó que el estudio del testimonio de una víctima hecho por un testigo experto fuera de la sala de justicia lograba resultados sumamente diferentes, en términos de fiabilidad y validez, a los alcanzados durante el juicio oral. Por ello, el Tribunal Supremo alemán dictaminó la necesaria intervención de un perito expert ante la falta de otras evidencias -salvo el testimonio de la víctimapara evaluar la veracidad del testimonio de menores víctimas de abuso sexual. Recientemente, un estudio elaborado por Arce y Fariña (2015), puso de relieve la importancia del testimonio experto. Concretamente, estos autores concluyeron que la evaluación de la credibilidad del testimonio a partir de un mandato judicial, permitió dotar de valor de prueba suficiente a la declaración en el 93.3% de los casos evaluados. Asimismo, encontraron que la mayoría de los casos sobreseídos se produjeron por la ausencia de informes periciales psicológicos, tanto de credibilidad de testimonio que avalaran la realidad de la declaración como del establecimiento de la huella psíquica consecuencia del hecho delictivo sufrido por la víctima. Igualmente, Van Koppen (2007) subrayó la importancia de la asistencia del perito experto en la toma de decisiones en casos penales graves que no son considerados “rutinarios” y que constituyen, aproximadamente, un 12% del total. La valoración de la prueba se convierte, por tanto, en el elemento esencial para la toma de decisiones judiciales, recayendo la motivación de la sentencia, en la fiabilidad y validez de la práctica de la prueba (Arce & Fariña, 2015). La fiabilidad estará condicionada por la credibilidad otorgada a la víctima o testigo e, incluso, por la consistencia interna del propio testimonio, mientras que la validez hace referencia a la relevancia de la prueba para el caso concreto. Cuando la declaración de la víctima constituye la única prueba de cargo, para que ésta sea suficiente, la Jurisprudencia española establece tres criterios jurídicos como apoyo al proceso de toma de decisión en la valoración de la prueba: ausencia de incredibilidad subjetiva (ausencia de motivaciones espurias para declarar), verosimilitud (corroboraciones periféricas objetivas que refuercen el testimonio) y persistencia en la incriminación (entendida como validez de la prueba, consistencia en el tiempo y ausencia de contradicciones). El estudio realizado por Seijo (2007) en el que se analizaron sentencias judiciales, reveló que las decisiones de magistrados sobre la verosimilitud de un testimonio recaían, en gran parte, en informes psicológico forenses que analizaban la credibilidad de la declaración de la víctima y/o el establecimiento del daño psíquico consecuencia de la Introducción 23 victimización. Estos resultados destacan la necesidad de elaborar pruebas objetivas, científicamente validadas, que estudien criterios empíricos de credibilidad, abandonando así los criterios tradicionales subjetivos de los que disponen los jueces para la estimación de la prueba (fiabilidad) en términos de realidad. No existe un consenso respecto de la estructura del informe pericial, sin embargo, los colegios profesiones de psicología han elaborado una serie de directrices (Seijo et al., 2014) para la redacción del contenido de un informe pericial psicológico, basándose en los requerimientos legales anteriormente expuestos: 1. Introducción. Debe incluir los datos de quien emite la pericial, su número de identificación y el número de expediente, en el caso que venga requerido por mandato judicial, así como los datos del demandante/denunciante y persona demandada o denunciada. Igualmente, se incluirá el objeto del informe. Finalmente, el número y fecha de las sesiones de evaluación y una breve mención a la intervención realizada en cada una de ellas. 2. Procedimiento o metodología de evaluación. Este punto constituye una de las partes más relevantes del informe pericial, puesto que en él se explicarán las técnicas y procedimientos utilizados para la evaluación. La mayor o menor credibilidad de nuestra pericial psicológica dependerá del rigor científico y concreción de los instrumentos seleccionados, así como las técnicas de evaluación forense utilizadas. De facto, los resultados y conclusiones del informe quedarán supeditadas a la metodología de evaluación. Por ello, el uso de técnicas forenses metódicas avaladas científicamente (criterio Daubert) permitirán tomar decisiones sobre la evaluación con un elevado nivel de certeza, evitando lo máximo posible la zona de incertidumbre. 3. Resultados. Se expondrán los resultados obtenidos tras la aplicación de las pruebas pertinentes y se incluirá la interpretación de las mismas. 4. Conclusiones. Este apartado es el resultado de la integración de los datos objetivos recabados en las anteriores fases, y deberá responder al mandato formulado por el juez o por alguna de las partes, de forma clara y concisa. Cuando el objeto del informe sea la determinación de la realidad de un testimonio, las conclusiones deben realizarse en términos de probabilidades, ajustándonos a las siguientes categorías: declaración (muy) probablemente cierta/real/creíble/de una memoria basada en hechos vividos; carente de criterios de realidad/de una memoria basada en hechos vividos; declaración o prueba inválida; e indeterminada (también puede referirse como prueba insuficiente (Arce & Fariña, 2006). La exhaustividad y rigurosidad durante todo el proceso de evaluación de la víctima determinará su fiabilidad y validez y, por tanto, su admisibilidad como prueba de cargo en el juicio. Por ello, resulta imprescindible la utilización de técnicas de evaluación en el contexto forense avaladas por la evidencia científica. BÁRBARA GONZÁLEZ AMADO 24 1.2. INSTRUMENTOS DE EVALUACIÓN DE LA REALIDAD DEL TESTIMONIO Los enfoques y herramientas que se han utilizado para la evaluación de la credibilidad del testimonio han sido muy variados. Por un lado, en el paradigma fisiológico, el polígrafo es el instrumento más extendido, especialmente en Norteamérica, donde tiene sus orígenes. El objetivo es detectar cuándo un sospechoso está mintiendo a partir de cambios comportamentales difíciles de controlar de modo consciente (e.g. cambios en el nivel de conductancia de la piel o sudoración, en la presión sanguínea). Esta premisa subyace a las distintas técnicas poligráficas, las más conocidas son: Relevant-Irrelevant Test, Comparison Question Test y el Guilty Knowledge Test. A pesar de su extendido uso a nivel experimental, la utilización del polígrafo en el proceso judicial, en general, es muy reducida, circunscribiéndose a determinados casos y tribunales, por incumplimiento de ciertos criterios Daubert de admisibilidad de pruebas científicas (Griesel & Yuille, 2007). Asimismo, como limitación de la técnica poligráfica, las respuestas fisiológicas pueden distorsionarse a través del entrenamiento y manipulación en las respuestas control, mermando así la precisión del instrumento. Por lo que respecta al paradigma de detección de engaño a través del estudio de señales no verbales, esto es, se relaciona la mentira con determinados movimientos corporales (e.g. movimientos de manos, de cabeza) y características del habla (e.g. titubeos), las investigaciones realizadas revelan la escala fiabilidad de estas medidas (DePaulo, et al., 2003). De facto, no permiten discriminar entre engaño y rasgos de personalidad inherentes al individuo. Tan solo identifican cambios comportamentales entre individuos, en un determinado contexto y, en ciertas ocasiones, el nivel de precisión en la detección se situaría por debajo del azar. No obstante, Vrij, Edward, Roberts y Bull (2000) consideran que la combinación de técnicas verbales y no verbales en la detección del engaño, podría alcanzar una tasa de correcta clasificación (verdaderos positivos y verdaderos negativos) de aproximadamente el 81% de los casos. Los instrumentos que evalúan la realidad del testimonio con base en indicadores verbales han sido los más utilizados y gozan de un mayor respaldo científico en la actualidad (Vrij, Akehurst, Soukara, & Bull, 2004). Concretamente, el Statement Validiy Assessment (SVA) y su componente principal, el Criteria-Based Content Analysis (CBCA). El eje central del CBCA es el análisis del contenido de la declaración y constituye el instrumento de evaluación de la validez del testimonio más estudiado y con mayor aceptación en la práctica forense. Otra herramienta que evalúa aspectos verbales del testimonio es el Reality Monitoring (RM; Johnson & Raye, 1981). La teoría que subyace a este procedimiento tiene su explicación en el proceso cognitivo que guía la capacidad para diferenciar una memoria externa basada en un hecho percibido de una memoria interna, es decir, inventada o imaginada. A diferencia del CBCA, el RM se considera una herramienta para la detección del engaño, mientras que el CBCA constituye un instrumento orientado a la búsqueda de señales de veracidad (truth bias). La investigación experimental en la psicología aplicada al testimonio ha recurrido al RM en numerosas ocasiones, sin embargo, su uso individualizado en la práctica forense es poco frecuente, no así su aplicación combinada con otros sistemas categoriales (Arce & Fariña, 2013; Sporer, 2004), incrementando la capacidad discriminativa o fiabilidad del sistema (Vrij, 2008). El Scientific Content Analysis (SCAN), sistema categorial de análisis de contenido del testimonio desarrollado por Smith (2001), es utilizado por las fuerzas policiales de distintos Introducción 25 continentes (Bogaard, Meijer, Vrij, & Merckelbach, 2016). La teoría memorística que subyace a esta herramienta es equivalente a la denominada Hipótesis Undeutsch: las memorias de un hecho experimentado difieren en contenido y calidad de aquellas basadas en la invención o fantasía. De facto, algunos de los criterios incluidos en el SCAN se corresponden con criterios de realidad del CBCA. La diferencia radica en la unidad de análisis o target, el SCAN evalúa el tipo de lenguaje que utilizan las personas entrevistadas que mienten en contraposición a las utilizadas por las que responden con sinceridad. No obstante, la ausencia de una base teórica que explique la capacidad discriminativa del SCAN así como los resultados contradictorios hallados en los escasos estudios experimentales publicados hasta la fecha (Bogaard et al., 2016), han derivado en una falta de apoyo empírico hacia este instrumento, por parte de la comunidad científica, como herramienta válida para la detección del engaño. 1.3. STATEMENT VALIDITY ASSESSMENT (SVA) 1.3.1. Origen y estructura Si bien originariamente el juicio sobre credibilidad se emitía con base en la personalidad o comportamiento de la víctima, Undeutsch desvió el foco atencional hacia las características del propio testimonio. Con ello se pretendía evitar una posición desfavorable de partida de la víctima y una evaluación puramente subjetiva basada en el sentido común. De este modo, independientemente de si la persona nos inspira confianza o no, la calidad de su declaración recaerá en las características de la propia narración. Para la evaluación su calidad, Undeutsch (1989) propone una serie de “criterios de realidad” que se esperan encontrar con mayor frecuencia en un testimonio real. Este autor se basa en la premisa de que las memorias de hechos vividos o auto-experimentados difieren significativamente de aquellas memorias falsas (inventadas, fabricadas, distorsionadas) siendo aquellas de mayor calidad en términos de criterios. A esta afirmación se le conocerá posteriormente como la Hipótesis Undeutsch (Steller, 1989) y al procedimiento criterial como Statement Reality Analysis (SRA). Este nuevo enfoque fue aceptado por investigadores de distintas partes del mundo, y muchos de ellos han colaborado en la modificación, redefinición y actualización de la propuesta (Undeutsch 1989). El procedimiento se componía de dos partes: una entrevista con preguntas abiertas y el análisis de la declaración mediante un listado de criterios de realidad. Esta estructura y metodología fue sistematizada posteriormente por Steller y Köhnken (1989) dando lugar a lo que hoy en día se conoce como Statement Validity Assessment (SVA), herramienta de evaluación de la realidad de la declaración más ampliamente utilizada en la actualidad (Vrij, 2008). El objetivo del SVA es determinar el origen de la declaración y si ésta contiene aspectos propios de experiencias realmente vividas por la persona entrevistada, sin emitir conclusiones sobre la credibilidad general de la víctima (Raskin & Esplin, 1991). El SVA se compone de cuatro fases: 1) el estudio del caso; 2) una entrevista semi-estructurada, en la que la persona declarante, sin influencias por parte del entrevistador/a, relata el suceso vivido; 3) el análisis de contenido de la declaración a través del Criteria-Based Content Analysis (CBCA); y 4) una combinación de los resultados anteriores con información derivada de un set de preguntas denominado Listado de Validez. BÁRBARA GONZÁLEZ AMADO 26 1) Estudio de caso: previamente a la realización de la entrevista, se estudia minuciosamente la información disponible relativa al caso: características de la víctima (e.g., edad, desarrollo cognitivo), declaraciones previas (con qué frecuencia ha sido entrevistada la víctima, ha habido o no inconsistencias entre las declaraciones) y otras características relacionadas con las motivaciones para testimoniar. Todo ello servirá para analizar el origen de la declaración y los posibles motivos para declarar falsamente. 2) Entrevista semi-estructurada en la que el relator o relatora narre libremente los hechos acaecidos (a partir de técnicas de entrevista que favorezcan el relato libre: Entrevista Cognitiva, Fisher & Geiselman (1992); Step-Wise Inteview Guidelines, Yuille, Marxsen, & Cooper, 1999). La pericia del entrevistador/a es un elemento crucial en la aplicación del SVA, ya que el modo en que se manejan los silencios, la forma en que se realicen las preguntas y las actitudes para con el entrevistado, pueden condicionar sus respuestas. Así, el procedimiento comenzaría con una narración libre de los hechos, seguida de una serie de preguntas abiertas, no dirigidas ni sugestivas, evitando así contaminar la declaración. Esto resulta de vital importancia en el caso de víctimas menores edad, ya que el nivel de sugestionabilidad es superior al de los adultos. La información de caso conjuntamente con los datos aportados por el relator durante la entrevista, son de especial relevancia para el proceso de falsación de hipótesis (Popper, 1968, citado en Köhnken, 2004), es decir, se trata de someter a evaluación las hipótesis alternativas a la hipótesis principal de partida (i.e., la validez de la declaración), procedimiento análogo al diagnóstico diferencial en el ámbito clínico. Las hipótesis alternativas (Köhnken, 2004; Raskin & Esplin, 1991) sometidas a contraste serían: 1) la declaración es completa e intencionalmente falsa, es decir, nunca le ha ocurrido a la persona declarante, con motivaciones de venganza o con la finalidad de ayudar a otra persona; 2) la mayor parte de la declaración es válida, sin embargo, contiene elementos relevantes que han sido inventados como, por ejemplo, a consecuencia de presiones; 3) la declaración es válida, no obstante, se acusa a un individuo distinto al que realmente cometió el delito (transferencia inconsciente); 4) la declaración es un testimonio “prestado”, es decir, no ha sido vivenciado por quien declara pero describe lo que le ha ocurrido a otras personas, como por ejemplo, en los casos de custodia durante un proceso de separación o divorcio; 5) el testimonio es fruto de la fantasía, como consecuencia de una enfermedad mental que le impide discernir entre realidad y ficción; y 6) la declaración es falsa, pero no de modo intencional, sino que puede tratarse de una falsa memoria del evento. El proceso de falsación de hipótesis alternativas comienza en una fase anterior a la entrevista, esto es, durante el análisis minucioso de caso (Köhnken, 2004). En esta fase se analiza la información relativa a la víctima o testigo, que es relevante para el caso y que podría explicar el origen de la narración (i.e., edad, habilidades cognitivas, hecho aislado o repetido en el tiempo, número de entrevistas a las que han sido sometidos la víctima o testigo, inconsistencias entre las declaraciones, etc.). Esta información, que será recuperada de nuevo en el análisis de la lista de validez, es determinante para la evaluación de la declaración como un todo. Finalmente, la totalidad de la entrevista deberá ser grabada en vídeo para proceder, posteriormente, a su transcripción y codificación. 3) Una vez transcrita la entrevista, se aplicará el CBCA con el objetivo de evaluar la realidad de la declaración, a través del análisis de su contenido. Se abordará de forma más pormenorizada el instrumento en un epígrafe posterior. Introducción 27 4) Finalmente, se aplica la Lista de Validez con el fin de conocer la plausibilidad de las hipótesis alternativas planteadas, a partir de un análisis holístico de la información disponible (estudio de caso), de los datos recopilados en la entrevista y del análisis de contenido del testimonio. Llegados a este punto, estaremos en disposición de tomar una decisión razonada sobre la realidad de la declaración, ya que la puntuación obtenida en el CBCA no es suficiente para concluir sobre la veracidad del testimonio (Vrij, 2008). Esta Lista de Validez está conformada por cuatro categorías generales de información que deben ser evaluadas: características psicológicas (e.g. capacidades cognitivas del declarante, susceptibilidad a la sugestión), características relativas a la entrevista (e.g. malas prácticas en el proceso de entrevista), motivaciones para declarar (i.e., presiones para hacer una falsa denuncia, algún tipo de relación entre la víctima y el acusado), otras cuestiones relativas a la investigación que pueden arrojar dudas sobre la validez de la declaración (e.g. descripción de hechos contraria a las leyes de la naturaleza, contradicciones intra e inter-declaración). Cuando la declaración es válida, esto es, de una alta calidad, se utiliza la checklist para desechar las hipótesis alternativas. Si la declaración es de baja calidad, el objetivo del listado es conocer si la información adicional apoya alguna de las hipótesis alternativas, o si la baja calidad de la declaración se debe a una pobre entrevista y/o a una limitada capacidad cognitiva de la víctima (Raskin & Esplin, 1991). 1.3.2. Admisibilidad del SVA en los distintos países La necesidad de evaluar la calidad de un testimonio, en términos de credibilidad, surge a mediados del siglo XX, cuando se puso de manifiesto la labor del perito como experto entrevistador. El perito poseía la capacidad de obtener mayor información y menos sesgada en comparación con la información obtenida en el contexto del juicio oral. El uso del CBCA como herramienta de evaluación de la realidad de un testimonio basado en criterios específicos de contenido es, hasta nuestro tiempo, el único método con apoyo científico aceptado en los tribunales alemanes. Especialmente, cuando se trata de delitos contra la libertad e indemnidad sexuales, ya sea en víctimas adultas o menores de edad (Steller & Böhm, 2006). El panorama en EEUU, como ya se ha indicado, era cualitativamente distinto. La admisibilidad de una prueba pericial de carácter científico requería del cumplimiento de los criterios Daubert o factores de cientificidad (Vázquez-Rojas, 2014): 1) ¿Puede probarse la hipótesis científica? 2) ¿Ha sido probada la hipótesis? 3) ¿Se conoce la tasa de error de la técnica? 4) ¿Ha sido sometida la hipótesis o la técnica a un proceso de revisión por pares y publicación? 5) ¿Se basa la hipótesis y/o técnica en una teoría generalmente aceptada como válida por la comunidad científica? A pesar del establecimiento de esta norma, seguía existiendo una desconfianza hacia la capacidad de los expertos para determinar la calidad de una declaración de abuso sexual con una alta fiabilidad (McGough, 1991) descartando, por tanto, la apreciación de la prueba en la sala de justicia. Los criterios Daubert surgen en el contexto judicial americano a raíz del caso Daubert vs. Merrel Dow Pharmaceuticals (1993). El Tribunal Supremo americano, con el objetivo de no caer en la perspectiva reduccionista de asimilar la cientificidad de la prueba con conocimiento fiable per se, dictaminó una serie de pautas que toda prueba pericial debe cumplir para ser BÁRBARA GONZÁLEZ AMADO 34 ocurre con frecuencia como consecuencia de los interrogatorios judiciales, especialmente en delitos contra la libertad e indemnidad sexuales y de violencia de género (Arce & Fariña, 2007). El entrenamiento en la técnica del CBCA también contribuye a la mejora en su manejo. De facto, el meta-análisis de Hauch, Sporer, Michael, y Meissner (2014) concluyó que el entrenamiento del personal evaluador en señales verbales que indican veracidad y, especialmente, indicadores de contenido de la declaración (en contraposición a señales no verbales de engaño), incrementa las tasas de acierto, esto es, de clasificar correctamente una declaración real, con un tamaño del efecto moderado (gu = 0.733). Del mismo modo, para lograr una declaración suficiente y válida en términos de análisis de contenido, el testimonio debe obtenerse siguiendo los estándares del SVA (Steller, 1989). 1.5. ABUSO SEXUAL EN LA INFANCIA Y ADOLESCENCIA. SECUELAS INTERNALIZANTES DE LA VICTIMIZACIÓN Si bien el CBCA, y por consiguiente el SVA en su conjunto, se erige como una técnica robusta para el análisis de la credibilidad del testimonio, en ocasiones resulta insuficiente como prueba de cargo para tomar una decisión sobre la victimización de los hechos. En este sentido, para la corroboración del testimonio, éste debe sustentarse, no solo en el análisis de su credibilidad sino también en la acreditación de su veracidad mediante la corroboración de circunstancias externas o periféricas (verosimilitud), esto es, el establecimiento de la relación causa-efecto entre el delito y las posibles lesiones (i.e., psicológicas) consecuencia de la victimización. Por ello, resulta de especial relevancia la medida del daño en casos donde el agresor reconoce la relación sexual pero niega que no haya sido consentida por la víctima (Echeburúa, Corral, & Amor, 2002). En respuesta a estas limitaciones prácticas, Arce y Fariña (2006), crearon y validaron una técnica de evaluación psicológico-forense que va más allá del análisis de la realidad del testimonio, denominado Sistema de Evaluación Global (SEG). Esta técnica es el resultado de la compilación y sistematización de diferentes instrumentos y técnicas de evaluación de la credibilidad del testimonio, utilizando la Entrevista Cognitiva Mejorada (Fisher & Geiselman, 1992) para obtener la declaración sobre los hechos en el marco del SVA; del estudio del daño psíquico consecuencia de la victimización, obteniéndose información del estado psicológico a través de la Entrevista Clínico-Forense (Arce & Fariña, 2001), y de la capacidad para prestar testimonio (evaluación de las capacidades cognitivas). El estudio de las propiedades psicométricas de la entrevista clínico-forense (Vilariño, Arce, & Fariña, 2013) mostró una alta fiabilidad tanto para los criterios diagnósticos del daño psicológico (α = .850) como para las estrategias de simulación (α = .744). Asimismo, corroboraron su validez predictiva (el diagnóstico del daño es similar al esperado), convergente en daño (el diagnóstico de daño en la entrevista correlaciona con la evaluación psicométrica de daño) y discriminante para las estrategias de simulación (las víctimas reales no fueron clasificadas como simuladoras). La evaluación de la lesión o daño psíquico se obtiene a través de la medida de los efectos del delito en la salud mental de la víctima, especialmente sintomatología internalizante (e.g. depresión, ansiedad), debiéndose descartar la simulación de síntomas (DSM-V; American Introducción 35 Psychiatric Association, 2013). El estudio de la huella psíquica y su identificación con el delito sufrido permiten establecer la relación causa-efecto entre ambos y, consecuentemente, respaldar y reforzar la veracidad del testimonio de la víctima. El Trastorno por Estrés Postraumático (TEP) se identifica con la huella psíquica propia de delitos como agresiones sexuales o violencia en general. Este trastorno raramente se da de forma aislada, de hecho, entre el 50% y el 60% de los casos en los que se diagnostica TEP, la depresión aparece como trastorno comórbido (Blanchard et al., 2004; O’Donnell, Creamer, & Pattison, 2004). Cuando la declaración es válida y fiable y cuando se determina la relación causa-efecto entre el delito y la huella psicológica, se puede concluir que la víctima presenta una secuela consecuencia de la victimización de los hechos relatados. En el estudio de la huella es crucial descartar otros posibles eventos vitales estresantes como causa de la sintomatología relatada por la víctima. La sintomatología o lesión psicológica asociada a la victimización de delitos de abuso o agresión sexual en la infancia es numerosa y variada (Echeburúa, Corral, Zubizarreta, & Sarasúa, 1995): ansiedad, conductas fóbicas y de evitación, baja autoestima, depresión o distimia, entre otras. Esta sintomatología constituye una medida indirecta del TEP, y puede actuar como potenciadora del diagnóstico de este trastorno (Arce & Fariña, 2005). Sin embargo, estos síntomas tomados de forma aislada no pueden justificarse como prueba de cargo, pues es posible que sean consecuencia de otras vivencias independientes del hecho delictivo (Arce & Fariña, 2005). Es por esto que el TEP se toma como la huella psíquica directamente relacionada con hechos traumáticos, permitiendo el establecimiento de la relación causa-efecto. Por otro lado, el daño puede expresarse a través de sintomatología externalizante que, en muchas ocasiones, se confunde con conductas propias de la vivencia del proceso judicial en el que se encuentra inmersa la víctima. El abuso sexual infantil (ASI) se define como la involucración de un menor en actividades sexuales que no alcanza a comprender o para las cuales está evolutivamente inmaduro y, por tanto, no tiene capacidad para dar su consentimiento expreso (World Health Organization, WHO, 1999). Estudios epidemiológicos y meta-analíticos han puesto de relieve el problema de salud pública que constituyen los abusos sexuales en la infancia y adolescencia, con una alta prevalencia a nivel mundial (Pereda, Guilera, Forns, & GómezBenito, 2009; Stoltenborgh, van Ijzendoorn, Euser, & Bakermans-Kranenburg, 2011). En un intento por acercar posturas a la concepción de minoría de edad cuando nos referimos a abusos sexuales a menores, determinados autores han diferenciado el ASI (< 14 años) del abuso sexual en la adolescencia (ASA, entre 14 y 18 años), identificando dos tramos diferenciados (cut-off). Aunque no hay acuerdo para definir el período de la infancia y la adolescencia, habitualmente se toman los 18 años, o mayoría de edad legal, como edad tope para englobar ASI y ASA. Las secuelas psicológicas, i.e., depresión y ansiedad, consecuencia de ASI/ASA, no solo se manifiestan a corto plazo, sino que, en muchas ocasiones, afectan al funcionamiento en la edad adulta (consecuencias a largo plazo), informando de un desajuste emocional en la adultez, variable en severidad, y que puede llegar a cronificarse (Hillberg, HamiltonGiachritsis, & Dixon, 2011). Diversos estudios han constatado que víctimas de ASI/ASA presentan una mayor probabilidad de sufrir sintomatología internalizante en la adultez en comparación con grupos de no víctimas (Fergusson, McLeod, & Hordwood, 2013; PérezFuentes, et al., 2013; Thompson et al., 2003). Para la evaluación de las lesiones en la salud mental durante la etapa adulta, investigaciones sobre las secuelas a largo plazo del ASI/ASA BÁRBARA GONZÁLEZ AMADO 36 se han valido de métodos de investigación retrospectivos. Estos procedimientos no están exentos de limitaciones (Briere, 1992): dificultades memorísticas para reproducir los abusos sufridos (e.g. la incorporación de información post-suceso que pueda afectar a la exactitud de los recuerdos) así como las dificultades para establecer una relación causal entre la victimización de ASI/ASA y las secuelas psicológicas en la adultez. Sin embargo, medidas de naturaleza retrospectiva tienden a infra-estimar la cifra de abuso sexual (falsos negativos) más que a sobre-estimar, siendo insuficiente este sesgo para invalidar los diseños retrospectivos de casos-controles (Hardt & Rutter, 2004) ya que, en muchas ocasiones, es la única forma de evaluar el abuso sexual. Una aproximación multimétodo (Arce, Fariña, & Vilariño, 2015; Fariña, Arce, Vilariño, & Novo, 2014) al caso concreto de abuso sexual, esto es, la combinación de la entrevista clínico-forense que favorece la aparición espontánea de sintomatología relacionada con la situación traumática, y el uso de instrumentos psicométricos válidos y fiables, que incorporen escalas de validez para el control del engaño, disminuyen el sesgo de la medida retrospectiva. Cuando la víctima tiene menos de 16 años, la victimización de actos de carácter sexual está considerada, en todo caso, un delito (siempre y cuando no se trate de relaciones consentidas con otro u otra menor con similar edad o desarrollo cognitivo). En estos casos, el diagnóstico de TEP para el establecimiento de la relación causa-efecto entre los hechos y las consecuencias psicológicas del delito es secundario, cobrando mayor relevancia la cuantificación de las secuelas clínicas, i.e, depresión y ansiedad, como medida indirecta de la victimización. Puesto que en el ámbito forense, se debe informar de todas las alteraciones pcicológicas que presenta la víctima con el fin de elaborar un perfil de daño moral y evaluar/cuantificar las consecuencias psicológicas consecuencia de los abusos, la identificación de un cuadro clínico de depresión y/o ansiedad en la víctima, constituye un indicador clave de victimización complementario a la medida del TEP y, por tanto, de detección de huella (Arce & Fariña, 2015). Cuando el daño psicológico sufrido en la infancia o adolescencia se prolonga hasta la edad adulta y afecta a diversas esferas de la vida de las víctimas (McKnight & Kashdan, 2009; McKnight, Monfort, Kashdan, Blalock, & Calton, 2016), debe ser operativizado y expresado en el informe pericial. Por ello, la técnica de Arce y Fariña (2009) de evaluación del daño moral en diferentes casuísticas (e.g., accidentes de tráfico), contempla la cuantificación, medida en porcentajes, de la lesión o huella psíquica consecuencia de la victimización. Para ello, se valen de la Escala de Evaluación de la Actividad Global (EEAG, eje V DSM-IV-TR) para evaluar el funcionamiento de la víctima. Todo ello, en el marco del contexto forense, donde debemos sospechar simulación de síntomas como resultado de una motivación espuria dentro del proceso penal (e.g., por cuestiones de venganza). En la última edición del DSM (2013) se eliminó la escala EEAG por considerarla subjetiva y se propuso, en su lugar, el uso del WHODAS 2.0 (World Health Organization Disability Assessment Schedule 2.0), un instrumento clínico, para cuantificar el nivel de incapacidad de la persona sometida a evaluación propuesto por la Organización Mundial de la Salud. Sin embargo, su carácter autoadministrado y su deficiencia para la evaluación de la simulación, lo hacen ineficaz para su uso en el contexto forense. La recomendación de diversos autores es continuar con el uso de la escala EEAG (Gold, 2014), que se había mostrado válida en sus dimensiones (Pedersen & Karterud, 2012). En definitiva, la cuantificación del daño en términos de desajuste psicológico en diferentes escenarios de la vida diaria, no solamente tiene implicaciones a nivel clínico (prevención e intervención), sino también a nivel legal. Introducción 37 1.6. EL PROCEDIMIENTO META-ANALÍTICO Hasta mediados del siglo pasado, la técnica utilizada para integrar resultados derivados de investigaciones primarias era la revisión narrativa o subjetiva. Su objetivo consistía en llegar a conclusiones generalizables, mediante la acumulación e interpretación de datos cualitativos. Sin embargo, su carácter asistemático y su falta de objetividad en el proceso de selección de estudios, las tornaron poco atractivas para su publicación en revistas científicas. Ante la falta de rigurosidad de las revisiones narrativas (la selección de los estudios se hace de forma subjetiva, al igual que la asignación de los pesos a cada artículo original), junto a otras limitaciones anteriormente mencionadas, comenzaron a surgir otros métodos de combinación de resultados que recibieron el nombre genérico de meta-análisis (Glass, 1976). Este autor definió el meta-análisis como el análisis estadístico de los resultados producto de una larga cantidad de investigaciones individuales con el propósito de integrar los hallazgos de todas ellas. Esta definición deja entrever los tres pilares definitorios del método metaanalítico: 1) integrar cuantitativamente los resultados obtenidos en investigaciones primarias sobre una misma problemática; 2) con el propósito de acumular conocimiento científico, de un modo objetivo y sistemático, sobre la temática bajo estudio; 3) a través de la aplicación de métodos estadísticos. Si bien inicialmente el término genérico ‘meta-análisis’ se utilizó para designar a toda revisión sistemática, en la práctica son dos métodos distintos. Todo meta-análisis constituye una revisión sistemática de la literatura, pero no toda revisión sistemática se puede considerar un meta-análisis, ya que la primera puede tener un carácter cualitativo. A colación de esto último, cabe decir que aunque las fases del proceso de ambos métodos son las mismas, el meta-análisis aplica técnicas estadísticas para integrar los resultados de las investigaciones empíricas y cuantificar el efecto de la relación bajo estudio. Las fases a seguir en la realización de un meta-análisis las desarrolla Sánchez-Meca (2010): 1. Formulación del problema. Al igual que ocurre en una investigación primaria, inicialmente se debe estipular el problema objeto de estudio. Para ello, nos formularemos una serie de preguntas que deben ser respondidas a partir de los resultados obtenidos. Una vez que tenemos claro cuál es el foco de nuestra investigación, se establecerán los objetivos a alcanzar, así como las hipótesis de partida. 2. Búsqueda de la literatura científica. Constituye la fase más relevante y laboriosa del proceso. Una vez tengamos clara la pregunta a responder, se hará una búsqueda, lo más exhaustiva posible, de la literatura científica para identificar aquellos estudios de carácter cuantitativo o empíricos cuyo objetivo se corresponda con el nuestro. Para ello, a partir de una serie de palabras clave que utilizaremos como motor de búsqueda, haremos uso de bases de datos electrónicas, así como de meta-buscadores para la identificación de estudios. Asimismo, se harán búsquedas manuales en libros y revistas científicas, se rastrearán las referencias bibliográficas de artículos primarios ya seleccionados para identificar otros potenciales estudios más antiguos (ancestry approach), y se contactará con investigadores especialistas en la temática bajo estudio con el objetivo de identificar artículos no publicados. Con motivo de ajustar la búsqueda, además del uso de palabras clave, se establecerán unos criterios BÁRBARA GONZÁLEZ AMADO 38 de inclusión (e.g., características de la muestra, criterio temporal de publicación en revistas científicas) que deben satisfacer los artículos para ser seleccionados, así como unos criterios de exclusión (e.g., artículos con deficiencias metodológicas o artículos de baja calidad) que les dejarían fuera del meta-análisis. Esta fase debe quedar descrita de forma detallada y rigurosa para que pueda ser sometida a réplica por parte de otros autores. 3. Codificación de los estudios primarios. Una vez seleccionados los estudios que formarán parte del meta-análisis, se elaborará un cuadro descriptivo en el que se recogerán las distintas características de las investigaciones primarias (e.g. tamaño de la muestra, existencia o no de grupos control, medida del efecto). La codificación de la información deberá llevarse a cabo por al menos dos jueces independientes para garantizar la calidad de dicha codificación, mediante el cálculo de la fiabilidad interjueces (Botella & Gambara, 2002). El número y tipo de características sometidas a codificación dependerá del área de investigación y de la información detallada por las investigaciones previas. La utilidad de la realización de un cuadro base de los estudios es la identificación de variables moderadoras o mediadoras que puedan incidir en los resultados y expliquen la variabiliad observada en los datos. 4. Cálculo del tamaño del efecto. Al mismo tiempo que codificamos los estudios, debemos convertir los índices estadísticos utilizados en cada estudio primario (e.g., medias y desviaciones típicas, t para la diferencia de medias, prueba Chi-cuadrado, proporciones) a una métrica común: el tamaño del efecto. El tamaño del efecto es un índice estadístico que nos informa en qué medida el fenómeno que estamos investigando está presente en la población o, lo que es lo mismo, el grado en que la hipótesis nula es falsa. El tamaño del efecto puede ser expresado en distintos índices, siendo los más comunes la d de Cohen y la correlación de Pearson (r). Las fórmulas para la conversión de los diferentes índices estadísticos a tamaños del efecto varían atendiendo a la relación existente entre ambos. Según Rosenthal (1984), la relación entre prueba de significación y tamaño del efecto responde a la siguiente ecuación general: Prueba de significación = Tamaño del efecto x Tamaño del estudio Las fórmulas para el cálculo del tamaño del efecto (Cohen, 1988; Rosenthal, 1984) deben seleccionarse meticulosamente, en función las pruebas estadísticas disponibles en los estudios primarios (e.g., medias y desviaciones típicas, t para una muestra, t para muestras relacionadas, tabla de contingencia 2X2, F con un grado de liberad en el numerador). 5. Análisis estadístico e interpretación. Finalizado el proceso de codificación y el cálculo del tamaño del efecto para cada estudio, se procederá a elaborar una base de datos para cada investigación primaria. Previamente a la realización del meta-análisis propiamente dicho, algunos autores aconsejan rastrear los datos incorporados en la base de datos en busca de valores outliers o extremos que puedan condicionar la variabilidad observada entre los estudios. Durante el proceso de depuración, debemos ser cautos ya que podemos estar eliminando estudios moderadores con apariencia de valores outliers. Las técnicas de análisis estadístico varían en función del procedimiento meta-analítico. Las modalidades de revisiones meta-analíticas se abordarán con mayor profundidad en un apartado posterior. Introducción 39 Por lo que respecta a la interpretación de los efectos obtenidos, no existen unas reglas absolutas para interpretar la magnitud del fenómeno, ya que depende del área de investigación en la que se encuadre nuestro estudio. Sin embargo, en la comunidad científica se aceptan unos valores establecidos convencionalmente que varían según el índice de tamaño del efecto que hayamos utilizado. Concretamente, ciñéndonos a la clasificación establecida por Cohen (1988), hablaremos de un tamaño del efecto pequeño (d = 0.20; r = .10), moderado (d = 0.50; r = .30) o grande (d = .80; r = .50). Adicionalmente, en los artículos meta-analíticos aquí presentados, se han calculado otros índices que facilitan al lector la interpretación de los efectos hallados: i. Binominal Effect Size Display (BESD). Se trata de un índice que permite conocer la importancia o utilidad práctica del efecto estimado a partir de una diferencia de proporciones. El BESD informa del efecto de una variable predictora sobre la tasa de éxito o mejora en la variable criterio, y se expresa como la diferencia entre el porcentaje de éxito de la tasa base y el porcentaje de éxito tras la aplicación de la variable a medir (Rosenthal, 1984). En nuestro caso concreto, permite calcular la probabilidad de falsos positivos (identificar como víctima a quien no lo ha sido) y falsos negativos (no identificar a víctimas reales). ii. Probabilidad de superioridad (PS). También denominado Common Language Effect Size Statistic (CLES) (McGraw & Wong, 1992), hace referencia a la probabilidad de que un individuo extraído al azar de una población obtenga un valor superior que otro individuo, perteneciente a otra población, extraído del mismo modo. En nuestro caso concreto, el cálculo de PS nos permitió conocer en porcentaje de declaraciones sobre hechos autoexperimentados que contenían más criterios de realidad que las declaraciones fabricadas. iii. Estadístico U1 (Cohen, 1988). Medida del área, expresada en porcentaje, que no se superpone entre dos distribuciones poblaciones. Es decir, el porcentaje de no-superposición (100% – U1) indica en qué medida un experimento o intervención han tenido un efecto separador entre las dos puntuaciones o poblaciones de interés. Así, cuando d = 0.00, U1 = 0.00%, las dos poblaciones se superponen al 100% lo que indicaría que ambas poblaciones son idénticas. Estos dos últimos índices ayudan al lector a comprender mejor la relación entre las distribuciones de las condiciones estudiadas, y se recomienda su inclusión en el informe meta-analítico conjuntamente al del tamaño del efecto para aquellos resultados más relevantes (Fritz, Morris, & Richler; 2012). 6. Publicación del estudio meta-analítico. La culminación de la revisión meta-analítica tiene lugar con su publicación en una revista científica. Puesto que estamos ante una investigación empírica, los apartados que debe contener el artículo son los mismos que los de una investigación primaria: introducción, método, resultados y discusión. BÁRBARA GONZÁLEZ AMADO 40 Como ya se explicó anteriormente, la revisión sistemática y el meta-análisis siguen el mismo proceso de elaboración. Sin embargo, a diferencia de lo que ocurre en el meta-análisis, las etapas relativas al cálculo del tamaño del efecto y al análisis estadístico no se llevan a cabo en una revisión sistemática. 1.6.1. Tipos de meta-análisis Los modelos estadísticos para la combinación de resultados que más atención han recibido en la teoría meta-analítica han sido los modelos de efectos fijos y de efectos aleatorios. El primero de ellos asume que no existe heterogeneidad entre los estudios primarios incluidos en la revisión meta-analítica, de modo que si se observa variabilidad, ésta se debe exclusivamente al error de muestreo intra-estudio. Por otro lado, en el modelo de efectos aleatorios se acepta la posibilidad de que los parámetros poblacionales varíen entre estudios, esto es, se asume a priori una heterogeneidad intra e inter-estudios. Existen diferentes métodos meta-analíticos en función del índice estadístico que se utilice para la acumulación de resultados. Hunter y Schmidt (2015) se centran en aquellos que utilizan correlaciones o tamaños del efecto, dejando de lado las metodologías basadas en los valores del test de significación (p). Atendiendo a esta distinción, Hunter y Schmidt (2015) proponen la siguiente clasificación de métodos meta-analíticos: puramente descriptivos, cuyo objetivo es dibujar el panorama sobre una determinada cuestión general; aquellos que sólo tienen en cuenta el error de muestreo (i.e., bare-bones meta-análisis); y métodos que tienen en cuenta y corrigen por el error de muestreo y otros artefactos (e.g., meta-análisis psicométrico). El meta-análisis psicométrico y el bare-bones meta-análisis, metodologías elaboradas por Hunter y Schmidt, se enmarcan en los modelos de efectos aleatorios. Estos modelos asumen que los parámetros poblacionales pueden variar entre estudios y que el propósito último es estimar dicha variabilidad. De hecho, el modelo de base es sustractivo, esto es, la varianza poblacional estimada es la varianza que resulta de la eliminación del error de muestreo y otros artefactos (Hunter & Schmidt, 2015). Según estos autores, el procedimiento denominado bare-bones es incompleto e insatisfactorio, debido a que solamente corrige el tamaño del efecto estimado por el error de muestreo cuando es habitual que existan otros errores artifactuales que alteran el valor de las medidas. Como consecuencia, Hunter y Schmidt proponen los modelos psicométricos como una alternativa más robusta, puesto que permiten conocer si la variabilidad observada entre los estudios se debe a la presencia de artefactos (e.g. error de muestreo, error de medida, restricción en el rango, dicotomización de una variable) o si se trata de variabilidad real. Además, los artefactos crean sesgos a la baja, es decir, infravaloran los resultados del meta-análisis lo que se traduce en una estimación del tamaño del efecto inferior al que es en realidad. Los métodos de meta-análisis psicométrico tienen dos variantes: a) Meta-analysis of correlations or experimental effect sizes corrected individually for artifacts: Meta-análisis que corrige individualmente cada r o d por los diferentes artefactos. Con frecuencia, no disponemos de información relativa a artefactos en todos los estudios primarios que van a formar parte del meta-análisis. Por ello, los autores plantean una segunda tipología de meta-análisis psicométrico: Introducción 41 b) Meta-analysis of correlations or experimental effects sizes using artifact distributions. Meta-análisis que corrige el tamaño del efecto por la distribución de artefactos. Ante la falta de información requerida para corregir individualmente cada estudio, se opta por corregir a partir de una distribución de los artefactos disponibles. Las revisiones meta-analíticas de las que consta esta tesis, siguen este modelo de procedimiento. Hunter y Schmidt (2015) identificaron una serie de errores artifactuales que alteran el efecto observado en comparación con el efecto real o verdadero: 1) Error de muestreo: está relacionado con el tamaño de la muestra, esto es, a menor tamaño muestral, mayor será el error de muestreo. Este artefacto provoca que el tamaño del efecto estimado varíe aleatoriamente del valor de la población. 2) Error de medida: puede observarse tanto en la variable dependiente (criterio) como en la variable predictora o independiente: se trata del error de medida aleatorio provocado por la falta de fiabilidad (unreliability) del instrumento de medida. Cuando existe este error el tamaño del efecto observado es sistemáticamente menor que la validez operativa. 3) Dicotomización de una variable continua: consiste en transformar los distintos valores de una variable continua a sólo dos, y puede ocurrir tanto en la variable independiente como en la dependiente. La consecuencia sería una atenuación del tamaño del efecto verdadero en comparación con el que obtendríamos con la variable continua. 4) Variaciones en el rango: cuanta menos variabilidad exista en la población, es decir, cuanto más homogéneas sean las puntuaciones, mayor será la restricción en el rango y menor el tamaño del efecto observado en relación con el efecto verdadero. Afecta a las variables predictora y criterio. 5) Errores de información o transcripción: hace referencia a los fallos cometidos durante el volcado de información en una base de datos, el uso de fórmulas incompletas o incorrectas (errores tipográficos), cuando hacemos una transcripción equivocada o la inexactitud de información al recuperarla de la memoria. Son errores frecuentes y, probablemente, los más difíciles de detectar y corregir posteriormente. 6) Varianza debida a factores extraños: se entiende por factores extraños aquellos casos que se desvían significativamente del resto y que pueden hacer variar los resultados. Se les conoce comúnmente como valores outliers y la dificultad estriba en diferenciarlos de valores extremos que son propios de la población de referencia. Al igual que cualquier otro procedimiento de análisis estadístico, las técnicas metaanalíticas no están exentas de inconvenientes. Primero, existe una alta propensión a cometer errores cuando incorporamos a nuestro meta-análisis investigaciones primarias con sesgos como, por ejemplo, errores cometidos durante el registro de los datos o en el proceso de análisis de los mismos. Segundo, la limitación anterior está íntimamente ligada con la calidad final de los datos. Estudios empíricos con deficiencias metodológicas pueden sesgar los resultados, por ello, ciertos autores son partidarios de eliminar a priori las investigaciones de baja calidad, esto es, tomando la calidad metodológica del estudio como un criterio de BÁRBARA GONZÁLEZ AMADO 42 inclusión. Sin embargo, en el otro polo, se apoya la inclusión inicial de todos los estudios empíricos seleccionados, independientemente de la calidad de sus datos y, sólo a posteriori, tomar la calidad de cada estudio como variable moderadora (Hunter & Schmidt, 2015) o bien se someten los resultados a un análisis de sensibilidad. Tercero, resulta imposible abarcar todo el material empírico existente sobre la temática de interés, ya sea porque no se han localizado algunos estudios con la metodología de búsqueda bibliográfica utilizada, o por las dificultades que entraña la localización de investigaciones no publicadas. Esta falta de representatividad de los estudios localizados en relación a la totalidad de los estudios existentes, nos puede llevar a incurrir en el denominado sesgo de publicación (File-Drawer Problem, FDA), constituyendo una amenaza a la validez de las conclusiones. Se recomienda, por tanto, el cálculo del número de seguridad o tolerancia, es decir, el número de estudios con resultado no significativo que serían necesarios para invertir la conclusión del efecto obtenido. De este modo se obtendrá una aproximación de la amenaza del sesgo de publicación a la validez de las conclusiones alcanzadas. Cuarto, si se toma más de un tamaño del efecto calculado a partir de la misma muestra, no garantizamos la independencia estadística de los datos, lo que conlleva una disminución de la fiabilidad de los resultados. Ante esta situación, se recomienda hallar un promedio de los tamaños del efecto del estudio correspondiente (Sánchez-Meca, 1986). Quinto, la inclusión de estudios muy heterogéneos entre sí, es decir, difícilmente comparables, pone en entredicho la generalización de los resultados. Por otro lado, las técnicas meta-analíticas poseen una serie de ventajas frente a las revisiones narrativas y a las investigaciones experimentales primarias. En primer lugar, se trata de un procedimiento eficiente en sí mismo, pues permite integrar los resultados de un gran número de estudios que versan sobre una misma temática de forma sistemática y objetiva, permitiendo así su replicación. En definitiva, el meta-análisis goza del mismo rigor científico que un estudio empírico, al exigírsele las mismas normas. En segundo lugar, la precisión (en términos de fiabilidad y validez) de las conclusiones es superior a la de las revisiones cualitativas, dotando al meta-análisis de capacidad para detectar pequeños efectos. En tercer lugar, ostentan una mayor potencia empírica al combinar los resultados de los artículos, aumentando el tamaño muestral y, por ende, la potencia estadística de la prueba aplicada (Sánchez-Meca, 1986). En cuarto lugar, no solamente permite conocer el estado de la cuestión de la problemática estudiada, sino que también detecta áreas de incertidumbre que es necesario abordar en futuras líneas de investigación. Por último, la existencia de resultados contradictorios en la literatura previa pueden ser abordados mediante el estudio de variables moderadoras, las cuales podrían explicar esas discrepancias inter-estudios. 43 2. OBJETIVOS La elección del procedimiento meta-analítico como el método estadístico más adecuado para la consecución de los objetivos establecidos en esta tesis doctoral radica en las ventajas y facilidades del propio método. A continuación se exponen, artículo por artículo, los objetivos específicos abordados por la revisión meta-analítica: 1. La investigación científica, a nivel cualitativo y cuantitativo, ha arrojado resultados aparentemente contradictorios e interpretaciones dispares sobre la capacidad del CBCA como herramienta válida para discriminar entre memorias de hechos autoexperimentados y memorias de hechos inventados o fabricados en poblaciones de menores. Asimismo, Vrij (2008) dejó patente en su revisión cualitativa de la literatura, la falta de consenso entre la comunidad científica (criterio Daubert número 5), en la aceptación del CBCA como prueba científica para ser admitida en los tribunales de justicia. Es decir, se ha puesto en entredicho que el CBCA cumpla los requerimientos legales de admisibilidad de toda prueba científica en la justicia (Daubert Standards y criterios jurisprudenciales). Ante este estado de la cuestión, la técnica meta-analítica fue la más pertinente para: a) acercar posiciones entre la comunidad científica sobre la aplicabilidad de la técnica de análisis criterial llegando a conclusiones generalizables a la práctica forense a partir de la interpretación de los datos cuantitativos; b) validar el CBCA como instrumento con capacidad para discriminar entre memorias de hechos auto-experimentados y fabricados (cuantificando el efecto de la relación bajo estudio) y, por consiguiente, afianzar su aplicabilidad en contextos forenses y en diferentes casuísticas (a partir de la generalización de los resultados); c) poner a prueba la cientificidad de la técnica a partir del estudio del cumplimiento de los criterios Daubert; d) contribuir al desarrollo de la teoría del testimonio, concretamente al estudio de la credibilidad a partir del análisis del contenido de la declaración, gracias al aporte de la técnica meta-analítica como instrumento de acumulación sistemática de conocimiento científico; d) proponer futuras líneas de investigación que aporten nuevo conocimiento sobre la materia bajo estudio. 2. La utilización de las técnicas de análisis de contenido del testimonio ha sufrido un aumento exponencial en la justicia española, especialmente cuando se trata de evaluar el testimonio de víctimas adultas. La ausencia de validación del CBCA con población adulta, ha generado la reactancia de ciertos sectores del ámbito judicial y forense para su aplicación en dicha población y en contextos distintos para que el que fue creado en origen (i.e., menores víctimas de abuso sexual). Este panorama puso en duda al CBCA como método categorial válido para ser aplicado en el contexto judicial, desamparando a las víctimas ante la ausencia de una metodología BÁRBARA GONZÁLEZ AMADO 50 Tabla 2. Estudio de la eficacia discriminativa diferencial de los criterios de realidad y del total, en muestras de menores y adultos Criterios k1 k2 N1 N2 δ1 δ2 DE1 DE2 ICδ1 95% ICδ2 95% Estructura lógica 16 30 1381 2265 0.52 0.62 0.272 0.849 [0.41, 0.63] [0.53, 0.70] Producción inestructurada 15 27 1217 1987 0.53 0.69 0.588 1.155 [0.42, 0.64] [0.60, 0.78] Cantidad de detalles 17 35 1477 2714 0.87 0.71 0.594 1.028 [0.76, 0.97] [0.63, 0.78] Engranaje contextual 15 29 1341 2137 0.78 0.24 0.571 0.741 [0.67, 0.89] [0.15, 0.32] Descripción de interacciones 16 29 1407 2243 0.50 0.36 0.434 0.379 [0.39, 0.60] [0.28, 0.44] Reproduction de conversaciones 16 34 1407 2528 0.59 0.44 0.415 0.607 [0.48, 0.69] [0.36, 0.52] Complicaciones inesperadas 11 29 1111 1956 0.33 0.32 0.000 0.370 [0.21, 0.45] [0.23, 0.41] Detalles inusuales 16 35 1437 2441 0.31 0.41 0.335 0.786 [0.21, 0.41] [0.33, 0.49] Detalles superfluos 13 27 1199 1863 0.47 0.18 0.274 0.667 [0.35, 0.58] [0.09, 0.27] Detalles incomprendidos 13 5 1062 376 0.35 0.28 0.385 0.000 [0.23, 0.47] [0.08, 0.48] Asociaciones externas 10 22 916 1612 0.32 0.34 0.328 0.538 [0.19, 0.45] [0.24, 0.44] Estado mental subjetivo 15 28 1194 2170 0.52 0.23 0.419 0.554 [0.40, 0.63] [0.14, 0.31] Estado mental autor del delito 10 31 1052 2232 0.21 0.11 0.222 0.747 [0.09, 0.33] [0.03, 0.19] Correciones espontáneas 15 29 1367 1842 0.23 0.20 0.383 0.601 [0.12, 0.33] [0.07, 0.33] Admisión de falta de memoria 13 34 1076 2305 0.17 0.32 0.318 0.377 [0.05, 0.29] [0.24, 0.40] Dudas propio testimonio 10 26 809 1755 0.22 0.26 0.238 0.492 [0.08, 0.36] [0.16, 0.35] Autodesaprobación 5 13 447 948 0.18 0.05 0.480 0.519 [-0.01, 0.37] [-0.08, 0.18] Perdón al autor del delito 6 8 517 680 0.25 -0.02 0.343 0.228 [0.08, 0.42] [-0.17, 0.13] Detalles característicos agresión 5 5 318 562 1.40 0.36 0.807 0.000 [1.15, 1.64] [0.19, 0.53] Total 18 31 1122 2124 0.79 0.56 0.275 0.638 [0.67, 0.91] [0.47, 0.65] Note. k1 = número de estudios con muestras de menores; k2 = número de estudios con muestras de adultos; N1 = tamaño de la muestra en menores; N2 = tamaño de la muestra en adultos; δ1 = tamaño del efecto en menores corregido por falta de fiabilidad en el criterio; δ2 = tamaño del efecto en adultos corregido por falta de fiabilidad en el cirterio; DE1 = desviación estándar de δ1; DE2 = desviación estándar de δ2; ICδ195% = Intervalo de confianza al 95% para δ1; CIδ2 95% = Intervalo de confianza al 95% para δ2 Resultados 51 3. Los criterios adicionales ‘estilo de la declaración’ y ‘dar muestras de inseguridad’, puestos a prueba en el meta-análisis con población adulta, mostraron un tamaño del efecto corregido positivo (δ = 0.48 y δ = 0.78, respectivamente) significativo y generalizable. No obstante, el criterio ‘dar explicaciones de la falta de memoria’ no resultó productivo (el intervalo de confianza contiene el valor cero). Finalmente, los restantes criterios adicionales analizados arrojaron un tamaño del efecto negativo, aunque sólo fue significativo para el criterio ‘repeticiones’ (δ = -0.54), apuntando que no se trata de un criterio de realidad, sino un indicador de engaño. 4. En tanto que la literatura analizada advertía de la existencia de características en los estudios primarios que podrían afectar a los resultados hallados, se llevaron a cabo nuevos meta-análisis con el objetivo de conocer cómo afectaban al efecto estimado. La primera de ellas y, desde el punto de vista de la práctica forense, la más relevante, se refiere al paradigma de investigación (investigación de campo vs. investigación experimental de laboratorio). 4.1. El tamaño del efecto hallado en los estudios de campo (tomando la puntuación total del CBCA) con muestras de menores fue positivo, significativo, generalizable, y de una magnitud más que grande δ = 2.71; p < .01. Igualmente, con la población de adultos se obtuvo un tamaño del efecto positivo y moderado (δ = 0.69), significativo y generalizable, con el promedio de criterios significativos (los criterios realmente discriminativos entre memorias de hechos vividos y fabricados son los que formarán parte del sistema categorial) como variable dependiente. En resumen, en estudios de campo, el total del instrumento y los criterios de realidad significativos, discriminan significativamente entre memorias de hechos autoexperimentados y memorias fabricadas, tanto en muestras de menores como en adultos, respectivamente. 4.2. Los resultados encontrados para el meta-análisis de estudios experimentales revela un tamaño del efecto positivo y significativo en ambos tipos de muestras (menores, δ = 0.56; y adultos, δ = 0.32), no resultando generalizable a otras poblaciones cuando se aplica el CBCA a adultos. Tomando en consideración las directrices de Steller (1989) para incrementar la validez ecológica de los paradigmas experimentales, se seleccionaron aquellos estudios primarios que cumplieran las condiciones de ‘high fidelity’ (involucración personal en los hechos, con un carácter emocional negativo y que implicaran una pérdida de control respecto del evento) y se llevó a cabo un nuevo meta-análisis. Los resultados hallados fueron los mismos que tomando todos los estudios experimentales (δ = 0.58 vs. δ = 0.56, respectivamente) en la población de menores. 4.3. Igualmente, se ha investigado el efecto del contexto en la capacidad discriminativa de los criterios de realidad en población de adultos. Para ello, se realizó un nuevo meta-análisis tomando únicamente los estudios de campo en casos de violencia sexual y de género (delitos cometidos en la esfera privada). Los resultados revelaron (con el promedio de criterios significativos) un tamaño del efecto elevado (δ = 0.96), positivo, BÁRBARA GONZÁLEZ AMADO 52 significativo y generalizable; siendo el efecto más pequeño esperable de δ = 0.64. La diferencia de tamaños del efecto, entre la condición de estudios de campo y de abuso sexual y violencia de género (con los criterios significativos) advierten de una mayor capacidad discriminativa de los criterios de realidad bajo la condición de contexto, qc = .168, p < .05 (0.69 vs. 0.96). 5. El análisis del impacto de otras variables en los resultados nos llevó a proceder con el cálculo de distintos meta-análisis (se tomó el promedio de criterios como variable dependiente) seleccionando como moderadores el criterio de publicación Daubert Estándar en revistas con proceso de revisión por pares (DSCP), la versión del sistema categorial (todos los criterios de realidad vs versión de 14 criterios, estatus del declarante, y tipo de evento [eventos auto-experimentados y eventos observados en vídeo]). En todos ellos se obtuvo un tamaño del efecto positivo y significativo, pero no generalizable (es necesario continuar con la búsqueda de moderadores). Los efectos hallados fueron todos pequeños (0.20 > δ < 0.50), a excepción del meta-análisis para los eventos observados en video (testigo) cuya magnitud fue moderada (δ = 0.51). En los anexos pueden consultarse los resultados de los meta-análisis (para cada criterio individual del CBCA, criterios adicionales, total del instrumento, y promedio) de cada variable moderadora en muestras de adultos (tablas 3-11). 6. El estudio del testimonio de agresores valida de nuevo la hipótesis Undeutsch: todos los criterios de realidad fueron positivos. No obstante, los criterios ‘detalles inusuales’, ‘detalles incomprendidos relatados con precisión’, ‘asociaciones externas relacionadas’, ‘estado mental subjetivo’, ‘correcciones espontáneas’, ‘dudas sobre el propio testimonio’, ‘auto-desaprobación’ y ‘perdón al autor del delito’ no fueron productivos y, por tanto, no pueden aplicarse a muestras de agresores. 7. La victimización de abuso sexual reveló un efecto global moderado (ρ = .34), en el padecimiento de secuelas psicológicas (depresión y ansiedad), significativo y generalizable. Concretamente, se obtuvo un tamaño del efecto pequeño y otro moderado en padecimiento de sintomatología depresiva (ρ = .28) y ansiosa (ρ = .31), respectivamente. 8. Los análisis realizados atendiendo al género de la víctima, revelaron un efecto pequeño, significativo y generalizable, para las mujeres, en la manifestación de depresión y ansiedad (ρ = .22). Sin embargo, para los varones se obtuvo un tamaño del efecto pequeño, positivo, significativo y generalizable en la medida de la depresión (ρ = .13), pero no generalizable a otras muestras en la medida de ansiedad (ρ = .15) (el intervalo de credibilidad contiene el cero). 9. Se llevaron a cabo, además, distintos meta-análisis tomando como variable moderadora el tipo de medida del daño psicológico, esto es, diagnóstico o sintomatología. Los resultados hallados informan de un tamaño del efecto moderado, positivo, significativo y generalizable en el diagnóstico de depresión (ρ = .31) y ansiedad (ρ = .35). Asimismo, la victimización de ASI/ASA también acarrea un desarrollo significativo, positivo, generalizable, y de un tamaño del efecto pequeño, de sintomatología (ρ = .21) ansiosa y depresiva. Resultados 53 10. Tomando en consideración la epidemiología de los trastornos depresivos y los trastornos de ansiedad para cada género (DSM-V), se identificó como moderador la interacción entre el tipo de medida del daño y el género de la víctima. 10.1. Los resultados mostraron un efecto positivo, significativo y generalizable para el diagnóstico de depresión (ρ = .42) y ansiedad (ρ = .24) en mujeres. El efecto hallado para los varones en ambos diagnósticos también resultó positivo y significativo (aunque de magnitud pequeña), pero no generalizable. Comparativamente, las mujeres víctimas de ASI/ASA son diagnosticadas, de forma significativa, de un trastorno depresivo, qs = 0.388, p < .01, así como de un trastorno de ansiedad qs = 0.104, p < .05, en mayor medida que los varones. 10.2. Los efectos hallados para la medida de sintomatología han sido positivos, de un tamaño del efecto pequeño, significativos y generalizables en todos los casos. Concretamente, varones y mujeres presentan efectos similares en el padecimiento de sintomatología depresiva (ρ = .22 y ρ = .20, respectivamente). No obstante, los varones refieren significativamente más sintomatología ansiosa que las mujeres, qs = 0.095, p < .05. 11. Los meta-análisis realizados tomando el tipo de abuso o severidad (sin contacto, con contacto y con penetración) como variable moderadora, resultaron en un tamaño del efecto significativo, positivo, de un tamaño pequeño y generalizable, en depresión y ansiedad. La comparación entre los tamaños del efecto mostró que el daño en el abuso con penetración (ρ = .19 en depresión, ρ = 0.15 en ansiedad), era significativamente mayor que en la condición de abuso sin contacto (ρ = .12, ρ = .08), tanto para depresión, qs = 0.093, p < .05; como para ansiedad, qs = 0.092, p < .05 55 5. DISCUSIÓN De los resultados hallados en los estudios meta-analíticos, se han extraído las siguientes conclusiones. En memorias de menores: 1. Se validó la Hipótesis Undeutsch en memorias de menores a partir del estudio de los criterios de realidad del CBCA. Esto es, el CBCA como sistema categorial para el análisis de contenido de declaraciones de menores discriminó significativamente entre memorias de hechos vividos y fabricados. Esto se traduce en un mayor número de criterios de realidad en declaraciones basadas en hechos reales frente a las declaraciones inventadas o fabricadas. Por tanto, los resultados apoyan la Hipótesis Undeutsch para niños de todas las edades y contextos (i.e., para contextos ajenos al abuso sexual). 2. La puntuación total en el CBCA discriminó entre declaraciones reales y fabricadas, independientemente del paradigma de investigación (estudios experimentales vs. campo). Sin embargo, la eficacia del sistema resultó significativamente superior en los estudios de campo que en los experimentales. En consecuencia, los resultados de estudios experimentales no son directamente generalizables para la práctica salvo que estén apoyados por estudios de campo. 3. Todos los criterios de realidad del CBCA discriminaron significativamente entre memorias de hechos vividos y fabricados. Esto es, los resultados validaron todos los criterios de realidad del CBCA. 4. Los resultados revelaron tamaños del efecto más elevados para los criterios de realidad del componente cognitivo (criterios 1 al 13 del CBCA) que para el componente motivacional (criterios 14 al 18). Además, los criterios motivacionales no fueron generalizables. 5. La primera gran categoría “características generales” i.e., los criterios ‘estructura lógica’, ‘producción inestructurada’, ‘cantidad de detalles’, han obtenido los tamaños del efecto más grandes. El tamaño del efecto más elevado fue para el criterio ‘detalles característicos de la agresión’. Una explicación radicaría en la dificultad para inventar estos detalles. 6. En cuanto a la práctica forense, la puntuación total en el CBCA clasifica correctamente al 68.5% de las memorias de hechos auto-experimentados (verdaderos positivos), mientras que falla en la detección del 31.5% de dichas memorias (falsos negativos). Respecto a los verdaderos negativos no pueden generalizarse a la práctica puesto que los criterios de realidad no clasifican BÁRBARA GONZÁLEZ AMADO 56 memorias falsas/fabricadas. En relación con los falsos positivos (declaraciones falsas identificadas como verdaderas), que en el campo forense debe ser el 0% puesto que transgrede el principio de presunción de inocencia (i.e., el Tribunal Constitucional y el Tribunal Supremo españoles han manifestado que ninguna persona inocente debe estar en prisión, mientras que algunas personas que son culpables puede estar fuera de prisión), la probabilidad potencial de falsos positivos es de .288 (1-.712, es decir, uno menos la probabilidad de obtener más criterios de realidad del CBCA en testimonios de menores relativos a hechos vividos). 7. El análisis de la variable moderadora paradigma de investigación mostró, con la puntuación total del CBCA, un tamaño del efecto y una correcta clasificación de los verdaderos positivos superior en los estudios de campo (90.2%) que en los experimentales (63.5%). Por consiguiente, los resultados derivados de los estudios experimentales no pueden generalizarse directamente a la práctica, pues solamente gozan de validez aparente (Fariña, Arce, & Real, 1994; Konecni & Ebbesen, 1992). Esto corrige también la probabilidad de falsos positivos pasando a ser .029. No obstante, se desconoce el criterio de decisión para alcanzar este resultado, lo que significa que este procedimiento no es una prueba válida cuya decisión se basa en una impresión global o juicio clínico (Arce & Fariña, 2013; Köhnken, 2004). En memorias de adultos: 8. Adicionalmente, la Hipótesis Undeutsch fue validada en adultos (puesto que la Hipótesis Undeutsch se sustenta en los contenidos de memoria, ya se había anticipado teóricamente que sería igualmente aplicable tanto en memorias de adultos como en otros contextos ajenos al abuso sexual) (Berliner & Conte, 1993). Asimismo, se validó la Hipótesis Undeutsch no solo para las memorias de víctimas, testigos y agresores, sino también para estudios experimentales (alta validez interna) y de campo (alta validez externa). Sin embargo, los criterios de realidad del CBCA clasificaron significativamente mejor las memorias vividas en estudios de campo que en estudios con diseños de simulación, no siendo éstos generalizables a la práctica (prueba suficiente) por sí mismos. 9. No todos los criterios de realidad fueron válidos (i.e., ‘auto-desaprobación’, ‘perdón al autor del delito’), ni generalizables (i.e., todos los criterios con excepción de ‘detalles incomprendidos relatados con precisión’ y ‘detalles característicos de la agresión’ que estaban afectados por un error de muestreo de segundo orden por lo que sus resultados no fueron válidos para esta estimación). No obstante, el criterio ‘perdón al autor del delito’ fue contrario a la Hipótesis Undeutsch (el tamaño del efecto medio fue negativo). 10. Aunque la Hipótesis Undeutsch y los criterios de realidad son también válidos para memorias de agresores, los resultados no pueden ser indubitablemente extrapolados a la práctica puesto que los datos proceden exclusivamente de estudios experimentales. Además de esto, en la evaluación forense de las memorias reales de agresores no existe la ground truth ya que los agresores pueden, por ley, mentir, no declarar, o simplemente negar los hechos alternativamente a contar la verdad. Bajo estas consideraciones, los criterios Discusión 57 ‘estructura lógica’, ‘producción inestructurada’, ‘cantidad de detalles’, ‘engranaje contextual’, ‘descripción de interacciones’, ‘reproducción de conversaciones’, ‘complicaciones inesperadas’, ‘detalles superfluos’, ‘estado mental del autor del delito’, ‘admisión de falta de memoria’, ‘detalles característicos de la agresión’ discriminaron significativamente entre memorias reales y fabricadas en agresores, siendo los dos últimos criterios conjuntamente con la puntuación total del CBCA, generalizables a cualquier condición. 11. Los resultados apoyan los criterios de realidad adicionales a los criterios de realidad que componen CBCA (i.e., ‘estilo de la declaración’, ‘dar muestras de inseguridad’, y ‘dar explicaciones de la falta de memoria’). En conclusión, los criterios de realidad del CBCA pueden complementarse con criterios de realidad adicionales. 12. Algunos de los criterios adiciones que fueron estudiados no son criterios de realidad, puesto que estaban inversamente relacionados con memorias vividas. Concretamente, los criterios ‘repeticiones’ y ‘clichés’ mostraron una relación negativa con las memorias vividas. Por tanto, constituyen atributos de memoria, esto es, características de memoria no directamente relacionadas con la realidad del evento, que deben ser estudiados con el objetivo de crear un sistema categorial que combine criterios de realidad y atributos de memoria. De facto, Arce y Fariña (2009; Vilariño, 2010; Vilariño, Novo, & Seijo, 2011) crearon y validaron un sistema categorial para discriminar memorias vividas y fabricadas en víctimas de violencia de género. 13. Los estudios de campo mostraron que no todos los criterios de realidad son válidos para discriminar entre memorias reales y fabricadas. Estos resultados sugieren que el sistema categorial debe reducirse a los criterios de realidad que discriminan significativamente, puesto que los no significativos solamente introducen ruido en la evaluación. 14. En consecuencia, la variable moderadora contexto de evaluación reveló que los criterios de realidad discriminan mejor en los casos reales de agresiones sexual y de violencia de género, clasificando correctamente al 71.5% de los verdaderos positivos y fallando en el 28.5% de los falsos positivos. 15. La eficacia de la puntuación total del CBCA en la discriminación entre memorias vividas y fabricadas fue significativamente superior en la muestra de menores que en la de adultos. 16. Comparativamente, los criterios ‘engranaje contextual’, ‘detalles superfluos’, ‘estado mental subjetivo’ y ‘detalles característicos de la agresión’ mejoran las ratios de clasificación en las muestras de menores. Los restantes criterios tienen la misma eficacia en la clasificación en ambas muestras. Sintomatología internalizante en casos de abuso sexual en la infancia y adolescencia 17. La victimización de abuso sexual infantil o adolescente (CSA/ASA) conlleva una probabilidad aproximada de .70 de padecer daño, .66 de depresión y .68 de ansiedad. La cuantificación estimada del daño causado fue de un 30% en la secuela en general, depresión y ansiedad. BÁRBARA GONZÁLEZ AMADO 58 18. En relación con el diagnóstico de trastornos internalizantes, las mujeres víctimas tienen un 9% más de probabilidades que los hombres de desarrollar depresión, pero no ansiedad. Esto deriva en un elevado daño en depresión (41 vs. 10% en mujeres y hombres, respectivamente) y ansiedad (24 vs. 14%) en mujeres en comparación con los hombres. Este hallazgo es consistente con los resultados de estudios epidemiológicos y de prevalencia informada por organismos internacionales de referencia en la clasificación de enfermedades (APA, 2013; WHO, 2000), señalando la vulnerabilidad de las mujeres en el desarrollo de secuelas internalizantes. 19. La probabilidad de desarrollar un daño crónico (distimia) fue significativamente superior a la probabilidad de un daño más grave o severo (trastorno de depresión mayor). 20. La cuantificación del daño para un trastorno depresivo persistente (distimia) y un trastorno de depresión mayor fue de un 46 y 31%, respectivamente, en las víctimas de ASI/ASA. 21. La victimización de ASI/ASA incrementa en un 43% la probabilidad de desarrollar un trastorno de ansiedad (trastorno de ansiedad generalizada, fobia específica, fobia social y trastorno de pánico). 22. En resumen, la victimización de ASI/ASA predispone a la aparición de problemas de ajuste psicológico en la adultez (Chiesa, Larsen-Paya, Martino, & Trinchieri, 2016). 23. El diagnóstico clínico es significativamente más adecuado (sensible) que la sintomatología clínica para la medida de victimización de ASI/ASA. En consecuencia, el diagnóstico clínico debería ser tomado como medida del daño consecuencia de la victimización de ASI/ASA, en lugar de la sintomatología. 24. La victimización de un delito sexual grave (i.e., con penetración), en contraste con la agresión sin contacto, conlleva un mayor daño en términos de depresión y ansiedad. Esto apoya la adjudicación de penas más altas en los Códigos Criminales cuando el delito de carácter sexual implica penetración, puesto que causa mayor daño psicológico que aquellos delitos contra la libertad e indemnidad sexuales donde la agresión no media contacto. Implicaciones de los resultados para la práctica forense La admisibilidad de pruebas científicas en el contexto judicial descansa en el cumplimiento de una serie de criterios establecidos por los tribunales y la jurisprudencia. A este respecto, se presentan a continuación las respuestas a dichos estándares sobre la admisibilidad de los criterios de realidad del CBCA a partir de los resultados alcanzados en los meta-análisis: 1. ¿Puede probarse la hipótesis científica? La hipótesis que subyace al sistema categorial CBCA, la denominada Hipótesis Undeutsch, ha sido sometida a prueba en numerosos estudios (todos los estudios incluidos en el primer, segundo y Discusión 59 cuarto meta-análisis de esta tesis doctoral, sometieron a prueba la hipótesis). Así, la hipótesis científica subyacente puede y ha sido sometida a contraste. 2. ¿Ha sido probada la hipótesis? La hipótesis evaluada a través del CBCA, ha sido probada y firmemente confirmada (con base en los resultados obtenidos en nuestros meta-análisis). 3. ¿Se conoce la tasa de error? Originalmente, no se proporcionaba ni se consideraba una tasa de error (Köhnken, 2004; Steller, 1989; Steller & Köhnken, 1989; Undeutsch, 1989) ya que la técnica era semi-objetiva o basada en el juicio clínico de los expertos. Vrij (2005, 2008) estimó (estudio cualitativo) una tasa de error próxima al 30% para las declaraciones verdaderas y falsas. Sin embargo, los resultados para las memorias falsas non son válidos puesto que la hipótesis y los criterios no clasifican dichas memorias. Nuestros resultados mostraron diferentes tasas de error en la clasificación de memorias reales (falsos positivos), siendo los más significativos, de un 9.8% en casos reales de abuso sexual infantil y de 28.5% en casos reales de violencia de género y delitos de naturaleza sexual en muestras de adultos. 4. ¿Ha sido sometida la hipótesis/técnica a un proceso de revisión por pares y publicación? Sí, muchos estudios que contrastan la hipótesis y la técnica han sido publicados. Muchos de ellos sometidos a un proceso de revisión por pares. 5. ¿Se basa la hipótesis y/o técnica en una teoría generalmente aceptada como válida por la comunidad científica? Aunque la comunidad científica no ha sido consultada como tal, la evidencia científica que procede de investigadores adecuados proporciona soporte (significativo y generalizable en los estudios de campo; ver los resultados de los meta-análisis) a la hipótesis y a la técnica. Puesto que no se trata de un soporte mínimo, en cuyo caso los resultados deben ser considerados con escepticismo (Daubert v. Merrell Dow Pharmaceuticals, Inc., 1993), la conclusión razonable es que la hipótesis y la técnica cumple este criterio. Conjuntamente a los criterios Frye, Kelley y Daubert y a los requerimientos legales, procedimentales y científicos para la práctica forense, Arce (2016) creó una lista de requerimientos para su cumplimiento por las pruebas de credibilidad de testimonio. Bajo estas premisas, la Hipótesis Undeutsch y la técnica (SVA y CBCA) fueron sometidas a prueba: 1. ¿Es la teoría científica que subyace a la evidencia, válida, comprobable y validada en una publicación científica? La Hipótesis Undeutsch no se derivó originalmente de una teoría científica, sino de información proveniente de archivos de caso, esto es, de un procedimiento bottom-up (fue originaria de un elevado número de casos reales de abuso sexual infantil). No obstante, la teoría que subyace a la Hipótesis Undeutsch puede explicarse desde diferentes teorías de la memoria y el engaño como las teorías de la racionalidad y constructivistas las cuales establecen que quien miente tiene una visión estereotipada de la mentira, es decir, la mentira es planeada y aprendida y, por tanto, consistente en el tiempo. Así, las personas que mienten construyen las mentiras a partir de scripts cognitivos (Schank & Abelson, 1977). Mientras, teorías sobre la memoria episódica y autobiográfica apoyan la diferente calidad de las memorias basadas en hechos experimentados. El modelo del reality monitoring también BÁRBARA GONZÁLEZ AMADO 66 7. The analysis of the research paradigm moderator showed a greater effect size and true positive classification accuracy of the total CBCA score for field studies (90.2%) than for experimental studies (63.5%). In consequence, results from experimental studies may not be generalized directly for practice, revealing face validity for experimental data (Fariña et al., 1994; Konecni & Ebbesen, 1992). This also corrects the false positive rate passing to .029. Nevertheless, the strict criterion for this remains unknown, which means that this procedure is not a valid/power proof since it is based on a global impression or a clinical judgment (Arce & Fariña, 2013; Köhnken, 2004). In adult memories: 8. Additionally, the Undeutsch Hypothesis was tested for adult population (since Undeutsch Hypothesis was grounded on memory content, it was already theoretically anticipated that the hypothesis would be also appropriate both in adults and in contexts other than child sexual abuse) (Berliner & Conte, 1993). Likewise, Undeutsch Hypothesis was validated not only for victims/claimants, witnesses and offenders’ memories but for experimental (high internal validity) and field (high external validity) studies. However, CBCA reality criteria classified vivid memories significantly superior in field rather than in simulation study designs, not being directly generalized by their own to practice. 9. Not all reality criteria were valid (i.e, ‘self-deprecation’, ‘pardoning the perpetrator’) nor generalizable (i.e., all criteria with the exception of ‘accurately reported details misunderstood’ and ‘details characteristic of the offence’ which are affected by a second-order sampling error and it is not possible to do an estimation). Notwithstanding, ‘pardoning the perpetrator’ criterion performed against Undeutsch Hypothesis (negative mean effect size). 10. Even though Undeutsch Hypothesis and reality criteria are also valid for offenders’ memories, results may not be undoubtedly extrapolated to practice as data has exclusively come from experimental studies. Even so, real offender memories in forensic evaluations have not a ground truth because offenders are able, by law, to lie, to not declare, or they simply may deny the alleged event alternatively to telling the truth. Under these considerations, the ‘logical structure’, ‘unstructured production’, ‘quantity of details’, ‘contextual embedding’, “descriptions of interactions’, ‘reproduction of conversations’, ‘unexpected complications’, ‘superfluous details’, ‘perpetrators mental state’, ‘admitting lack of memory’ and ‘details characteristic of the offence’ criteria distinguished significantly between the offender’s vivid and fabricated memories, being the last two criteria, as well as the total CBCA score, generalized to any condition. 11. Results support the additional reality criteria to those involving CBCA reality criteria (i.e., ‘reporting style’, ‘display insecurities’, ‘providing reasons for lack of memory’). In consequence, CBCA criteria may be supplemented with additional reality criteria. Discussion 67 12. Some of the studied additional reality criteria are not in fact reality criteria, as they are inversely related to vivid memories. Thus, ‘repetitions’ and ‘clichés’ criteria have a negative relation with vivid memories. Therefore, these are memory attributes i.e., memory characteristics not directly related to reality, that should be studied to create a categorical system combining reality criteria and memory attributes. On this, Arce and Fariña (2009; Vilariño et al., 2011) created and validated a categorical system to discriminate vivid and fabricated memories of intimate partner violence claimants. 13. Overall, field studies exhibited that not all reality criteria are valid to discriminate between true and fabricated memories. These results suggest reducing the categorical system to significant discriminant criteria, as the non-significant criteria only introduces noise. 14. Consequently, the moderator variable ‘evaluation setting’ revealed the best performance of the reality criteria for statements of sexual offences and intimate partner violence field studies, classifying 71.5% of true positives correctly and failing in 28.5% of false positives. 15. The efficiency of the total CBCA score to distinguish between fabricated and vivid memories was significantly higher for children than for adults. 16. Comparatively, ‘contextual embedding’, ‘superfluous details’, ‘subjective mental state’ and ‘details characteristic of the offence’ criteria enhanced classification rates in children. Remaining criteria performed equally in adults and children. Child and adolescent sexual abuse internalizing symptomatology 17. Child or adolescent sexual abuse victimization (CSA/ASA) implies approximately a probability of .70 of suffering internalizing injury, of .66 of depression and of .68 of anxiety. Damage quantification was estimated at about 30% in general sequelae, depression, and anxiety. 18. In relation to the diagnosis of internalizing disorders, female victims have 9% more probability of developing depression than male victims, but not anxiety. This derives in large damage in depression (41 vs. 10%, for females and males, respectively) and anxiety (24 vs. 14%) for females than for males. This result is linked to epidemiological studies findings and international organizations classifying mental disorders prevalence informed (APA, 2013; WHO, 2000) pointing out the vulnerability of women in developing internalizing sequelae. 19. The probability of chronic injury (dysthymia) was significantly greater than more severe injury (major depressive disorder). 20. Damage quantification for a persistent depressive disorder (dysthymia) and for a major depressive disorder was about 46 and 31% for CSA/ASA victims, respectively. BÁRBARA GONZÁLEZ AMADO 68 21. CSA/ASA victimization increases 43% the probability of developing an anxiety disorder (generalized anxiety disorder, specific phobia, social phobia, or panic disorder). 22. In short, suffering CSA/ASA predisposes victims to psychological adjustment problems in adulthood (Chiesa, Larsen-Paya, Martino, & Trinchieri, 2016). 23. Clinical diagnosis is a significantly more adequate (sensitive) injury measure than symptomatology for CSA/ASA victims. In consequence, clinical diagnosis should be preferred as a benchmark, rather than symptoms to assess the CSA/ASA victimization sequelae. 24. Severe sexual aggression victimization (i.e., with penetration), in contrast to noncontact aggression, causes a higher anxious and depressive injury. This finding gives support to Criminal Codes adjudicating more severe punishment when the sexual crime implies a penetration, as it causes a greater psychological damage than sexual crimes with non-contact aggression. Implications of the results for forensic practice Admissibility of scientific evidence in judicial setting rests on meeting standards established by courts and the law of precedent. On this, the answers to such standards about the admissibility of the CBCA reality criteria throughout meta-analyses results are presented below: 1. Is the scientific hypothesis testable? The underlying hypothesis of the CBCA categorical system, the Undeutsch Hypothesis, has been submitted to proof in many studies (all the studies included in the first, second and fourth metaanalyses involved in the present doctoral dissertation submitted to test the hypothesis). Thus, the underlying scientific hypothesis may and has been submitted to contrast. 2. Has the hypothesis been tested? The hypothesis assessed through the CBCA has been tested and firmly confirmed (on the basis of our meta-analyses). 3. Is there a known error rate? Originally, the error rate was neither provided, nor considered (Köhnken, 2004; Steller, 1989; Steller & Köhnken, 1989; Undeutsch, 1967, 1988) as the resulting technique was semi-objective or based on expert clinical judgments. Vrij (2005, 2008) estimated (qualitative study) the error rate around 30% for true and false statements. Nevertheless, the results for false memories are not valid as the hypothesis and criteria do not classify them. Our results stated different error rates in truthful memories classification (false positives), being the most noteworthy, 9.8% in real child sexual abuse cases and 28.5% in real sexual and intimate partner violence adult cases. 4. Has the hypothesis/technique been subjected to peer review publication? Yes, many studies contrasting the hypothesis and technique have been published. Most of them submitted for peer review. Discussion 69 5. Is the theory upon which the hypothesis and/or technique is based on generally accepted among appropriate scientific community? Although the scientific community has not been consulted as such, scientific evidence coming from the appropriate researchers give support (significant and generalized in field context; see the results of our meta-analyses) to the hypothesis and technique. As this is not a minimal support, in which case results should be considered with skepticism (Daubert v. Merrell Dow Pharmaceuticals, Inc., 1993), the reasonable conclusion is that the hypothesis and technique meet this standard. Joining the Frye, Kelley and Daubert Standards with scientific, procedural and legal requirements to forensic proofs, Arce (2016) created a list of testimony credibility proof requirements for which the Undeutsch Hypothesis and technique (SVA and CBCA) are submitted: 1. Is the scientific theory, upon which the evidence-based is valid, testable and tested in a scientific publication? The Undeutsch Hypothesis was not originally derived from a scientific theory, but from data, that is, bottom-up (it was originated from an elevated number of real-life child sexual abuse criminal cases). However, the underpinning Undeutsch Hypothesis theory can be explained by different memory and deception theories such as rationality and constructivism which state that liars have a stereotypical notion of lies, that is, a lie is planned and learned and consequently consistent in time. Thus, liars construct their lies from cognitive scripts (Schank & Abelson, 1977). Meanwhile, theories about episodic and autobiographical memories support different quality of memories from self-experienced events. The reality monitoring model also gives theoretical support (memory content categories discriminate between internal ―imagined/thought― and external ―perceived― memories) to the hypothesis (memory attributes discriminate between both memories) (Johnson & Raye, 1981). Finally, sensorial and narrative memories may explain the differences between self-experienced and fabricated memories, being the content criteria the media for this. The results of our meta-analyses validate the hypothesis, that is, the theory i.e., Undeutsch Hypothesis, is testable and was tested and verified by the scientific publications, that is, the scientific community research agrees on its validity (see Daubert standards 1, 2, 4 and 5 for more details). 2. Is the technique derived from the hypothesis valid? The technique derived from the hypothesis, the reality criteria of the CBCA, discriminate significantly between-contexts (e.g., children, adults, different crimes, victims, witnesses) between vivid and fabricated memories. In consequence, scientific evidence validates the technique. Nonetheless, for forensic practice it is not valid because a strict decision criterion to classify statements is not provided, leaving the forensic decision resting on the expert impression or clinical judgment. Thus, the resulting decision is semi-objective i.e., based on general scientific evidence (objective) but when applied to a specific case sustained in the expert impression (subjective), while an objective decision criterion is expected in court. 3. Has been the technique applied correctly in a particular case? The SVA, as the whole procedural system where the application of the CBCA reality criteria is BÁRBARA GONZÁLEZ AMADO 70 described, does not meet this standard as it does not have any measure. In consequence, forensic practice results (N = 1) are not reliable. 4. Is it possible for the methods to be testable and revised by other experts? Although the SVA does not include this standard, if the statements are recorded and maintained, the methods may be tested. 5. Is it possible for the evidence to be replicated (second expert testimony)? If the witness statement could be obtained once again, the evidence may be replicated. 6. Is it possible for the persistence in the testimony (test-retest) to be assessed? As the statements obtained with non-invasive interview techniques i.e., centered in a free narrative recall, may be applied in repeated measures (Memon, Meissner, & Fraser, 2010; Memon, Wark, Bull, & Köhnken, 1997), testimony persistence may be assessed. As for this, Undeutsch created a second hypothesis: the central elements of the account are persistent; meanwhile the peripheral ones may not be (expected results for vivid memories). Memory recall theories and evidence support this (Fariña, Arce, & Real, 1994). 7. Meanwhile, the error rate is known (Is the error rate known? Daubert standard 3, see above), it does not meet the law of precedent standard of the presumption of innocence (no one innocent person shall be imprisoned, whereas it is enough that a guilty person generally stays imprisoned; 832/2000 28th February, 2000 Spanish Supreme Court Sentence; 213/2002 14th February, 2002 Spanish Supreme Court Sentence). That is, given that the classification of a fabricated memory as a real memory involves a false accusation (false positive), in forensic evaluation its rate must be zero. The technique fails in the classification of false positives among 9.8% in children and 28.5% in adults. In consequence, the technique is not valid for forensic purposes. 8. Are the conclusions about the testimony valid in a judicial setting? SVA establishes three decision options: a) in a qualitative judgment guided by rules; b) in a five-point Likert scale, ranging from incredible (-2) to credible (+2), passing through indeterminate (0) (Steller, 1989); and c) in an expert clinical judgement i.e., qualitative (Köhnken 2004). These three options are not valid for courts as they lack objectivity, replicability, and consistency, and are even confusing for courts (e.g., Spanish Supreme Court advertises about this). 9. Is the content analysis system valid? Psychometric properties of the CBCA are unknown. In order for a categorical system to be a methodic measure (Bardin, 1977), category analysis must be mutually exclusive (overlapping between categories has been found, producing duplicity of measures; Horowitz et al., 1997; Roma et al., 2011); homogeneous (unknown); exhaustive (CBCA reality criteria system is not exhaustive as additional criteria may be added, see adult meta-analysis; Arce & Fariña, 2009); fidelity (SVA has not an estimation of the fidelity of its application in forensic setting; in scientific research, fidelity is variable); objective (categories are not defined with precision giving to different interpretations and codings); pertinence (as the system was formulated for child sexual abuse memories, some categories are not pertinent in other evaluation contexts ―see adult meta-analysis―, and deception strategies depend on the Discussion 71 context; Volbert & Steller, 2014); and productive (all the CBCA reality criteria were productive and significant to discriminate memories). 10. Expert witness qualification. Although SVA does not realize the expert qualification, authors advertise that training in this technique is necessary and effective, in both content analysis and interview techniques, eliciting reality criteria (Fisher, Geiselman, & Raymond, 1987; Köhkhen, 2004; Vrij, 2005). In consequence, training is required and, when available, an official qualification as a forensic psychologist (professional accreditation may substitute official if it is not available). Expert psychological evidence must contemplate sequelae quantification measures in addition to the testimony credibility assessment, consequence of child sexual abuse or adolescent sexual abuse victimization. It has been estimated that there is an elevated probability of developing depression and anxiety disorders in adulthood as a consequence of victimization (long-term sequelae). Consequently, quantifying psychological injury from sexual crime victimization emerges as a way to compensate victim economically (as well as crime responsibility) and this civil responsibility must be proportional to the injury caused. Even though we must study each isolated case, meta-analytic results with moderator variables give us a guide to the extent of the injury expected according to each case (i.e., sex of the victim, severity of the abuse). As clinical diagnosis emerges as a significantly more sensitive injury measure than the symptoms, and psychometric instruments, as a recognition task, they favor symptomatology malingering (in the judicial setting we must suspect malingering, DSM). Measuring damage from clinical diagnosis through an interview is recommended because it allows both linking injury and victimization and to control malingering. Forensicclinical interview (Arce & Fariña, 2001) has arisen as an efficient interview in forensic practice for this issue because it elicits spontaneous symptomatology reports related to the crime (knowledge task). In short, the main objective of this doctoral dissertation was to provide answers to the concerns about the admissibility in court of the credibility assessment categorical system. Meta-analytic technique with its integrative labor of empirical results permitted reaching generalizable conclusions that contribute to the Evidence-Based Psychology (BEP) approach. Findings found the creation and validation of the methodic credibility assessment categorical system, which should be made up of reliable and valid categories of content analysis, conformed by significant reality criteria and contextually specific (i.e., sexual violence, intimate partner violence). In fact, CBCA does not fulfil some of these legal (e.g., elevated classification error rate, lack of a strict decision criterion) and methodological requirements (e.g., some reality criteria are non-mutually exclusive, and the system is neither objective nor pertinent) in its present state. Likewise, the assessment method must allow studying psychological injury victimization and its measure as a consequence of a sexual crime, with the objective of repairing damage proportionally to the distress suffered even retrospectively (injury may appear after a long period of time and even become chronic). Meta-analysis results provide juries and judges a prospectively estimated measure of injury, which shed some light the extent of the cost of psychological damage in sexual crimes victimization. In addition, after diagnosing an internalizing disorder, Global Assessment of Functioning (GAF) scale must be applied to the particular case in order to quantify the individual’s overall level of distress (psychological adjustment in different life facets). This measure allows repair BÁRBARA GONZÁLEZ AMADO 72 damage proportionally to the distress suffered, and it would be helpful for judges to reach a more accurate and fair judicial sentence. Future research must focus their efforts on the following investigation lines. First, improving system effectiveness through identifying new productive categories (real cases bottom-up process), regarding specific contexts. Likewise, combining CBCA reality criteria with SRA criteria, memory attributes (RM) and additional criteria coming from real-cases, the categorical system is expected to improve its discriminative capacity. Even more, while the inclusion of “deception criteria” in CBCA had not been accepted because of its original nature, according to Steller (1988), their integration might be a path to a better performance of the categorical system. Second, motivational criteria showed the lowest discriminative capacity but they were nonetheless productive. They are related to a mindful deception strategy, deceivers avoid consciously motivational criteria in his or her statements in order to appear honest. For this reason, it remains relevant to study deeply those criteria, because they seem to be more influenced by context or situation assessment than criteria, which are part of the cognitive component of the hypothesis. Concretely, ‘self-deprecation’ and ‘pardoning the perpetrator’ must be tested in real forensic cases because findings have pointed out that they seem not to be memory-related content criteria but self-presentation strategies, consistent with other proposals (Niehaus, 2008). Third, research in criminological characteristics of different type of sexual offenses will help not only to create a specific testimony assessment protocol but also to define ‘details characteristics of the offence’ criterion that could be highly relevant in a truthful decision-making procedure. Fourth, CBCA differentiates between memories of video-observed events (non-self-experienced) and fabricated events, even though the hypothesis has not been formulated for that aim. Accordingly, with these originally nonhypothesized findings, it would be of great interest to analyze the CBCA capacity to discriminate between statements based on self-experienced events and video-observed events. Finally, more field studies with real victims/witnesses should be performed in order to obtain a strict decision criterion to make an accurate decision about the reality of a given testimony. Experimental studies designed under high fidelity conditions may engage internal validity and contribute significantly to research in credibility. Therefore, future laboratory research might include the following criteria: personal involvement of the participant in the scenario, negative emotional tone of an event, loss of control over the situation (Steller, 1989), and giving a testimony in a face-to-face situation under which the implication or motivation of the participants is controlled, eliciting long statements not only by a free recall but asking questions, and being questioned repeatedly (Volbert & Steller, 2014). The meta-analytic studies that have tested CBCA validity in both children and adults’ samples share certain limitations that must be taken into account. First, interviews’ reliability prior to the content analysis is unknown. A poor interviewers training may lead to obtain insufficient statements for conducting a credibility analysis or to get biased testimonies as a result of suggestive questions stemmed from the interviewer’s style (e.g., leading and/or suggestive questioning style). Moreover, most of the primary investigations did not inform about the SVA standards compliance. Concretely, there was a lack of information regarding the type of interview used, which is a relevant aspect for the statement’s quality evaluation. These considerations may favor Undeutsch Hypothesis rejection as a result of bad practices. Second, experimental studies had no proven external validity (Sarwar, Allwood, & Innes-Ker, 2014), but only ‘face validity’ (Konecni & Ebbesen, 1992). Researchers in laboratory context manipulate experimental conditions in order to know the veracity of statements, that is, the Discussion 73 ground truth. This leads them to the use of the CBCA as a deception detection tool, classifying false statements with confidence. Nevertheless, in the forensic practice, other alternative hypotheses to the total and conscious fabrication of the testimony must be considered as a possible explanation for the poor quality of the statement: lack of victims’ collaboration, insufficient memories regarding the incident to analyze the content of the statement, or diminished cognitive capacities (Köhnken, 2004). Third, the trend towards to not publish non-significant results and to remove non-efficient reality categories to discriminate between real and fabricated memories (conflicting data), favors Undeutsch Hypothesis fulfillment. Fourth, reality criteria reliability was estimated since most of the primary studies did not report the raters’ consistency. Additionally, the estimated reliability based on the presence of the reality criteria frequency without identifying the exact correspondence on statement overestimates their reliability. Furthermore, content analysis reliability is related not only to coders’ consistency but also to within-coder consistency (reliability in time) and between-contexts consistency (with other raters independent to the study, and other materials). Conducting a more reliable codification would offset the potential effect of a truth bias or a response bias associated with the application of reality criteria (Griesel, Ternes, Schraml, Cooper, & Yuille, 2013; Rassin, 1999; Sporer, 2004). Finally, some meta-analytic results may be subject to certain variability when Ns < 400 so they do not guarantee the stability of sampling estimates (Hunter & Schmidt, 2015). Several limitations were also found in the meta-analytic studies of the psychological consequences of sexual abuse victimization. First, since primary investigations were real cases of child and adolescent sexual abuse, the ground truth rests on self-reports of a retrospective nature in most cases. This aspect hinders veracity validation of facts, underestimates the relationship between victimization and injury and facilitates classification errors, especially false negatives. Second, the symptomatology and the diagnosis of a depressive and/or anxiety disorder as a consequence of a victimization of a sexual crime was assumed, however, cause-effect relation between facts and psychological damage has not been explicitly established. Third, some primary investigations informed about an abuse measure without differentiating between typologies (e.g., sexual abuse, physical abuse, negligence). Fourth, since some studies did not report a control group, normative population was taken as the contrast group. 75 7. REFERENCIAS American Psychiatric Association. (APA, 2013). Diagnostic and statistical manual of mental disorders (5th ed.). Washington, DC: American Psychiatric Association. Arce, R. (2016). Evaluación del SVA/CBCA y el SEG: Criterios Frye, Kelly, Daubert, jurisprudenciales, procesales y legales. Manuscrito Inédito. Santiago de Compostela, España: Unidad de Psicología Forense. Arce, R., & Fariña, F. (2001). Construcción y validación de un procedimiento basado en una tarea de reconocimiento para la medida de la huella psíquica en víctimas de delitos: La entrevista forense. Manuscrito inédito. Universidad de Santiago de Compostela. Arce, R., & Fariña, F. (2005). Psychological evidence in court on statement credibility, psychological injury and malingering: The Global Evaluation System (GES). Papeles del Psicólogo, 26, 59-77. Retrieved from http://www.papelesdelpsicologo.es/English/1247.pdf Arce, R., & Fariña, F. (2006). Psicología del testimonio: Evaluación de la credibilidad y de la huella psíquica en el contexto penal. In Consejo General del Poder Judicial (Ed.), Psicología del testimonio y prueba pericial (pp. 39-103). Madrid: Consejo General del Poder Judicial. Arce, R., & Fariña, F. (2009). Evaluación psicológica forense de la credibilidad y daño psíquico en casos de violencia de género mediante el Sistema de Evaluación Global. In F. Fariña, R. Arce, & G. Buela-Casal (Eds.), Violencia de género. Tratado psicológico y legal (pp. 147-168). Madrid: Biblioteca Nueva. Arce, R., & Fariña, F. (2013). Psicología forense experimental. Testigos y testimonio. Evaluación cognitiva de la veracidad de testimonios y declaraciones. In S. Delgado (Dir. Tratado), S. Delgado, & J. M. Maza (Coords. Vol.), Tratado de medicina legal y ciencias forenses: Vol. V. Psiquiatría legal y forense (pp. 21-46). Barcelona, España: Bosch. Arce, R., & Fariña, F. (2015). Evaluación psicológico-forense de la credibilidad y daño psíquico mediante el Sistema de Evaluación Global. In P. Rivas & G. L. Barrios (Dirs.), Violencia de género: Perspectiva multidisciplinar y práctica forense (pp. 657-367). Navarra, España: Thomson Aranzadi. Arce, R., Fariña, F., & Vilariño, M. (2010). Contraste de la efectividad del CBCA en la evaluación de la credibilidad en casos de violencia de género [Contrasting the efficiency of the CBCA in the assessment of credibility in violence against women cases]. Intervención Psicosocial, 19, 109-119. doi: 10.5093/in2010v19n2a2 BÁRBARA GONZÁLEZ AMADO 82 Volbert, R., & Steller, M. (2014). Is this testimony truthful, fabricated, or based on false memory? Credibility assessment 25 years after Steller and Köhnken (1989). European Psychologist, 19, 207 -220. doi: 10.1027/1016-9040/a000200 Vrij, A. (2005). Criteria-Based Content Analysis: A qualitative review of the first 37 studies. Psychology, Public Policy, and Law, 11, 3–41. doi: 10.1037/1076-8971.11.1.3 Vrij, A. (2008). Detecting lies and deceit: Pitfalls and opportunities (2nd ed.). Chichester, UK: John Wiley and Sons. Vrij, A., Akehurst, L., Soukara, R., & Bull, R. (2004). Detecting deceit via analyses of verbal and nonverbal behavior in children and adults. Human Communication Research, 30, 841. doi:10.1111/j.1468-2958.2004.tb00723.x Vrij, A., Edward, K., Roberts, K.P., & Bull, R. (2000). Detecting deceit via analysis of verbal and nonverbal behavior. Journal of Nonverbal Behavior, 24, 239-263. doi:10.1023/A:1006610329284 World Health Organization. (WHO, 1999, March 29-31). Report of the consultation on child abuse prevention. Geneva, Switzerland: Author. Retrieved from http://www.who.int/iris/handle/10665/65900#sthash.YgBC2YVe.dpuf World Health Organization. (2000). Women’s mental health: An evidence based review. Geneva, Switzerland: Author. Yuille, J. C., Marxsen, D., & Cooper, B. S. (1999). Training investigative interviewers: Adherence to the spirit, as well as the letter. International Journal of Law and Psychiatry, 22 (3-4), 323-336. doi: 10.1016/S0160-2527(99)00012-6 8. ARTÍCULOS DE INVESTIGACIÓN 85 8.1. UNDEUTSCH HYPOTHESIS AND CRITERIA BASED CONTENT ANALYSIS: A METAANALYTIC REVIEW Bárbara G. Amado*, Ramón Arce*, & Francisca Fariña** (2015) *Department of Psicología Organizacional, Jurídico-Forense y Metodología de las Ciencias de Comportamiento, University of Santiago de Compostela (Spain). **AIPSE Department, University of Vigo (Spain). Abstract The credibility of a testimony is a crucial component of judicial decision-making. Checklists of testimony credibility criteria are extensively used by forensic psychologists to assess the credibility of a testimony, and in many countries they are admitted as valid scientific evidence in a court of law. These checklists are based on the Undeutsch hypothesis asserting that statements derived from the memory of real-life experiences differ significantly in content and quality from fabricated or fictitious accounts. Notwithstanding, there is considerable controversy regarding the degree to which these checklists comply with the legal standards for scientific evidence to be admitted in a court of law (e.g., Daubert standards). In several countries, these checklists are not admitted as valid evidence in court, particularly in view of the inconsistent results reported in the scientific literature. Bearing in mind these issues, a meta-analysis was designed to test the Undeutsch hypothesis using the CBCA Checklist of criteria to discern between memories of self-experienced real-life events and fabricated or fictitious accounts. As the original hypothesis was formulated for populations of children, only quantitative studies with samples of children were considered for this study. In line with the Undeutsch hypothesis, the results showed a significant positive effect size that is generalizable to the total CBCA score, δ = 0.79. Moreover, a significant positive effect size was observed in each and all of the credibility criteria. In conclusion, the results corroborated the validity of the Undeutsch hypothesis and the CBCA criteria for discriminating between the memory of real self-experienced events and false or invented accounts. The results are discussed in terms of the implications for forensic practice. Key words: meta-analysis, CBCA, credibility, testimony, sexual abuse, child. Hipótesis Undeutsch y Criteria Based Content Analysis: Una revision meta-analítica. Resumen Con frecuencia, la evaluación de la fiabilidad de un testimonio se lleva a cabo mediante el uso de sistemas categoriales de análisis de contenido. Concretamente, el instrumento más utilizado para determinar la credibilidad del testimonio es el Criteria Based Content Analysis BÁRBARA GONZÁLEZ AMADO 86 (CBCA), el cual se sustenta en la hipótesis Undeutsch, que establece que las memorias de un hecho auto-experimentado difieren en contenido y calidad de las memorias fabricadas o imaginadas. Las opiniones y resultados contradictorios encontrados en la literatura científica respecto al cumplimiento de los criterios judiciales (Daubert standards) así como el abundante número de trabajos existentes sobre la materia, nos llevó a diseñar un meta-análisis para someter a prueba la hipótesis Undeutsch, a través de la validez de los criterios de realidad del CBCA para discriminar entre la memoria de lo auto-experimentado y lo fabricado. Se tomaron aquellos estudios cuantitativos que incluían muestras de menores, esto es, con edades comprendidas entre los 2 y 18 años. En línea con la hipótesis Undeutsch, los resultados mostraron un tamaño del efecto positivo, significativo y generalizable para la puntuación total del CBCA, δ = 0.79. Asimismo, en todos los criterios de realidad se encontró un tamaño del efecto positivo y significativo. En conclusión, los resultados avalan la validez de la hipótesis Undeutsch y de los criterios del CBCA para discriminar entre memorias de hechos autoexperimentados y fabricados. Se discuten las implicaciones de los resultados para la práctica forense. Palabras clave: meta-análisis, CBCA, credibilidad, testimonio, abusos sexuales, menores. Hipótese Undeutsch e Criteria Based Content Analysis: Unha revisión meta-analítica Resumo Con frecuencia, a avaliación da fiabilidade da testemuña lévase a cabo a partires do uso de sistemas categoriais de análise de contido. Concretamente, o instrumento máis usado para determinar a credibilidade do testemuño é o Criteria Based Content Analysis (CBCA), o cal susténtase na Hipótese Undeutsch, que establece que as memorias de feitos autoexperimentados difiren en contido e calidade das memorias fabricadas ou imaxinadas. As opinións e resultados contraditorios atopados na literatura científica respecto ó cumprimento dos criterios xudiciais (Daubert standards) así coma o abundante número de traballo existentes sobre a materia, levounos a deseñar un meta-análise para someter a proba a hipótese Undeutsch. A partir da validez dos criterios de realidade do CBCA para discriminar entre a memoria do auto-experimentado e o fabricado. Tomáronse aqueles estudios cuantitativos que incluían mostras de menores, isto é, con idades comprendidas entre os 2 e 18 anos. En liña coa Hipótese Undeutsch, os resultados amosaron un tamaño do efecto positivo, significativo e xeralizable para a puntuación total do CBCA, δ = 0.79. asemade, en todos os criterios de realidade atopouse un tamaño do efecto positivo e significativo. En conclusión, os resultados avalan a validez da hipótese Undeutsch e dos criterios do CBCA para discriminar entre memorias de feitos auto-experimentados e fabricados. Discútense as implicacións dos resultados para a práctica forense. Palabras chave: meta-análise; CBCA; credibilidade; testemuño; abusos sexuais; menores. Artículos de investigación 87 Acknowledgements This research has been carried out within the framework of research project with the Reference Ref: GPC2014/022, funded by the Xunta de Galicia [Galician Autonomous Government] (Spain). The authors acknowledge and thank to Professor Jesús F. Salgado for his methodological support and the revision of this manuscript. Introduction Hans and Vidmar (1986) estimated that in around 85% of judicial cases, the evidence bearing most weight is the testimony, which underscores that the evaluation of a testimony is crucial for judicial judgement-making. In terms of the application of Information Integration Models to legal judgements (Kaplan, 1982), the reliability and validity of the testimony are the mechanisms underlying the evaluation of a testimony. The validity of a testimony, i.e., the value of a testimony for judgement-making is easily estimated and is to be determined by the rulings of judges and the courts. As for the reliability of a testimony, the courts and scientific studies have tended to estimate it in terms of the credibility of a testimony (Arce, Fariña, & Fraga, 2000), which entails the design of methods for its estimation. Traditionally, judges and the courts have performed this function on the basis of legal criteria, jurisprudence, and their own value judgements. Alternatively, numerous scientific techniques (and pseudoscientific) have been proposed such as non-verbal indicators of deception; paraverbal indicators of deception; physiological indicators (e.g., polygraph tests or functional magnetic resonance imaging); and categorical systems of content analysis. Of these, categorical systems of content analysis are currently the most systemically used technique by the courts. Thus, the courts in countries such as Germany, Sweden, Holland, and several states in the USA admit these categorical systems as scientific evidence (Steller & Böhm, 2006; Vrij, 2008). In Spain, where they are also admitted as legally admissible evidence and extensively used by the courts, an analysis of legal judgements showed that when a forensic psychological report based on a categorical system of content analysis (i.e., Statement Validity Analysis, SVA) confirmed the credibility of a testimony, the conviction rate was 93.3%, but when it failed to do so, the acquittal rate was 100%. In contrast, in other countries such as the UK, the US, and Canada these checklists are not admitted as legally valid evidence (Novo & Seijo, 2010). Underlying categorical content systems is what is commonly referred to as the Undeutsch hypothesis that asserts that the memory of a real-life self-experienced event differs in content and quality from a fabricated or imagined event (Undeutsch, 1967, 1989). On the basis of this hypothesis, Steller and Köhnken (1989) have integrated all the categorical systems (e.g., Arntzen, 1970; Dettenborn, Froehlich, & Szewczyk, 1984; Szewczyk, 1973; Undeutsch, 1967) into what is known as Criteria Based Content Analysis (CBCA), which has become the leading categorical system for evaluating the credibility of a testimony (Griesel, Ternes, Schraml, Cooper, & Yuille, 2013; Vrij, 2008). CBCA, which is part of SVA, consists of three elements: 1) semi-structured interview, i.e., the free narrative interview; 2) content analysis on CBCA criteria; and 3) evaluation of CBCA outcomes using the Validity Checklist. The semi-structured interview involves a narrative format that, unlike other types of interview such as standard, interrogative or structured interviews, facilitates the emergence of criteria (Vrij, 2005). Moreover, this type of BÁRBARA GONZÁLEZ AMADO 88 interview generates more information (Memon, Meissner, & Fraser, 2010), which meets the requirement that CBCA criteria content analysis be performed on sufficient material (Köhnken, 2004; Steller, 1989). The Checklist of CBCA criteria (Steller & Köhnken, 1989) consists of 19 criteria structured around 5 major categories: general characteristics, specific contents, peculiarities of the content, contents related to motivation, and specific elements of aggression (see Table 1). These criteria of reality do not constitute a methodic categorical system (Bardin, 1977; Weick, 1985), but rather stem from the authors’ personal experiences of cases (Steller & Köhnken, 1989). Though this checklist was originally developed as a comprehensive system of credibility criteria grounded on the Undeutsch hypothesis, Raskin, Esplin, and Horowitz (1991) highlighted that only the first 14 criteria are related to the Undeutsch hypothesis, and the remaining 5 criteria are not associated to the aforementioned hypothesis as they are not linked to the concept of memory of actual events. This reclassification overlaps, though not entirely, with the theoretical model proposed by Köhnken (1996) who regroups these major categories into two main factors: cognitive (criteria 1 to 13), and motivational (criteria 14 to 18). The cognitive factor encompasses cognitive and verbal skills, and implies a selfexperienced statement contains CBCA criteria from 1 to 13. The motivational factor, however, relies on the individual’s ability to avoid appearing deceitful and ways of managing a positive self-impression of oneself as an honest witness. Thus, the motivational factor covers criteria 14 to 18, which are contrary-to-truthfulness-stereotype criteria though they really appear in true statements. Thus, these criteria have been suggested to be useful for assessing the hypothesis of the (partial) fabrication of statements (Köhnken, 1996, 2004). Table 1.CBCA-Criteria (adapted from Steller & Köhnken, 1989) GENERAL CHARACTERISTICS 1. Logical structure 2. Unstructured production 3. Quantity of details SPECIFIC CONTENTS 4. Contextual embedding 5. Descriptions of interactions 6. Reproduction of conversation 7. Unexpected complications during the incident PECULIARITIES OF CONTENT 8. Unusual details 9. Superfluous details 10. Accurately reported details misunderstood 11. Related external associations 12. Accounts of subjective mental states 13. Attribution of perpetrator’s mental state MOTIVATION-RELATED CONTENTS 14. Spontaneous corrections 15. Admitting lack of memory 16. Raising doubts about one’s own testimony 17. Self-deprecation 18. Pardoning the perpetrator OFFENCE-SPECIFIC ELEMENTS 19. Details characteristic of the offence Artículos de investigación 89 Initially, the CBCA criteria were intended for populations of child alleged victims of sexual abuse. However, CBCA criteria have been applied to other types of events and age ranges. This generalization has been extended to professional practice too. Thus, the guidelines of the Institute of Forensic Medicine in Spain, which is the official public institution responsible for forensic evidence, recommends SVA as part of the protocol for women alleging intimate partner violence (Arce & Fariña, 2012). Moreover, there is no consensus regarding the term minor, particularly since studies using the term range from 2 to 18-year-olds and the concept of minor is generally associated to the legal age of criminal responsibility. In relation to the context of application, only field studies involve real cases of sexual abuse, since it would be unethical to subject children to conditions or instructions of victims of sexual abuse. Hence, most research is experimental and certain authors have expressed their reservations regarding validity (Konecni & Ebbesen, 1992). Moreover, real eyewitnesses and subjects under high fidelity laboratory conditions have been found to perform different tasks (Fariña, Arce, & Real, 1994). In order to overcome this limitation, some experimental studies have recreated high fidelity simulated conditions in order to mimic the context of recall of child alleged victims of sexual abuse. These conditions have been defined as personal involvement, negative emotional tone of an event, and extensive loss of control over the situation (Steller, 1989). Accordingly, this achieves face validity, with external validity remaining entirely untested (Konecni & Ebbesen, 1992). Nevertheless, in spite of the weaknesses of this experimental paradigm in generalizing CBCA outcomes to the forensic context, experimental laboratory studies are useful for assessing certain variables that may lead to further internal validity research (Griesel et al., 2013). In contrast, the limitation of field studies resides in their difficulty in sustaining the ground truth in accurate objective criteria. These differences in research paradigms imply more credibility criteria are observed in field studies than in experimental ones (Vrij, 2005). The CBCA criteria are measured on two response scales, presence vs. absence, and the degree of presence. The unit of analysis is the full statement for the first major category, general characteristics, and for the remainder frequency counts. The presence of reality criteria is assumed to be indicative of memory based on real-life events, but the absence of criteria does not imply recall is based on fabricated accounts. Additionally, fictitious memory may contain reality criteria. Thus, the evaluation rests on a clinical judgement (Köhnken, 2004), which is semi-objective. Nonetheless, a replicable objective evaluation system, i.e., stringent is a fundamental standard for forensic practice. Alternatively a variety of decision rules have been proposed such as the presence of 3 criteria to judge a statement as true (Arntzen, 1983); of 7 criteria, criteria 1 to 5, plus 2 others; or the first three ones, plus 4 others (Zaparniuk, Yuille, & Taylor, 1995). Unfortunately, these decision rules lack empirical rigour. Furthermore, the application of criteria derived from the Undeutsch hypothesis to the forensic assessment of the credibility of a testimony (to be more precise, the truthfulness of a testimony given that credibility is a legal concept defined by the ruling of judges and the courts) must fulfil the legal standards for scientific evidence to be admitted as such in court. The standards governing the admission of expert testimony in court have been laid down by the Supreme Court of the United States in Daubert v. Merrel Dow Pharmaceuticals (1993) and are as follows: 1) Is the scientific hypothesis testable?; 2) has the proposition been tested?; 3) is there a known error rate?; 4) has the hypothesis and/or technique been subjected to peer review and publication; and 5) is the theory upon which the hypothesis and/or BÁRBARA GONZÁLEZ AMADO 90 technique is based on generally accepted among the appropriate scientific community? Several authors have raised their doubts on whether SVA/CBCA fulfil the above criteria, and this has led to seemingly contradictory viewpoints (Honts, 1994; Vrij, 2008). Bearing in mind these contradictory results and interpretations, and the wealth of scientific evidence, the aim of this meta-analysis was to review the literature on both experimental and field studies in order to assess the degree to which the Undeutsch hypothesis meets the judicial standards of evidence by estimating the effect size of the CBCA criteria. As initially the hypothesis was formulated for a population of children, though it has also been applied to adults, this review was restricted to studies on samples of children. Method Literature search The aim of the scientific literature search was to identify all of the empirical studies assessing the efficacy of CBCA criteria in discriminating between true statements of actual experiences, and the invented, imagined, fictitious, fabricated or false accounts of children. An exhaustive multi-method search was undertaken in the following international psychology databases of reference: PsycInfo and all of the databases of Web of Science, the Spanish language databases of reference Psicodoc (database of the Official Spanish College of Psychology), the Italian databases (ACPN, Archivio Collettivo Nazionale dei Periodici), the German Psychlinker databases, and the French human, social sciences, and economics Francis databases; the Google Scholar meta-search engine of scientific articles; a manual search in books; crosschecking all the references included in published reviews of articles and manuals; and directly contacting authors to request copies of unavailable studies. The keywords entered in the search engines were: Criteria Based Content Analysis or CBCA (kriteriumbasierte inhaltsanalyse, análisis de contenido basado en criterios, analisi del contenuto basata su criteri), credibility (glaubwürdigkeit, credibilidad, credibilità), content analysis (inhaltsanalyse, análisis de contenido, analisi del contenuto), child sexual abuse (kindesmissbrauch, abusos sexuales a menores, abuso sessuale perpetrato su minori), child testimony (kindliche zeugenaussage, testimonio del menor, bambini testimoni). The searches in the databases of non-English speaking countries were undertaken in the corresponding language. In line with the method of successive approximations, all of the keywords in the selected articles were revised in search of other potential descriptors. However, successive searches with these new descriptors failed to produce any further studies for the metaanalysis. Inclusion and exclusion criteria Though the Undeutsch hypothesis was initially formulated for children, the exact age group it encompasses has never been clearly specified. Nevertheless, as it was intended for judicial contexts and victims of sexual abuse, it is understood that it refers to children under the age of consent. First, the literature has taken the legal concept of minor as below the age of criminal responsibility (< 18 years) (Raskin & Esplin, 1991). Likewise, the lower age group was not related to the model, but the hypothesis is supported in that the memory of genuine life experiences differs in content and quality to memory of fictitious accounts, with Artículos de investigación 91 the criteria for discriminating between both types of memory being derived from the witness’ verbal account. Thus, the child is expected to have the sufficient narrative capability to express these criteria (Köhnken, 2004). Once again, the lower age limit was crosschecked in the studies reviewed, observing the lowest was a 2-year-old child (Buck, Warren, Betman, & Brigham, 2002; Lamers-Winkelman & Buffing, 1996), whose narrative skills, memory, gaps in memory, and recovery may be insufficient. Notwithstanding, given that the scientific literature has set a minimum age of 2 years and a maximum of 18 years, all of the studies with witnesses between these ages were included. Thus, studies with samples of children or that calculated the effect size for the subsamples of children were included. Second, delimiting the testimony to sexual abuse would compel studies to focus on this type of victim, as it would be unethical to subject children to memories of feigned victims. Consequently, the studies reviewed can be subdivided into low fidelity experimental studies (i.e., the scenarios neither involved sexual abuse nor was the implication or motivation of the participants controlled), high fidelity experiments (i.e, the scenarios do not involve sexual abuse, but they create an emotionally charged contexts close to the victimization of sexual abuse, and the implication of the participants is controlled), and field studies (i.e., real cases of sexual abuse where the ground truth is based on judicial judgements, the confession of the accused, medical evidence, and polygraph tests). The effects of the context of the research (i.e., field vs. laboratory high fidelity studies) on the results of the quality of an eyewitness’ identification have been found to be significant, and even contradictory (Fariña et al., 1994). Thus, it would be plausible to believe that this same bias may also affect the testimony of children alleged victims of sexual abuse. Hence, according to the circumstances, the context of the research was considered as a moderator. Third, for the criteria to discriminate between memories of real events and fabricated accounts, SVA proposes statements should be obtained using a free narrative interview e.g., step-wise interview, cognitive interview, Memorandum of Good Practices (Köhnken, 2004; Steller, 1989; Undeutsch, 1989), as they facilitate the emergence of criteria (Vrij, 2005). Likewise, compliance with other SVA criteria was strictly observed, i.e., studies noncompliant with any of the basic characteristics of the interview, such as inappropriate prompts or suggestions, were excluded. Fourth, the studies admitted as legal evidence should be published in scientific peerreviewed journals (Daubert v. Merrell Dow Pharmaceuticals, 1993). Notwithstanding, the literature has identified as key references studies that have not been published in these journals (i.e., Boychuk, 1991; Esplin, Houed, & Raskin, 1988), and these have been included in the meta-analysis. Thus, according to the circumstances, compliance with the peer-review publication Daubert standard was taken as a moderator. Fifth, the effect size was calculated from the data obtained and, if required, the authors were contacted to request the effect size or data for computing it as well as to clarify errors or queries regarding the data. Sixth, studies with samples shared with other studies were excluded (i.e., Hershkowitz, 1999), to avoid empirical redundancy (duplicity in publishing data) ‒ only the original study and the outlier values [IQR±1.5] were included. An independent analysis and control of outliers was carried out for each meta-analysis. BÁRBARA GONZÁLEZ AMADO 98 In terms of forensic applications, the analysis of the credibility of a testimony is admissible as incriminating evidence, being inoperative in classifying false statements. Thus, in terms of fidelity with the Undeutsch hypothesis and its application to forensic assessment, the results should be in the direction of the classification of true statements. Fifth, the results of some meta-analysis may be subject to a degree of variability given that Ns < 400 do not guarantee the stability of sampling estimates (Hunter & Schmidt, 2004). Nevertheless, as the results of the inter-meta-analysis were consistent, these may affect the statistical data, with expected effects in terms of the Undeutsch hypothesis. Sixth, the reliability of each criterion was an estimate, given the aforementioned reporting problems in the primary studies. Seventh, the results were somewhat biased towards supporting the hypothesis since some studies failed to publish so called conflicting data i.e., data with criteria that failed to discriminate significantly between real-life self-experienced events and fabricated accounts were excluded (e.g., Akehurst, Manton, & Quandte, 2011). In any case, none of these studies reported results that contradicted the Undeutsch hypothesis, but rather criteria that failed to discriminate significantly between both types of memory. Nevertheless, most of these limitations are reflected in the increase in the error variance, reducing the estimated effect sizes which means the true effect sizes would have been greater, lending even more support to the Undeutsch hypothesis. As for the practical implications of this meta-analysis, the findings support the Undeutsch hypothesis and several Daubert standards, but this does not imply the use of Checklist of CBCA criteria can be directly generalized to the context of forensic evaluation. First, the categorical system proposed is not methodic, i.e., it fails to comply with stringent methodic conditions: mutual exclusion, homogeneity, pertinence, objectivity, fidelity and productivity (Bardin, 1997). For instances, as non-mutual exclusion between categories is guaranteed, the duplicity of measures may arise; the criteria are neither objective nor exhaustive ‒ e.g., Roma et al. (2011) and Horowitz et al. (1997) have proposed the integration or redefinition of criteria due to rating difficulties ‒; the checklist may need additional criteria; or the checklist lacks internal consistency, in other words, it is not reliable. Second, the forensic application of a checklist of categories is driven by clinical judgements (Köhnken, 2004), or on quantitative decision rules that are not supported by empirical data (Arntzen, 1983; Zaparnuik et al., 1995). However, in the field of forensics an objective and strict decision criterion based on stringent standards of evidence should prevail over subjective clinical judgements, i.e., the rate of classification of false statements as true (false positives) should be 0 (i.e., the burden of proof is on the prosecution; it is entirely inadmissible to present incriminating expert forensic testimony on the basis of unsubstantiated evidences). It would be a decision rule based on data for controlling false positives, and to ensure reliable coding (i.e., within-raters, betweenraters, and between-context consistency), which would offset the potential effects of a truth bias or a response bias associated to the application of reality criteria (Griesel et al., 2013; Rassin, 1999; Sporer, 2004). The results of previous meta-analyses have shown it is possible, i.e., in field studies approximately 97% of truthful statements contained more reality criteria than fabricated accounts, with an approximately 90% total independence between the distributions of both groups of statements. Artículos de investigación 99 References (References marked with an asterisk indicate studies included in the meta-analysis) *Akehurst, L., Bull, R., Vrij, A., & Köhnken, G. (2004). The effects of training professional groups and lay persons to use Criteria-Based Content Analysis to detect deception. Applied Cognitive Psychology, 18, 877–891. doi: 10.1002/acp.1057 *Akehurst, L., Manton, S., & Quandte, S. (2011). Careful calculation or a leap of faith? A field study of the translation of CBCA ratings to final credibility judgements. Applied Cognitive Psychology, 25, 236–243. doi: 10.1002/acp.1669 Anson, D. A., Golding, S. L., & Gully, K. J. (1993). Child sexual abuse allegations: Reliability of Criteria-Based Content Analysis. Law and Human Behavior, 17, 331–341. doi: 10.1007/BF01044512 Arce, R., & Fariña, F. (2012). Psicología social aplicada al ámbito jurídico [Applied social psychology to the legal context]. In A. V. Arias, J. F. Morales, E. Nouvilas, & J. L. Martínez (Eds.), Psicología social aplicada (pp. 157–182). Madrid, Spain: Panamericana. Arce, R., Fariña, F., & Fraga, A. (2000). Género y formación de juicios en un caso de violación [Gnder and juror judgment making in a case of rape]. Psicothema, 12, 623– 628. Arce, R., Velasco, J., Novo, M., & Fariña, F. (2014). Elaboración y validación de una escala para la evaluación del acoso escolar [Development and validation of a scale to assess bullying]. Revista Iberoamericana de Psicología y Salud, 5, 71–104. Arntzen, F. (1970). Psychologie der zeugenaussage. Einführung in die forensische aussagepsychologie [Psychology of eyewitness testimony. Introduction to forensic psychology of statement analysis]. Göttingen, Germany: Hogrefe. Arntzen, F. (1983). Psychologie der zeugenaussage: Systematik der glaubwürdigkeitsmerkmale [Psychology of witness statements: The system of reality criteria]. Munich, Germany: C. H. Beck. Bardin, L. (1977). L'Analyse de contenu [Content analysis]. Paris, France: Presses Universitaires de France. Bembibre, J., & Higueras, L. (2010). Eficacia diferencial de la entrevista cognitiva en función de la profesión del entrevistador: Policías frente a psicólogos [Differential efficacy of the cognitive interview as a function of the interviewers’ profession: Polices vs. psychologists]. In F. Expósito, M. C. Herrera, G. Buela-Casal, M. Novo, & F. Fariña (Eds.), Psicología jurídica. Áreas de intervención (pp. 141–149). Santiago de Compostela, Spain: Consellería de Presidencia, Xustiza e Administracións Públicas. Bensi, L., Gambetti, E., Nori, R., & Giusberti, F. (2009). Discerning truth from deception: The sincere witness profile. European Journal of Psychology Applied to Legal Context, 1, 101–121. *Blandón-Gitlin, I., Pezdek, K., Rogers, M., & Brodie, L. (2005). Detecting deception in children: An experimental study of the effect of event familiarity on CBCA ratings. Law and Human Behavior, 29, 187–197. doi: 10.1007/s10979-005-2417-8 *Boychuk, T. D. (1991). Criteria Based Content Analysis of children’s statements about sexual abuse: A field-based validation study. Unpublished doctoral dissertation, Arizona State University. Buck, J. A., Warren, A. R., Betman, S. I., & Brigham, J. C. (2002). Age differences in Criteria-Based Content Analysis scores in typical child sexual abuse interviews. Journal BÁRBARA GONZÁLEZ AMADO 100 of Applied Developmental Psychology, 23, 267–283. doi: 10.1016/S01933973(02)00107-7 *Casado del Pozo, A. M., Romera, R. M., Vázquez, B., Vecina, M., & De Paúl, P. (2004). Análisis estadístico de una muestra de 100 casos de abuso sexual infantil [Statistical analysis of a one hundred real cases of child sexual abuse]. In B. Vázquez (Ed.), Abuso sexual infantil. Evaluación de la credibilidad del testimonio (pp. 73–105). Valencia, Spain: Centro Reina Sofía para el Estudio de la Violencia. Cohen, J. (1988). Statistical power analysis for the behavioral sciences (2nd ed.). Hillsdale, NJ: LEA. *Craig, R. A., Scheibe, R., Raskin, D. C., Kircher, J. C., & Dodd, D. H. (1999). Interviewer questions and content analysis of children’s statements of sexual abuse. Applied Developmental Science, 3, 77–85. doi: 10.1207/s1532480xads0302_2 Daubert v. Merrell Dow Pharmaceuticals, 509 U.S. 579 (1993). Dettenborn, H., Froehlich, H., & Szewczyk, H. (1984). Forensische psychologie [Forensic Psychology]. Berlin, Germany: Deutscher Verlag der Wissenschaften. *Di Blasio, P., & Conti, A. (2000). L'applicazione del "Criteria-Based Content Analysis" (C.B.C.A.) a racconti di storie vere e inventate [Application of Criteria-Based Content Analysis to accounts of real and fabricated stories]. Maltrattamento e Abuso all’infanzia, 2(3). 57-78. doi: 10.1400/62968 *Erdmann, K., Volbert, R., & Böhm, C. (2004). Children report suggested events even when interviewed in a non-suggestive manner: What are its implications for credibility assessment? Applied Cognitive Psychology, 18, 589–611. doi: 10.1002/acp.1012 *Esplin, P. W., Houed, T., & Raskin, D. C. (1988). Application of statement validity assessment. Paper presented at the NATO Advanced Study Institute on Credibility Assessment, Maratea, Italy. Fariña, F., Arce, R., & Real, S. (1994). Ruedas de identificación: De la simulación y la realidad [Linepus: A comparison of high-fidelity research and research in a real context]. Psicothema, 6, 395–402. Fisher, R. P., Geiselman, R. E., & Amador, M. (1989). Field test of the cognitive interview: Enhancing the recollection of actual victims and witness of crime. Journal of Applied Psychology, 74, 722–727. doi: 10.1037/0021-9010.74.5.722 Fritz, C. O., Morris, P. E., & Richler, J. J. (2012). Effect size estimates: Current use, calculations, and interpretation. Journal of Experimental Psychology: General, 141, 2– 18.doi: 10.1037/a0024338 *Granhag, P. A., Strömwall, L. A., & Landström, S. (2006). Children recalling an event repeatedly: Effects on RM and CBCA scores. Legal and Criminological Psychology, 11, 81–98. doi: 10.1348/135532505X49620 Griesel, D., Ternes, M., Schraml, D., Cooper, B. S., & Yuille, J. C. (2013). The ABC’s of CBCA: Verbal credibility assessment in practice. In B. S., Cooper, D. Griesel, & M. Ternes (Eds.), Applied issues in investigative interviewing, eyewitness memory, and credibility assessment (pp. 293–323). New York, NY: Springer. doi: 10.1007/978-14614-5547-9_12 Hans, V. P., & Vidmar, N. (1986). Judging the jury. New York, NY: Plenum Press. Hedges, L. V., & Olkin, I. (1985). Statistical methods for meta-analysis. San Diego, CA: Academic Press. Hershkowitz, I. (1999). The dynamics of interviews involving plausible and implausible allegations of child sexual abuse. Applied Developmental Science, 3, 86–91. doi: 10.1207/s1532480xads0302_3 Artículos de investigación 101 Honts, C. R. (1994). Assessing children’s credibility: Scientific and legal issues in 1994. North Dakota Law Review, 70, 879–903. Horowitz, S. W., Lamb, M. E., Esplin, P. W., Boychuk, T. D., Krispin, O., & Reiter-Lavery, L. (1997). Reliability of criteria-based content analysis of child witness statements. Legal and Criminological Psychology, 2, 11–21. doi: 10.1111/j.20448333.1997.tb00329.x Hunter, J. E., & Schmidt, F. L. (2004). Methods of meta-analysis: Correcting errors and bias in research findings. Thousand Oaks, CA: Sage. Kaplan, M. F. (1982). Cognitive processes in the individual juror. In N. L. Kerr & R. M. Bray (Eds.), The psychology of the courtroom (pp. 197–220). New York, NY: Academic Press. Köhnken, G. (1996). Social psychology and the law. In G. R. Semin & K. Fiedler (Eds.), Applied social psychology (pp. 257–282). Thousand Oaks, CA: Sage. Köhnken, G. (2004). Statement Validity Analysis and the detection of the truth. In P. A. Granhag & L. A. Strömwall (Eds.), The detection of deception in forensic contexts (pp. 41–63). Cambridge, UK: Cambridge University Press. doi: 10.1017/CBO9780511490071.003 Köhnken, G., & Steller, M. (1988). The evaluation of the credibility of child witness statements in the German procedural System. Issues in Legal and Criminological Psychology, 13, 37–45. Konecni, V. J., & Ebbesen, E. B. (1992). Methodological issues on legal decision-making, with special reference to experimental simulations. In F. Lösel, D. Bender, & T. Bliesener (Eds.), Psychology and law. International perspectives (pp. 413–423). Berlin, Germany: Walter de Gruyter. Kraemer, H. C., & Andrews, G. (1982). A non-parametric technique for meta-analysis effect size calculation. Psychological Bulletin, 91, 404-412. doi: 10.1037/0033-2909.91.2.404 *Lamb, M. E., Sternberg, K. J., Esplin, P. W., Hershkowitz, I., & Orbach, Y. (1997). Assessing the credibility of children’s allegations of sexual abuse: A survey of recent research. Learning and Individual Differences, 9, 175–194. doi: 10.1016/S10416080(97)90005-4 Lamers-Winkelman, F., & Buffing, F. (1996). Children’s testimony in the Netherlands: A study of Statement Validity Analysis. Criminal Justice and Behavior, 23, 304–321. doi: 10.1177/0093854896023002004 *Mazzoni, G., & Ambrosio, K. (2002). L’analisi del resoconto testimoniale in bambini: Impiego del metodo di analisi del contenuto C.B.C.A. in bambini di 7 anni [Assessment of child witness statements: Application of CBCA method in a 7 year old children sample]. Psicologia e Giustizia, 3 (2). *Mazzoni, G., & Pezzati, S. (2002). Esame della validità del C.B.C.A. in racconti di bambini di 4-5 anni [An exam of validity of the CBCA in the account’s four-five year old children.]. Eta’ Evolutiva, 73, 5–17. Memon, A., Meissner, C. A., & Fraser, J. (2010). Cognitive interview. A meta-analytic review and study space analysis of the past 25 years. Psychology, Public Policy, and Law, 16, 340–372. doi: 10.1037/a0020518 Novo, M., & Seijo, D. (2010). Judicial judgement-making and legal criteria of testimonial credibility. European Journal of Psychology Applied to Legal Context, 2, 91–115. Raskin, D. C., & Esplin, P. W. (1991). Statement Validity Assessment: Interview procedures and content analysis of children’s statements of sexual abuse. Behavioral Assessment, 13, 265–291. BÁRBARA GONZÁLEZ AMADO 102 Raskin, D. C., Esplin, F. W., & Horowitz, S. (1991). Investigative interviews and assessment of children in sexual abuse cases. Unpublished manuscript, Department of Psychology, University of Utah. Rassin, E. (1999). Criteria Based Content Analysis: The less scientific road to truth. Expert Evidence, 7, 265–278. doi: 10.1023/A:1016627527082 *Roma, P., San Martini, P., Sabatello, U., Tatarelli, R., & Ferracuti, S. (2011). Validity of Criteria-Based Content Analysis (CBCA) at trial in free-narrative interviews. Child Abuse & Neglect, 35, 613–620. doi: 10.1016/j.chiabu.2011.04.004 Rosenthal, R. (1994). Parametric measures of effect size. In H. Cooper & L. V. Hedges (Eds.), The handbook of research synthesis (pp. 231–244). New York, NY: Russell Sage Foundation. *Santtila, P., Roppola, H., Runtti, M., & Niemi, P. (2000). Assessment of child witness statements using Criteria-Based Content Analysis (CBCA): The effects of age, verbal ability, and interviewer’s emotional style. Psychology, Crime & Law, 6, 159–179. doi: 10.1080/10683160008409802 Sporer, S. L. (2004). Reality monitoring and detection of deception. In P. A. Granhag & L. A. Strömwall (Eds.), The detection of deception in forensic contexts (pp. 64-102). Cambridge, UK: Cambridge University Press. doi: 10.1017/CBO9780511490071.004 Steller, M. (1989). Recent developments in statement analysis. In J. C. Yuille (Ed.), Credibility assessment (pp. 135–154). Deventer, Holland: Kluwer. doi: 10.1007/978-94015-7856-1_8 Steller, M., & Böhm, C. (2006). Cincuenta años de jurisprudencia del Tribunal Federal Supremo alemán sobre la psicología del testimonio. Balance y perspectiva [Fifty years of the Federal German Court jurisprudence about forensic psychology]. In T. Fabian, C. Böhm, & J. Romero (Eds.), Nuevos caminos y conceptos en la psicología jurídica (pp. 53–67). Münster, Germany: LIT Verlag. Steller, M., & Köhnken, G. (1989). Criteria-Based Content Analysis. In D. C. Raskin (Ed.), Psychological methods in criminal investigation and evidence (pp. 217–245). New York, NY: Springer-Verlag. *Steller, M., Wellershaus, P., & Wolf, T. (1988). Empirical validation of Criteria-Based Content Analysis. Paper presented at the NATO Advanced Study Institute on Credibility Assessment, Maratea, Italy. *Strömwall, L. A., Bengtsson, L., Leander, L., & Granhag, P. A. (2004). Assessing children’s statements: The impact of a repeated experience on CBCA and RM ratings. Applied Cognitive Psychology, 18, 653–668. doi: 10.1002/acp.1021 Szewczyk, H. (1973). Kriterien der Beurteilung kindlicher zeugenaussagen [Criteria for the evaluation of child witnesses]. Probleme und Ergebnisse der Psychologie, 46, 47–66. *Tye, M. C., Amato, S. L., Honts, C. R., Devitt, M. K., & Peters, D. (1999). The willingness of children to lie and the assessment of credibility in an ecologically relevant laboratory setting. Applied Developmental Science, 3, 92–109. doi: 10.1207/s1532480xads0302_4 Undeutsch, U. (1967). Beurteilung der glaubhaftigkeit von aussagen [Evaluation of statement credibility/ Statement validity assessment]. In U. Undeutsch (Ed.), Handbuch der Psychologie, Vol. 11: Forensische Psychologie (pp. 26–181). Göttingen, Germany: Hogrefe. Undeutsch, U. (1989). The development of statement reality analysis. In J. Yuille (Ed.). Credibility assessment (pp.101–119). Dordrech, Holland: Kluwer Academic Publishers. Vrij, A. (2005). Criteria-Based Content Analysis: A qualitative review of the first 37 studies. Psychology, Public Policy, and Law, 11, 3–41. doi: 10.1037/1076-8971.11.1.3 Artículos de investigación 103 Vrij, A. (2008). Detecting lies and deceit: Pitfalls and opportunities (2nd ed.). Chichester, UK: John Wiley and Sons. *Vrij, A., Akehurst, L., Soukara, S., & Bull, R. (2002). Will the truth come out? The effect of deception, age, status, coaching and social skills on CBCA scores. Law and Human Behavior, 26, 261–283. doi: 10.1023/A:1015313120905 *Vrij, A., Akehurst, L., Soukara, R., & Bull, R. (2004). Detecting deceit via analyses of verbal and nonverbal behavior in children and adults. Human Communication Research, 30, 8–41. doi: 10.1111/j.1468-2958.2004.tb00723.x Weick, K. E. (1985). Systematic observational methods. In G. Lindzey & E. Aronson (Eds.), The handbook of social psychology (Vol. 1, pp. 567–634). Hillsdale, NJ: LEA. Zaparniuk, J., Yuille, J. C., & Taylor, S. (1995). Assessing the credibility of true and false statements. International Journal of Law and Psychiatry, 18, 343–352. doi: 10.1016/0160-2527(95)00016-B BÁRBARA GONZÁLEZ AMADO 104 Appendix 1. Moderator Variables Primary studies N Age Sex Paradigm Design Raters Coding Scale Coders training Type of interview Criteria Transcript/ video Akehurst, Bull, Vrij, and Köhnken (2004) 151 7-11 M: 23 F: 26 Experimental: active Within 58 0-4 Extensive training in CBCA coding Step-wise interview 13 Transcript Akehurst, Bull, Vrij, and Köhnken (2004) 132 Experimental: video Within 13 Transcript Akehurst, Manton, and Quandte (2011)2,6 31 6-17 M: 5 F: 26 Field Between 2 1-5 Extensive training in CBCA coding Step-wise interview 3 Transcript Blandon-Gitlin, Pezdek, Rogers, and Brodie (2005)5 94 9-12 – Experimental: active Between 2 0-1 Extensive training in CBCA coding Step-wise interview 16 Transcript Boychuk (1991) 75 4-16 M: 15 F: 60 Field Between 2 0-1 Expert raters No standardized interview procedures 19 Transcript Casado del Pozo, Romera, Vázquez Mezquita, Vecina, and de Paúl (2002)6 96 4-18 M: 28 F: 72 Field Between 2 0-1 Expert raters SVA guidelines 19 – Craig, R. A., Scheibe, R., Raskin, D. C., Kircher, J. C., y Dodd, D. H. (1999)1 48 3-16 M: 11 F: 37 Field Between 4 0-1 8 hours training SVA guidelines 14 Transcript Di Blasio and Conti (2000) 88 9 M: 25 F: 19 Experimental: memory Within 2 0-1 Expert raters Step-wise interview 19 – Erdmann, Volbert, and Böhm (2004) 70 6-8 M: 36 F: 31 Experimental: memory Within 2 0-1/0-2 Experts raters Step-wise interview 15 Transcript Esplin, Houed, and Raskin (1988) 40 3-15 – Field Between 1 0-2 Intensive training in CBCA coding – 19 Transcript Granhag, Strömwall, and Landström (2006) 80 12-13 M: 42 F: 38 Experimental: staged Between 2 0-2 Intensive training in CBCA coding Cognitive Interview 10 – Lamb, Sternberg, Esplin, Orbach, and Hovav (1997)6 89 4-13 M: 28 F: 70 Field Between 2 0-1 Intensive training in CBCA coding No standardized interview procedures. 141 Transcript Primary studies N Age Sex Paradigm Design Raters Coding Scale Coders training Type of interview Criteria Transcript/ video Mazzoni and Pezzati (2002)5 60 4-5 M: 20 F: 21 Experimental: memory Within 2 0-2 Training in CBCA Step-wise interview 19 Transcript Artículos de investigación 105 Mazzoni and Ambrosio (2002)5 60 7 – Experimental: memory Within 2 0-2 – Step-wise interview 19 Transcript Roma, San Martini, Sabatello, Tatarelli, and Ferracuti (2011)1,6 109 4-14 M: 23 F: 86 Field Between 2 0-1 Expert raters Step-wise interview 14 Transcript Santtila, Roppola, Runtti, and Niemi (2000)1,5 136 7-14 M: 34 F: 34 Experimental: memory Within 2 0-1/0-2 Expert raters Step-wise interview 14 Transcript Steller, Wellershaus, and Wolf (1988) 176 10-13 – Experimental: memory Within 3 0-3 – SVA guidelines 16 – Strömwall, Bengtsson, Leander, and Granhag (2004)3,5 41 10-13 M: 45 F: 42 Experimental: active Between 2 0-1/0-2 Extensive training in CBCA coding Cognitive interview 15 Transcript Strömwall, Bengtsson, Leander, and Granhag (2004)4,5 46 Tye, Amato, Honts, Devitt, and Peters (1999) 28 6-10 M: 21 F: 27 Experimental: active Between 3 0-2 Expert raters SVA guidelines 12 Transcript Vrij, Akehurst, Soukara, and Bull (2002)5 36 5-6 M: 16 F: 20 Experimental: active Between 2 1-5 Extensive training in CBCA coding Step-wise interview 9 Transcript 56 10-11 M: 22 F: 34 57 14-15 M: 33 F: 24 Vrij, Akehurst, Soukara, and Bull (2004)5 35 5-6 M: 16 F: 19 Experimental: active Between 2 1-5 Extensive training in CBCA coding Step-wise interview 9 Transcript 54 10-11 M: 22 F: 32 55 14-15 M: 32 F: 23 Note. 1CBCA14 criteria version; 2limited to those criteria discriminating significantly between self-experienced and fabricated statements; 3event experienced once, 4event experienced four times; 5high fidelity study, 6restricted field study 8.2. CRITERIA-BASED CONTENT ANALYSIS (CBCA) REALITY CRITERIA IN ADULTS: A META-ANALYTIC REVIEW Bárbara G. Amado*, Ramón Arce*, Francisca Fariña**, & Manuel Vilariño* (2016) * University of Santiago de Compostela (Spain). ** University of Vigo (Spain). Abstract Criteria-Based Content Analysis (CBCA) is the tool most extensively used worldwide for evaluating the veracity of a testimony. CBCA, initially designed for evaluating the testimonies of victims of child sexual abuse, has been empirically validated. Moreover, CBCA has been generalized to adult populations and other contexts though this generalization has not been endorsed by the scientific literature. Thus, a meta-analysis was performed to assess the Undeutsch hypothesis and the CBCA checklist of criteria in discerning in adults between memories of self-experienced real-life events and fabricated or fictitious memories. Though the results corroborated the Undeutsch Hypothesis, and CBCA as a valid technique, the results were not generalizable, and the self-deprecation and pardoning the perpetrator criteria failed to discriminate between both memories. The technique can be complemented with additional reality criteria. The study of moderators revealed discriminating efficacy was significantly higher in field studies on sexual offences and intimate partner violence. The findings are discussed in terms of their implications as well as the limitations and conditions for applying these results to forensic settings. Keywords: Criteria-Based Content Analysis; adults; statements; credibility; meta-analysis. Criterios de realidad del CBCA en adultos: Una revisión meta-analítica Resumen El Criteria-Based Content Analysis (CBCA) constituye la herramienta mundialmente más utilizada para la evaluación de la credibilidad del testimonio. Originalmente fue creado para testimonios de menores víctimas de abuso sexual, gozando de amparo científico. Sin embargo, se ha generalizado su práctica a poblaciones de adultos y otros contextos sin un aval de la literatura para tal generalización. Por ello, nos planteamos una revisión meta-analítica con el objetivo de contrastar la Hipótesis Undeutsch y los criterios de realidad del CBCA para conocer su potencial capacidad discriminativa entre memorias de eventos autoexperimentados y fabricados en adultos. Los resultados confirman la hipótesis Undeutsch y validan el CBCA como técnica. No obstante, los resultados no son generalizables y los criterios autodesaprobación y perdón al autor del delito no discriminan entre ambas memorias. Además, se encontró que la técnica puede ser complementada con criterios adicionales de realidad. El estudio de moderadores mostró que la eficacia discriminativa era significativamente superior en estudios de campo en casos de violencia sexual y de género. Se BÁRBARA GONZÁLEZ AMADO 114 Study of moderators The study of moderators (criteria average as dependent variable; Table 3) showed a positive and significant mean true effect size, but not generalizable, in all of the moderators analysed. As for the magnitude of the effect sizes, excluding the witness condition with a medium effect size (δ > 0.50), all were small (0.20 > δ < 0.50). Arce and Fariña (2009) have suggested (and designed) the specifications of categorical systems based on bottom-up rather than ‘top-down’ procedures to ensure only categories that effectively discriminate between memories of experienced events and fabricated memories form part of the system. This maximizes the efficacy of the resulting categorical system by eliminating the noise produced by non-discriminating ‘top-down’ categories. Thus, the meta-analyses were repeated with the categories of content analysis with significant effect size i.e., the confidence interval for d did not contain zero. The results (Table 3) revealed a significant increase in the effect size of field studies, qc = .119, p < .05 (one-tailed; a larger effect size was expected with significant criteria), thus the effect size was significantly larger with significant criteria. Moreover, for significant criteria, the results (not all of the reality criteria were generalizable) became generalizable (the credibility interval had no zero). As for the experimental studies on the remaining moderators, the results did not corroborate the Hypothesis as the reality categories had been initially or subsequently screened to eliminate the non-significant ones. The meta-analytical technique does not take into account the theoretical foundations or the reliability of the studies included in the original theories, that is, all of the studies on categories of reality are included. Moreover, the experimental designs of studies on witnesses are not really on witnesses of self-experienced events, but on non-self-experienced events i.e., video-observed events (watched on video, not involving self-experienced events) that do not fulfil the original theory hypothesizing that reality criteria discern between memories of selfexperienced real-life events and fabricated or fictitious memories. Only one of the studies on witnesses involved self-experienced events (Gödert, Gamer, Rill, & Vossel, 2005), and for the total reality criteria, were found to discriminate significantly real witness from real offenders giving false memory, d = 0.59, 1-β = .78, and from uninvolved participants, d = 0.83, 1-β = .96. Nevertheless, reality criteria also discriminated between both memories of videoobserved events and fabricated events. The only study (Lee, Klaver, & Hart, 2008) comparing memories of self-experienced events (truth condition) and video-observed events (lie condition) found CBCA reality criteria, and the total CBCA score discriminated significantly between both memories in line with the Undeutsch Hypothesis. Artículos de investigación 115 Table 3. Results of the Meta-analysis of Moderators MODERATOR k n dw SDd SDpre SDres δ SDδ %Var 95% CId 80% CVδ CBCA significant criteria (17) 46 3,223 0.27 0.5187 0.2380 0.4433 0.36 0.5835 31 0.19, 0.35 -0.39, 1.11 14-criteria version 45 3,143 0.28 0.5567 0.2394 0.4906 0.36 0.6465 25 0.22, 0.34 -0.47, 1.19 DAUBERT STANDARD PUBLICATION CRITERION All criteria (22) 35 2,256 0.20 0.4575 0.2407 0.3733 0.26 0.4786 39 0.12, 0.28 -0.35, 0.87 SELF-EXPERIENCED EVENTS All criteria (22) 34 2,277 0.26 0.4647 0.2371 0.3879 0.33 0.5022 40 0.18, 0.34 -0.31, 0.97 NON SELF-EXPERIENCED EVENTS (WITNESS) All criteria (13) 11 625 0.39 0.5835 0.2707 0.5032 0.51 0.6548 65 0.23, 0.55 -0.33, 1.35 OFFENDERS All criteria (21) 11 1,067 0.27 0.4662 0.2024 0.3743 0.35 0.4975 41 0.15, 0.39 -0.29, 0.99 VICTIMS All criteria (18) 11 840 0.27 0.4781 0.2355 0.4012 0.35 0.5221 35 0.13, 0.41 -0.32, 1.02 FIELD STUDIES All field studies (18) 6 422 0.34 0.4948 0.2385 0.4153 0.45 0.5404 35 0.14, 0.54 -0.24, 1.14 Significant criteria (10)a 6 422 0.53 0.4774 0.2458 0.3834 0.69 0.4989 42 0.33, 0.73 0.05, 1.33 SEXUAL AND IPV FIELD STUDIES All criteria (17)b 5 263 0.67 0.3587 0.2871 0.1957 0.87 0.2459 72 0.41, 0.92 0.55, 1.18 Significant criteria (15)c 5 263 0.74 0.3654 0.2892 0.2134 0.96 0.2478 72 0.48, 0.99 0.64, 1.28 EXPERIMENTAL STUDIES All criteria (22) 39 2,721 0.25 0.4497 0.2336 0.3934 0.32 0.4933 37 0.17, 0.33 -0.31, 0.95 Note. asignificant criteria (CBCA criteria, as for additional criteria, studies were insufficient): 1-3, 5-8, 11, 12 and 19; bsignificant criteria (CBCA criteria): 1-9, 11-18; csignificant criteria (CBCA criteria): 1-9, 11-12, 14-17. BÁRBARA GONZÁLEZ AMADO 116 The high observed variability in effect sizes in field studies, which was mostly due to one study alone, suggested differences in experimental design (the crime context in this study was found to be different to the other studies). As the effect of context has been hypothesized (Köhnken, 1996; Volbert & Steller, 2014), and found (Arce, Fariña, & Vilariño, 2010; Vilariño et al. 2011) to mediate the discriminating efficacy of reality categories, the metaanalysis was repeated in field studies on sexual offences and intimate partner violence (IPV) cases (crimes committed in the privacy of one’s home according to the categorization of Arce and Fariña, 2005). The results showed a positive, significant and generalizable (not generalizable in all field studies) mean true effect size for studies under this condition. Moreover, the magnitude of the effect sizes were significantly larger in sexual offences and IPV cases than in all field studies in all the reality criteria (0.45 for all field studies vs. 0.87 for sexual offences and IPV cases), qc = .199, p < .01 (one-tailed; a higher effect size was expected in specific criminal contexts), and in the significant criteria, qc = .168, p < .05 (0.69 vs. 0.96). Likewise, reality criteria were significantly more efficacious, qc = .2622, p < .01, in sexual offences and IPV cases than in all other types of cases (0.32 vs. 0.87). Results (meta-analysis could not be performed because ks and ns were insufficient and research designs incomparable) for the comparison between statements of participants instructed to lie (lie coaching condition) with truthful statements were inconclusive2 in relation to the effectiveness of reality criteria to discriminate between truthful and false statements. Discussion The following conclusions may be drawn from the results of this study. First, the results confirmed the Undeutsch Hypothesis, that is, reality criteria discriminated between memories of self-experienced and fabricated events [File Drawer Analysis (FDA): to bring down this hypothesis to a trivial effect (McNatt, 2000), .05, for the average of the CBCA criteria, it would be necessary 184 studies with null effect; Hunter & Schmidt, 2015. It is unlikely to happen]. Besides fulfilling the DSPC, this Hypothesis was also valid for memories of victims/claimants and offenders (for witness of self-experienced events further research is required); and robust in both experimental studies (high internal validity), and field studies (high external validity). Notwithstanding, the reality criteria also discriminated between memories of video-observed events i.e., non-self-experienced events, and fabricated events for which the Hypothesis was not formulated, and research findings are inconclusive as to the validity of the Hypothesis with lie coached subjects. Second, though the results validated CBCA as a categorical system based on the Undeutsch Hypothesis, neither were all of the criteria validated, nor were they generalizable, and some even contradicted the Hypothesis. Thus, these criteria can be used neither in all types of contexts, nor indiscriminately. Both versions of the CBCA (all criteria or 14 criteria) were exactly the same (δ = 0.36) in discriminating between memories of self-experienced and fabricated events. Though the results open the door to the inclusion of new reality criteria, additional criteria have been proposed that fail to fulfil the Undeutsch Hypothesis (significant negative effect sizes i.e., not 2 Conclusions in the primary studies about non-significant results are inconclusive as the statistical power, 1β<.80, is insufficient to conclude (d = -0.44, 1-β = .41, Bogaard, Meijer, & Vrij, 2013; d = 0.37, 1-β = .26, Vrij, Akehurst, Soukara, & Bull, 2002; d = 0.11, 1-β = .06, Vrij, Kneller, & Mann, 2000). Artículos de investigación 117 reality criteria), so they cannot be included in the CBCA. Third, in field studies the discriminating power of reality criteria was significantly higher in sexual offences and IPV cases (FDA: to bring the results in sexual offences and IPV cases down to a trivial effect, it would be necessary 62 and 69 studies with null effect for all criteria and significant criteria, respectively. It is unlikely to occur) in comparison to other types of contexts (FDA: to reduce the efficacy of the reality criteria to discriminate between real and fabricated memories in any context of field studies to a trivial effect it would be necessary 35 studies with null effect. It is unlikely to happen). Succinctly, the areas of both populations do not overlap in 54% (U1 = 0.54), that is, they were totally independent, thus the efficacy of the reality criteria in discriminating between memories of self-experienced and fabricated events in sexual and IPV cases was total in 54% of the evaluations of credibility. Moreover, 75% of statements of selfexperienced events contained more reality criteria than fabricated events (probability of superiority, PS = 0.75), the probability of false positives was 28% (BESD). These results were highly robust i.e., not only establishing a positive and significant relation between reality criteria and true statements, but were also generalizable to all types of sexual offences and IPV cases, and were homogeneous (i.e., subject to little variability since the correlation between the effect sizes was .72). As for the implications for forensic practice, the results of the present meta-analysis reveal that the reality criteria were statistically effective for discriminating between memories of self-experienced and fabricated events, but this does not imply they are directly generalizable to forensic practice. Even under the best discriminating conditions i.e., field studies in sexual and IPV cases, the probability of false positives may reach .22, whilst this probability must be zero in forensic settings (Arce, Fariña, & Fraga, 2000). In general, only significant reality criteria i.e., scientifically attested evidence, were admissible for forensic practice (see note of Table 3), since the results were generalizable, whereas for all criteria they were not. However, as the credibility interval lower limit was 0.05, the practical utility of these categories was almost negligible (PS=.51), that is, in only 51% of true statements there were more reality criteria than in false statements, and under what specific conditions this contingency occurred remains unknown. However, the credibility interval lower limit of the reality criteria applied to cases of sexual offences and IPV, which were also generalizable both in terms of all the criteria and the significant criteria, was larger, PS = .73 and .75 (Hedges and Olkin’s δ = 0.59 and 0.65, test value = .51), for all the reality criteria and the significant criteria, respectively. However, these conclusions are not directly applicable to forensic practice as the decision criteria which in the forensic context must the ‘strict decision criterion’ in which a type II error (classify a false statement as true) is not admissible i.e., must be equal to zero. Regarding the strict decision criterion, Arce et al. (2010) found up to 13 CBCA reality criteria in fabricated statements of IPV cases, which means that at least 14 reality criteria would have to be detected in a statement to conclude that the testimony was true, with a correct classification of true positives (true statements classified as such) of 36%. Succinctly, the CBCA reality criteria were a poor tool for assigning the credibility of IPV victim testimony. Thus, to enhance efficacy, CBCA reality criteria must be complemented with additional criteria. In this line, Arce and Fariña (2009), Vilariño, (2010) and Vilariño et al. (2011) combined CBCA and SRA criteria, memory attributes, and additional reality criteria specific to IPV cases derived from real statements (judicial judgements as ground truth), to create and validate a categorical system specific for IPV cases, including sexual offences, with a strict decision criterion to reduce the rate of false negatives to 2%. In any way, only results with a strict decision criterion can be translated into forensic practice. BÁRBARA GONZÁLEZ AMADO 118 In terms of future research, the results of the present meta-analysis underscored the need for further studies with experimental designs assessing the efficacy of reality criteria in discriminating between memories of self-experienced events and video witnessed non-selfexperienced events; between self-experienced witnessed events vs. fabricated events; between memories of participants coached to lie and honest; and research driven to find new reality categories (bottom-up), mainly for a specific context i.e., crime victimization. This meta-analysis is subject to the following limitations. First, previous publications have biased the results in that the non-significant results or predictably inefficacious categories were eliminated (favouring the validation of the Undeutsch Hypothesis). Second, the feigning methodology (experimental studies) had no proven external validity (Sarwar, Allwood, & Innes-Ker, 2014), but only ‘face validity’ (Konecni & Ebbesen, 1992). Third, for some experimental literature, statements are insufficient material for reality content analysis (Köhnken, 2004), which favours the rejection of the Undeutsch Hypothesis. Fourth, there was no control on the effects of the interviewer on the contents of the statement, or on the reliability of the interviews, which were often carried out by poorly trained interviewers. Fifth, few studies comply with SVA standards that are a requirement for applying CBCA. Sixth, the results of some meta-analysis may be subject to a degree of variability, given that Ns < 400, did not guarantee stability in sample estimates (Hunter & Schmidt, 2015). Seventh, primary studies did not estimate the reliability of the codings, thus results’ reliability is uncertainty. References *Akehurst, L., Easton, S., Fullar, E., Drane, G., Kuzmin, K., & Litchfield, S. (2015). An evaluation of a new tool to aid judgements of credibility in the medico-legal setting. Legal and Criminological Psichology. Advance online publication. doi: 10.1111/lcrp.12079 Amado, B.G., Arce, R., & Fariña, F. (2015). Undeutsch hypothesis and Criteria Based Content Analysis: A meta-analytic review. European Journal of Psychology Applied to Legal Context, 7, 3-12. doi:10.1016/j.ejpal.2014.11.002 Anson, D.A., Golding, S.L., & Gully, K.J. (1993). Child sexual abuse allegations: Reliability of Criteria-Based Content Analysis. Law and Human Behavior, 17, 331-341. doi:10.1007/BF01044512 Arce, R., & Fariña, F. (2005). Peritación psicológica de la credibilidad del testimonio, la huella psíquica y la simulación: El Sistema de Evaluación Global (SEG) [Psychological evidence in court on statement credibility, psychological injury and malingering: The Global Evaluation System (GES)]. Papeles del Psicólogo, 26, 59-77. Arce, R., & Fariña, F. (2009). Evaluación psicológica forense de la credibilidad y daño psíquico en casos de violencia de género mediante el Sistema de Evaluación Global. In F. Fariña, R. Arce, & G. Buela-Casal (Eds.), Violencia de género. Tratado psicológico y legal (pp. 147-168). Madrid: Biblioteca Nueva. Arce, R., & Fariña, F. (2012). Psicología social aplicada al ámbito jurídico. In A.V. Arias, J.F. Morales, E. Nouvilas, & J.L. Martínez (Eds.), Psicología social aplicada (pp. 157182). Madrid, Spain: Panamericana. Artículos de investigación 119 Arce, R., Fariña, F., & Fraga, A. (2000). Género y formación de juicios en un caso de violación [Gender and juror judgment making in a case of rape]. Psicothema, 12, 623628. *Arce, R., Fariña, F., & Vilariño, M. (2010). Contraste de la efectividad del CBCA en la evaluación de la credibilidad en casos de violencia de género [Contrasting the efficiency of the CBCA in the assessment of credibility in violence against women cases]. Intervención Psicosocial, 19, 109-119. doi:10.5093/in2010v19n2a2 *Beaulieu-Prévost, D. (2001). Analyse de validité de la déclaration (SVA), mensonge et faux souvenirs: Validité et efficacité chez les adultes. (Doctoral dissertation). Retrieved from ProQuest Dissertations & Theses Global. (Order No. MQ60609) *Bensi, L., Gambetti, E., Nori, R., & Giusberti, F. (2009). Discerning truth from deception: The sincere witness profile. European Journal of Psychology Applied to Legal Context, 1, 101-121. Berliner, L., & Conte, J.R. (1993). Sexual abuse evaluation: Conceptual and empirical obstacles. Child Abuse and Neglect, 17, 111-125. doi:10.1016/0145-2134(93)90012-T *Biland, C., Py, J., & Rimboud, S. (1999). Evaluer la sincérité d’un témoin grâce à trois techniques d’analyse, verbales et non verbale [Three verbal or nonverbal techniques for evaluating sincerity]. European Review of Applied Psychology, 49, 115-122. *Blandón-Gitlin, I., Pezdek, K., Lindsay, D.S., & Hagen, L. (2009). Criteria-Based Content Analysis of true and suggested accounts of events. Applied Cognitive Psychology, 23, 901-917. doi:10.1002/acp.1504 *Bogaard, G., Meijer, E.H., & Vrij, A. (2013). Using an example statement increases information but does not increase accuracy of CBCA, RM, and SCAN. Journal of Investigative Psychology and Offender Profiling, 11, 151-163. doi:10.1002/jip.1409 Botella, J. & Gambara, H. (2006). Doing and reporting a meta-analysis. International Journal of Clinical & Health Psychology, 6, 425-440. *Caso, L., Vrij, A., Mann, S., & de Leo, G. (2006). Deceptive responses: The impact of verbal and non-verbal countermeasures. Legal and Criminological Psychology, 11, 99111. doi:10.1348/135532505X49936 Cohen, J. (1988). Statistical power analysis for the behavioral sciences (2nd ed.). Hillsdale, NJ: LEA. *Critchlow, N. (2011). Applying Criteria Based Content Analysis to assessing the veracity of rape statements (Unpublished doctoral dissertation). Manchester Metropolitan University, Manchester, UK. *Critchlow, N. (2011). [A field validation of CBCA when assessing authentic police rape statements: evidence for discriminant validity to prescribe veracity to adult narrative]. Unpublished raw data. *Dana-Kirby, L. (1997). Discerning truth from deception: Is Criteria-Based Content Analysis effective with adult statements? (Unpublished doctoral thesis). University of Oregon, Oregon. Daubert v. Merrell Dow Pharmaceuticals, 509 U.S. 579 (1993). *Evans, J., Michael, S.W., Meissner, C.A., & Brandon, S.E. (2013). Validating a new assessment method for deception detection: Introducing a psychologically based credibility assessment tool. Journal of Applied Research in Memory and Cognition, 2, 33-41. doi:10.1016/j.jarmac.2013.02.002 Fariña, F., Arce, R., & Novo, M. (2002). Heurístico de anclaje en las decisiones judiciales [Anchorage in judicial decision making]. Psicothema, 14, 39-46. BÁRBARA GONZÁLEZ AMADO 120 Fariña, F., Arce, R., & Real, S. (1994). Ruedas de identificación: De la simulación y la realidad [Linepus: A comparison of high-fidelity research and research in a real context]. Psicothema, 6, 395-402. *Gödert, H.W., Gamer, M., Rill, H.G., & Vossel, G. (2005). Statement Validity Assessment: Inter-rater reliability of Criteria-Based Content Analysis in the mock-crime paradigm. Legal and Criminological Psychology, 10, 225-245. doi:10.1348/135532505X52680 *Godoy, V., & Higueras, L. (2008). El análisis de contenido basado en criterios (CBCA) y la entrevista cognitiva aplicados a la credibilidad del testimonio en adultos. In F. Rodríguez, C. Bringas, F. Fariña, R. Arce, & A. Bernardo (Eds.), Psicología Jurídica: Entorno judicial y delincuencia (pp. 117-125). Retrieved from http://gip.uniovi.es/T5EJD.pdf Hans, V.P., & Vidmar, N. (1986). Judging the jury. New York: Plenum Press. Hedges, L.V., & Olkin, I. (1985). Statistical methods for meta-analysis. Orlando, FL: Academic Press. Höfer, E., Köhnken, G., Hanewinkel, R., & Bruhn, C. (1993). Diagnostik und attribution von glaubwürdigkeit. Unpublished final report. University of Kiel, Germany. *Honts, C.R., & Devitt, M.K. (1993). Credibility Assessment of Verbatim Statements (CAVS). Retrieved from http://www.dtic.mil/dtic/tr/fulltext/u2/a271575.pdf Horowitz, S.W., Lamb, M.E., Esplin, P.W., Boychuk, T.D., Krispin, O., & Reiter-Lavery, L. (1997). Reliability of criteria-based content analysis of child witness statements. Legal and Criminological Psychology, 2, 11-21. doi:10.1111/j.2044-8333.1997.tb00329.x Hunter, J.E., & Schmidt, F.L. (2015). Methods of meta-analysis: Correcting error and bias in research findings. Newbury Park, CA: Sage. *Johnston, S., Candelier, A., Powers-Green, D., & Rahmani, S. (2014). Attributes of truthful versus deceitful statements in the evaluation of accused child molesters. Sage Open, 4(3), 1-10. doi:10.1177/2158244014548849 *Juárez, J.R., Mateu, A., & Sala, E. (2007). Criterios de evaluación de la credibilidad en las denuncias de violencia de género. Retrieved from http://justicia.gencat.cat/web/.content/documents/arxius/sc-3-143-07-cas.pdf Köhnken, G. (1996). Social psychology and the law. In G.R. Semin & K. Fiedler (Eds.), Applied social psychology (pp. 257-282). Thousand Oaks, CA: Sage. Köhnken, G. (2004). Statement Validity Analysis and the detection of the truth. In P. A. Granhag & L. A. Strömwall (Eds.), The detection of deception in forensic contexts (pp. 41-63). Cambridge, UK: Cambridge University Press. doi:10.1017/CBO9780511490071.003 *Köhnken, G., Schimossek, E., Aschermann, E., & Höfer, E. (1995). The cognitive interview and the assessment of the credibility of adults’ statements. Journal of Applied Psychology, 80, 671-684. doi:10.1037/0021-9010.80.6.671 Konecni, V.J., & Ebbesen, E.B. (1992). Methodological issues on legal decision-making, with special reference to experimental simulations. In F. Lösel, D. Bender, & T. Bliesener (Eds.), Psychology and law. International perspectives (pp. 413-423). Berlin, Germany: Walter de Gruyter *Leal, S., Vrij, A., Warmelink, L., Vernham, Z., & Fisher, R.P. (2015). You cannot hide your telephone lies: Providing a model statement as an aid to detect deception in insurance telephone calls. Legal and Criminological Psychology, 20, 129-146. doi:10.1111/lcrp.12017 *Lee, Z., Klaver, J.R., & Hart, S.D. (2008). Psychopathy and verbal indicators of deception in offenders. Psychology, Crime & Law, 14, 73-84. doi:10.1080/10683160701423738 Artículos de investigación 121 McNatt, D.B. (2000). Ancient Pygmalion joins contemporary management: A meta-analysis of the result. Journal of Applied Psychology, 85, 314-322. doi:10.1037/0021-9010.85.2.314 *Merckelbach, H. (2004). Telling a good story: Fantasy proneness and the quality of fabricated memories. Personality and Individual Differences, 37, 1371-1382. doi:10.1016/j.paid.2004.01.007 Novo, M., & Seijo, D. (2010). Judicial judgement-making and legal criteria of testimonial credibility. The European Journal of Psychology Applied to Legal Context, 2, 91-115. *Porter, S., & Yuille, J.C. (1996). The language of deceit: An investigation of the verbal clues to deception in the interrogation context. Law and Human Behavior, 20, 443-458. doi:10.1007/BF01498980 *Porter, S., Yuille, J.C., & Lehman, D.R. (1999). The nature of real, implanted, and fabricated memories for emotional childhood events: Implications for the recovered memory debate. Law and Human Behavior, 23, 517-537. doi:10.1023/A:1022344128649 Raskin, D.C., Esplin, F.W., & Horowitz, S. (1991). Investigative interviews and assessment of children in sexual abuse cases. Unpublished manuscript, Department of Psychology, University of Utah, Utah. *Rassin, E., & van-der-Sleen, J. (2005). Characteristics of true versus false allegations of sexual offences. Psychological Reports, 97, 589-598. doi:10.2466/pr0.97.2.589-598 Sarwar, F., Allwood, C. M., & Innes-Ker, A. (2014). Effects of different types of forensic information on eyewitness’ memory and confidence accuracy. European Journal of Psychology Applied to Legal Context, 6, 17-27. doi: 10.5093/ejpalc2014a3 *Schelleman-Offermans, K., & Merckelbach, H. (2010). Fantasy proneness as a confounder of verbal lie detection tools. Journal of Investigative Psychology and Offender Profiling, 7, 247-260. doi:10.1002/jip.121 *Sporer, S.L. (1997). The less travelled road to truth: Verbal cues in deception detection in accounts of fabricated and self-experienced events. Applied Cognitive Psychology, 11, 373-397. doi:10.1002/(SICI)1099-0720(199710)11:5<373::AID-ACP461>3.0.CO;2-0 Steller, M., & Böhm, C. (2006). Cincuenta años de jurisprudencia del Tribunal Federal Supremo alemán sobre la psicología del testimonio. Balance y perspectiva. In T. Fabian, C. Böhm, & J. Romero (Eds.), Nuevos caminos y conceptos en la psicología jurídica (pp. 53-67). Münster, Germany: LIT Verlag. Steller, M., & Köhnken, G. (1989). Criteria-Based Content Analysis. In D.C. Raskin (Ed.), Psychological methods in criminal investigation and evidence (pp. 217-245). New York: Springer-Verlag. *Ternes, M. (2009). Verbal credibility assessment of incarcerated violent offenders’ memory reports (Unpublished doctoral thesis). University of British Columbia, Vancouver. Tukey, J.W. (1960). A survey of sampling from contaminated distributions. In I. Olkin, J.G. Ghurye, W. Hoeffding, W.G. Madoo, & H. Mann (Eds.), Contributions to probability and statistics (pp. 448-485). Stanford, CA: Stanford University Press. Undeutsch, U. (1967). Beurteilung der glaubhaftigkeit von aussagen. In U. Undeutsch (Ed.), Handbuch der psychologie, Vol. 11: Forensische psychologie (pp. 26-181). Göttingen, Germany: Hogrefe. *Vilariño, M. (2010). ¿Es posible discriminar declaraciones reales de imaginadas y huella psíquica real de simulada en casos de violencia de género? (Doctoral thesis, Universidad de Santiago de Compostela, Spain). Retrieved from http://hdl.handle.net/10347/2831 Vilariño, M., Novo, M., & Seijo, D. (2011). Estudio de la eficacia de las categorías de realidad del testimonio del Sistema de Evaluación Global (SEG) en casos de violencia de género BÁRBARA GONZÁLEZ AMADO 122 [Study of the efficacy of the testimony reality categories of the Global Evaluation System (GES) in violence against women cases]. Revista Iberoamericana de Psicología y Salud, 2, 1-26. Volbert, R., & Steller, M. (2014). Is this testimony truthful, fabricated, or based on false memory? Credibility assessment 25 years after Steller and Köhnken (1989). European Psychologist, 19, 207-220. doi:10.1027/1016-9040/a000200 Vrij, A. (2005). Criteria-Based Content Analysis: A qualitative review of the first 37 studies. Psychology, Public Policy, and Law, 11, 3-41. doi:10.1037/1076-8971.11.1.3 Vrij, A. (2008). Detecting lies and deceit: Pitfalls and opportunities (2nd ed.). Chichester, UK: John Wiley and Sons Vrij, A., Akehurst, L., Soukara, S., & Bull, R. (2002). Will the truth come out? The effect of deception, age, status, coaching and social skills on CBCA scores. Law and Human Behavior, 26, 261-283. doi:10.1023/A:1015313120905 *Vrij, A., Akehurst, L., Soukara, R., & Bull, R. (2004). Detecting deceit via analyses of verbal and nonverbal behavior in children and adults. Human Communication Research, 30, 841. doi:10.1111/j.1468-2958.2004.tb00723.x *Vrij, A., Edward, K., & Bull, R. (2001). People’s insight into their own behaviour and speech content while lying. British Journal of Psychology, 92, 373-389. doi:10.1348/000712601162248 *Vrij, A., Edward, K., Roberts, K.P., & Bull, R. (2000). Detecting deceit via analysis of verbal and nonverbal behavior. Journal of Nonverbal Behavior, 24, 239-263. doi:10.1023/A:1006610329284 *Vrij, A., Evans, H., Akehurst, L., & Mann, S. (2004). Rapid judgements in assessing verbal and nonverbal cues: Their potential for deception researchers and lie detection. Applied Cognitive Psychology, 18, 283-296. doi:10.1002/acp.964 *Vrij, A., & Heaven, S. (1999). Vocal and verbal indicators of deception as a function of lie complexity. Psychology, Crime and Law, 5, 203-215. doi:10.1080/10683169908401767 *Vrij, A., Kneller, W., & Mann, S. (2000). The effect of informing liars about Criteria-Based Content Analysis on their ability to deceive CBCA-raters. Legal and Criminological Psychology, 5, 57-70. doi:10.1348/135532500167976 *Vrij, A., & Mann, S., (2006). Criteria-Based Content Analysis: An empirical test of its underlying processes. Psychology, Crime and Law, 12, 337-349. doi:10.1080/10683160500129007 *Vrij, A., Mann, S., & Edward, K. (2000). I think it was a green scarf but I am not sure. Raising doubts about one’s own testimony during lying and truth telling. In A. Czerederecka, T. Jaskiewicz-Obydzinska, & J. Wójcikiewicz (Eds.), Forensic, psychology and law. Traditional questions and new ideas (pp. 205-207). Institute of forensic research, Poland: Krakow. *Vrij, A., Mann, S., Kristen, S., & Fisher, R.P. (2007). Cues to deception and ability to detect lies as a function of police interview styles. Law and Human Behavior, 31, 449-518.doi: 10.1007/s10979-006-9066-4 *Willén, R.M., & Strömwall, L.A. (2011). Offender’s uncoerced false confessions: A new application of statement analysis? Legal and Criminological Psychology, 17, 346-359. doi:10.1111/j.2044-8333.2011.02018.x. *Wojciechowski, B.W. (2014). Content analysis algorithms: An innovative and accurate approach to statement veracity assessment. European Polygraph, 8, 119-128. doi:0.2478/ep-2014-0010 8.3. PSYCHOLOGICAL INJURY IN VICTIMS OF CHILD SEXUAL ABUSE: A META-ANALYTIC REVIEW Bárbara G. Amado, Ramón Arce, & Andrés Herraiz Departament of Psicología Organizacional, Jurídico-Forense y Metodología. University of Santiago de Compostela Abstract In order to assess the effects of child/adolescent sexual abuse (CSA/ASA) on the victim’s probability of developing symptoms of depression and anxiety, to quantify injury in populational terms, to establish the probability of injury, and to determine the different effects of moderators on the severity of injury, a meta-analysis was performed. Given the abundant literature, only studies indexed in the scientific database of reference, the Web of Science, were selected. A total of 78 studies met the inclusion criteria: measured CSA/ASA victims; measured injury in terms of depression or anxiety symptoms; measured the effect size or included data for computing them; and provided a description of the sample. The results showed that CSA/ASA victims suffered significant injury, generally of a medium effect size, and generalizable; victims had 70% more probabilities of suffering from injury; and clinical diagnosis was a significantly more adequate measure of injury than symptoms. The probability of chronic injury (dysthymia) was greater than developing more severe injury i.e., major depressive disorder (MDD). In the category of anxiety disorders, injury was expressed with a higher probability in specific phobia. In terms of the victim’s gender, females had significantly higher rates of developing a depressive disorder (DD) and/or an anxiety disorder (AD), quantified in a 42% and 24% over the baseline, for a DD and AD, respectively. As for the type of abuse, the meta-analysis revealed that abuse involving penetration was linked to severe injury whereas abuse with no contact was associated to less serious injury. The clinical, social, and legal implications of the results are discussed. Keywords: child sexual abuse, adolescent sexual abuse, psychological injury, victimization, meta-analysis. Daño psicológico en víctimas de abuso sexual infantil: Una revision meta-analítica Resumen Con el objetivo de conocer los potenciales efectos de la victimización de abuso sexual infantil/adolescente (ASI/ASA) en el desarrollo de sintomatología depresiva y ansiosa así como cuantificar, en su caso, el potencial daño en términos poblacionales, la probabilidad de manifestación de daño y el efecto diferencial de moderadores en la severidad del daño manifestado, se planificó una revisión meta-analítica. Dada la gran proliferación de literatura se seleccionaron aquellos estudios indexados en la base de datos de referencia de calidad científica, la Web of Science. 78 estudios cumplieron los criterios de inclusión: medida de la victimización de ASI/ASA; medida del daño en sintomatología depresiva o ansiosa; medida del tamaño del efecto o inclusión de datos que permitieran computarlo; y descripción de la muestra. Los resultados mostraron que la victimización de ASI/ASA conlleva a un daño BÁRBARA GONZÁLEZ AMADO 130 Study of moderators Gender effects The results of the meta-analysis showed (see Table 4) a significant effect, positive, of a small size (Cohen’s category: r = .10), and generalizable in depression and anxiety in female CSA/ASA victims. In comparison, the meta-analysis revealed for male CSA/ASA victims a significant effect, positive, and of a small size in depression and anxiety, being generalizable in depression, but not so in anxiety (see Table 4). Thus, in the latter case, the results exhibited moderators mediated the direction of the effects. Having contrasted the significance of the differences between the effect sizes, the true correlation between female participants and male participants, sequelae in depression was found to be significantly higher, qs = 0.093, p < .05, in females, but not so for anxiety, qs = 0.073, ns. This translates into quantifying injury in females as suffering from 9% more injury in depression than males. In prevalence rates (odds ratio), injury in depression for female and male victims was 2.26 and 1.60 times greater than for non-victims, and 2.26 and 1.73 times greater in anxiety for female and male victims, respectively, in contrast to non-victims. Effects of the type of measure The results of the meta-analysis showed a significant effect, positive, of a medium size, and generalizable in the diagnosis of depressive and anxiety disorders in victims of CSA/ASA (see Table 5). Likewise, the results of the meta-analysis displayed a significant effect, positive, of a small effect size and generalizable in anxiety and depression symptoms in CSA/ASA victims (see Table 5). Comparatively, injury was significantly higher, qs = 0.152, p < .01, in the clinical diagnosis of anxiety disorder than reported symptoms of anxiety. Likewise, the diagnosis of a depressive disorder was significantly more sensitive, qs = 0.108, p < .05, for CSA/ASA victims than the report of depressive symptoms. These results, due to the type of measure, explained the differences attributed to sample type (Rind, Bauserman, & Tromovitch, 1998): university population (measure of symptoms), and clinical population (clinical diagnosis). The meta-analysis of major depressive disorder and dysthymia (persistent depressive disorder) nesting in the diagnosis of depression (see Table 5), confirmed a significant and positive effect, of a medium size, and generalizable. Thus, the prevalence of a dysthymic disorder in victims of abuse was 6.59 (odds ratio) times higher than for non-victims, 3.25 higher for major depression. In terms of injury quantification, it was of 46% for dysthymia, and 31% for major depression. The difference between effect sizes was significant, qs = 0.176, p < .01, thus the effect size for the diagnosis of dysthymia was significantly larger than for major depressive disorder. Artículos de investigación 131 Similarly, the meta-analysis on generalized anxiety disorder, specific phobia, social phobia and panic disorder were nested in anxiety disorders (see Table 5) found a significant and positive effect, of a medium to large, and generalizable for every diagnosis. Indeed, this implied CSA/ASA victims had 5.12, 7.62, 4.85, and 5.60 (odds ratio) greater probability of developing generalized anxiety disorder, specific phobia, social phobia, and panic disorder, respectively, than CSA/ASA non-victims. Injury was quantified as 41, 49, 40, and 43% for generalized anxiety disorder, specific phobia, social phobia and panic disorder, respectively. The probability of developing these disorders as sequelae was similar except for specific phobia that was significantly higher than social phobia, qs = 0.112, p < .05, and generalized anxiety disorder, qs = 0.100, p < .05. Effects of the type of abuse The results of the meta-analysis on the type of abuse suffered (no contact, contact, and penetration) revealed a significant and positive effect, of small size, and generalizable in depression and anxiety (see Table 6). The comparison of sizes, showed injury derived from abuse with penetration, both in depression and anxiety, was significantly higher than injury in the no contact abuse condition for depression, qs = 0.093, p < .05; and anxiety, qs = 0.092, p < .05. Effects of the interaction type of measure and gender In the diagnosis of depressive disorders (see Table 7), the effect sizes were positive and significant for both males and females, of a small size for males and a medium one for females, which were generalizable for females, but not for males (the effects of the moderators could not be assessed in this case due to the very small k). In comparison, the effect size for females was significantly higher, qs = 0.388, p < .01, than for males, quantifying injury in 42% for female victims and 10% for males. As for prevalence, female CSA/ASA victims had a 5.40 (odds ratio) more probability of meeting the criteria of depressive disorders than female non-victims, whereas males had a 1.44 more probability than male non-victims. Moreover, for the diagnosis of anxiety disorders the effect sizes were positive, significant, and of a small size for both males and females, generalizable for females, but not so for males (once again, moderators could not be found due to the very small k). Once again, the effect size found in females was significantly higher, qs = 0.104, p < .05, than in males, with injury of 24% for female victims and 14% for males. This reveals that female victims had 2.43 (odds ratio) more probability of developing anxiety disorders than female non-victims, and male victims 1.66 more probability than male non-victims. The effect sizes in depressive symptoms were significant, positive, of small sizes, generalizable, and similar (ns) for both males and females. As for anxiety symptoms, the effect sizes were significant, positive, of small sizes, and generalizable for both males and females. Nonetheless, the effect size was significantly higher, qs = 0.095, p < .05, in males. Thus, the results highlight that male CSA/ASA victims had a 2.76 probability (odds ratio) of developing significantly more anxiety symptoms than male non-victims, and female victims a 1.95 probability than female non-victims. BÁRBARA GONZÁLEZ AMADO 132 Table 3. Results of the Meta-Analyses of Sexual Abuse Victimization in General Sequelae, Depression and Anxiety k NE NC NT rw SDr ρ SDρ %VE 95% CIr 90% CIρ General Sequelae 91 19360 93988 125555 .28 .17 .34 .20 2.82 [.27, .29] [.08, .59] Depression 87 18910 92618 123735 .24 .14 .28 .16 4.10 [.23, .25] [.08, .49] Anxiety 62 14587 77494 93075 .26 .14 .31 .17 3.73 [.25, .27] [.10, .52] Note. k = number of studies; NE = experimental group sample size; NC = control group sample size; NT = total sample size; rw = observed correlation (observed validity) weighted for sample size; SDr = standard deviation of the observed correlation; ρ = true correlation (operational validity corrected for criterion and predictor unreliability); SDρ = standard deviation of true correlation; %VE = percentage of variance accounted for by artifactual errors; 95% CIr = 95% confidence interval; 90% CIρ = 90% credibility interval. When NT ≠ NE + NC, it means that experimental or control group sample size in primary studies were unknown. Table 4. Results of the Meta-Analyses of Sexual Abuse Victimization in Depression and Anxiety by Gender k NE NC NT rw SDr ρ SDρ %VE 95% CIr 90% CIρ Depression Measure Females 42 8074 20127 39498 .18 .09 .22 .09 14.56 [.17, .19] [.10, .34] Males 12 1830 13843 15673 .11 .08 .13 .10 10.60 [.09, .13] [.01, .26] Anxiety Measure Females 27 4926 12542 17706 .18 .12 .22 .13 10.82 [.17, .19] [.05, .39] Males 8 998 7380 8378 .12 .13 .15 .15 5.56 [.09, .14] [-.04, .35] Note. k = number of studies; NE = experimental group sample size; NC = control group sample size; NT = total sample size; rw = observed correlation (observed validity) weighted for sample size; SDr = standard deviation of the observed correlation; ρ = true correlation (operational validity corrected for criterion and predictor unreliability); SDρ = standard deviation of true correlation; %VE = percentage of variance accounted for by artifactual errors; 95% CIr = 95% confidence interval; 90% CIρ = 90% credibility interval. When NT ≠ NE + NC, it means that experimental or control group sample size in primary studies were unknown. Artículos de investigación 133 Table 5. Results of the Meta-Analyses of Sexual Abuse Victimization in Depression and Anxiety by Type of Measure. k NE NC NT rw SDr ρ SDρ %VE 95% CIr 90% CIρ Depressive Disorder 28 12131 66986 90220 .26 .14 .31 .03 2.32 [.25, .27] [.10, .52] Dysthymia 8 4524 36668 41192 .38 .08 .46 .09 8.91 [.37, .39] [.34, .56] Major Depressive Disorder 24 9406 64284 84793 .26 .14 .31 .16 2.21 [.25, .27] [.11, .52] Depressive symptomatology 59 6668 25533 33293 .18 .12 .21 .13 12.75 [.17, .19] [.04, .37] Anxiety Disorder 21 10133 58784 68917 .29 .14 .35 .16 2.40 [.28, .30] [.14, .56] Generalized Anxiety Disorder 8 5808 43403 49211 .34 .11 .41 .13 3.63 [.33, .35] [.25, .57] Specific Phobia 3 3830 30616 34446 .41 .03 .49 .02 70.97 [.40, .42] [.46, .51] Social Phobia 10 4901 39701 44602 .34 .13 .40 .15 2.82 [.33, .35] [.21, .60] Panic Disorder 8 4932 37321 42253 .36 .11 .43 .12 4.30 [.35, .37] [.27, .58] Anxiety symptomatology 41 4510 18845 24270 .18 .12 .21 .13 12.53 [.17, .19] [.05, .37] Note. k = number of studies; NE = experimental group sample size; NC = control group sample size; NT = total sample size; rw = observed correlation (observed validity) weighted for sample size; SDr = standard deviation of the observed correlation; ρ = true correlation (operational validity corrected for criterion and predictor unreliability); SDρ = standard deviation of true correlation; %VE = percentage of variance accounted for by artifactual errors; 95% CIr = 95% confidence interval; 90% CIρ = 90% credibility interval. When NT ≠ NE + NC, it means that experimental or control group sample size in primary studies were unknown. Effect size for agoraphobia has not been obtained due to insufficient k. Table 6. Results of the Meta-analyses of Sexual Abuse Victimization in Depression and Anxiety by Type of Abuse. k NE NC NT rw SDr ρ SDρ %VE 95% CIr 90% CIρ Depression Measure Non-Contact 7 278 5431 5709 .12 .03 .14 0 100 [.09, .15] [.14] Contact 4 171 3228 3399 .16 .07 .18 .07 25.97 [.13, .19] [.10, .27] Intercourse 4 184 3228 3412 .19 .08 .23 .09 17.60 [.16, .22] [.11, .34] Anxiety Measure Non-Contact 4 101 3225 3326 .08 .03 .09 0 100 [.05, .11] [.09] Contact 4 170 3225 3395 .11 .06 .14 .06 34.73 [.08, .14] [.06, .21] Intercourse 4 184 3225 3409 .15 .07 .18 .08 22.02 [.12, .18] [.08, .28] Note. k = number of studies; NE = experimental group sample size; NC = control group sample size; NT = total sample size; rw = observed correlation (observed validity) weighted for sample size; SDr = standard deviation of the observed correlation; ρ = true correlation (operational validity corrected for criterion and predictor unreliability); SDρ = standard deviation of true correlation; %VE = percentage of variance accounted for by artifactual errors; 95% CIr = 95% confidence interval; 90% CIρ = 90% credibility interval. When NT ≠ NE + NC, it means that experimental or control group sample size in primary studies were unknown. BÁRBARA GONZÁLEZ AMADO 134 Table 7. Results of the Meta-analyses of the Sexual Abuse Victimization in Depression and Anxiety by Type of Measure and Gender. k NE NC NT rw SDr ρ SDρ %VE 95% CIr 90% CIρ Depressive Disorder Diagnosis Females 11 4421 11594 27130 .33 .19 .42 .23 1.79 [.32, .34] [.13, .71] Males 5 1446 10220 11666 .08 .08 .10 .09 6.90 [.06, .10] [-.02, .21] Anxiety Disorder Diagnosis Females 8 3192 8018 11210 .20 .13 .24 .15 4.75 [.18, .22] [.05, .43] Males 4 845 6806 7651 .11 .13 .14 .15 3.26 [.09, .13] [-.06, .33] Depressive Symptomatology Females 32 3709 8530 12482 .17 .09 .20 .09 27.79 [.15, .19] [.07, .32] Males 7 458 3577 4035 .18 .05 .22 .04 62.72 [.15, .20] [.17, .26] Anxiety Symptomatology Females 20 1790 4580 6608 .15 .09 .18 .09 33.68 [.13, .17] [.07, .30] Males 4 153 594 727 .23 .09 .27 .05 69.51 [.16, .30] [.20, .35] Note. k = number of studies; NE = experimental group sample size; NC = control group sample size; NT = total sample size; rw = observed correlation (observed validity) weighted for sample size; SDr = standard deviation of the observed correlation; ρ = true correlation (operational validity corrected for criterion and predictor unreliability); SDρ = standard deviation of true correlation; %VE = percentage of variance accounted for by artifactual errors; 95% CIr = 95% confidence interval; 90% CIρ = 90% credibility interval. When NT ≠ NE + NC, it means that experimental or control group sample size in primary studies were unknown. Artículos de investigación 135 Discussion As a whole, the results of this study support undoubtedly (a total of 93 studies with nonsignificant results would be required to annul the effect) that CSA/ASA victimization had a significant and positive effect (injury) on mental health, of a smallto large-size, and generalizable. This was demonstrated in the following: a) A higher probability, around 70% in each of the different measures, of suffering from internalized injury, depression, and anxiety. b) Injury caused to the victim’s mental health, that is, mental injury and/or emotional suffering (United Nations, 1988), was calculated to be around 30% (34, 28 and 31% in general sequelae, depression, and anxiety, respectively). This finding implies that offenders are not only criminally responsible for their deeds, but are also liable to civil compensation payments for injuries caused to victims. With this aim in mind a forensic technique has been developed for quantifying injury in specific cases (Arce & Fariña, 2009). c) Injury to mental health in terms of depression and anxiety associated to victims of CSA/ASA was significant in males and females, but with 9% more depression in female, leading to a higher probability of developing a depressive or anxiety disorders in females with injury of 42 and 24%, respectively. In contrast, injury involved significantly more anxiety symptoms in males with 27% injury. However, symptoms are not an optimum indicator of injury. d) Clinical diagnosis was a measure of injury significantly more adequate than symptoms. The evaluation techniques characteristic to clinical diagnosis and clinical symptoms may explain these differences. In the interview, indeed, injury was linked to cause; whereas in the psychometric measures of CSA/ASA victimization it was not, allowing for other causes. Furthermore, the diagnostic threshold was much stricter than for symptoms, which underscores its greater sensitivity and specificity. Thus, the benchmark for future research should be the diagnosed measure of injury based on an interview task, rather than symptoms based on a psychometric measure. e) Injury was calculated to be 46 and 31% for persistent depressive disorder (dysthymia) and major depressive disorder, respectively. Moreover, the expression of injury as dysthymia was significantly greater than for major depressive disorder, that is, the probability of chronic injury (dysthymic) was greater than more serious injury (major depressive disorder). f) Injury to CSA/ASA victims was expressed in anxiety disorders, estimated to be around 40 to 49%, being the highest in specific phobias. g) Abuse with penetration led to injury in depression and anxiety significantly greater than abuse with no contact. These results lend support to the distinction in the legal classification of both criminal typologies. Limitations of the study This meta-analysis entails certain limitations that should be borne in mind when interpreting the data. First, the ground truth of the primary studies for the classification of abuse generally rests on self-reports of a retrospective nature that relies on individual memory capabilities, and are related to false positives or false alarms (Amado, Arce, & Fariña. 2015; BÁRBARA GONZÁLEZ AMADO 136 Schoedl et al., 2010). Moreover, victim self-reports of sexual abuse may bias the results towards concealing them (false negatives), in particular for males (Stoltenborgh et al., 2011). Second, primary studies assume that injury to mental health is sequelae to abuse, without appraising other possible causes (cause-effect relationship) (Jumper, 1995; Vilariño et al., 2013). Third, the effect of the variable under analysis in primary studies was not completely isolated as in many studies victims of sexual abuse, physical abuse, neglect, and other categories appear under the same umbrella. Fourth, as some studies had no control group, the normative population was taken as the contrast group; or it was not equivalent to the experimental one with the subsequent potential for distortion in the calculated effect sizes (Briere, 1992). Alternatively, the results of the meta-analysis were subject to little variability, that is, Ns > 400 and a large k (Hunter & Schmidt, 2004), were highly generalizable (entirely for the female population, and for males with the exception of the diagnosis of a disorder and the general measure of anxiety for the male population), whereas 93 studies with no significant results would be required to annul the evidence supporting the claim that CSA/ASA leads to mental health injuries. Further research is required to determine which moderators inhibit the generalization of the effects in the general measure of anxiety in the male population and in the diagnosis of depression and anxiety. References [References marked with an asterisk indicate studies included in the meta-analysis] Amado, B. G., Arce, R., & Fariña (2015). Undeutsch hypothesis and Criteria Based Content Analysis: A meta-analytic review. The European Journal of Psychology Applied to Legal Context, 7, 3-12. doi: 10.1016/j.ejpal.2014.11.002 American Psychiatric Association. (2013). Diagnostic and statistical manual of mental disorders (5th ed.). Washington, DC: American Psychiatric Association. Arce, R., & Fariña, F. (2009). Evaluación psicológica forense de la credibilidad y daño psíquico en casos de violencia de género mediante el Sistema de Evaluación Global. In F. Fariña, R. Arce, & G. Buela-Casal (Eds.), Violencia de género. Tratado psicológico y legal (pp. 147-168). Madrid: Biblioteca Nueva. Arce, R., Velasco, J., Novo, M., & Farina, F. (2014). Elaboración y validación de una escala para la evaluación del acoso escolar [Development and validation of a scale to assess bullying]. Revista Iberoamericana de Psicología y Salud, 5, 71-104. *Balsam, K. F., Lehavot, K., Beadnell, B., & Circo, E. (2010). Childhood abuse and mental health indicators among ethnically diverse lesbian, gay, and bisexual adults. Journal of Consulting and Clinical Psychology, 74, 459-468. doi: 10.1037/a0018661 *Bonomi, A. E., Cannon, E. A., Anderson, M. L., Rivara, F. P., & Thompson, R. S. (2008). Association between self-reported health and physical and/or sexual abuse experienced before age 18. Child Abuse & Neglect, 32, 693-701. doi: 10.1016/j.chiabu.2007.10.004 Briere, J. (1992). Methodological issues in the study of sexual abuse effects. Journal of Consulting and Clinical Psychology, 60, 196-203. doi: 10.1037/0022-006X.60.2.196 Artículos de investigación 137 *Briere, J., & Elliott, D. M. (2003). Prevalence and psychological sequelae of self-reported childhood physical and sexual abuse in a general population simple of men and women. Child Abuse & Neglect, 27, 1025-1222. doi: 10.1016/j.chiabu.2003.09.008 *Brown, J., Cohen, P., Johnson, F., & Smailes, E. (1999). Childhood abuse and neglect: Specificity of effects on adolescent and young adult depression and suicidality. Journal of the American of child & Adolescent Psychiatry, 38, 1490-1496. doi: 10.1097/00004583-199912000-00009 Bulik, C. M., Prescott, C. A., & Kendler, K. S. (2001). Features of childhood sexual abuse and the development of psychiatric and substance use disorders. British Journal of Psychiatry, 179, 444-449. doi: 10.1192/bjp.179.5.444 *Cantón-Cortés, D., Cortés, M. R., & Cantón, J. (2012). The role of traumagenic dynamics on the psychological adjustment of survivors of child sexual abuse. European Journal of Developmental Psychology, 9, 665-680. doi: 10.1080/17405629.2012.660789 *Cantón-Cortés, D., & Justicia-Justicia, F. (2008). Afrontamiento del abuso sexual infantil y ajuste psicológico a largo plazo [Child sexual abuse coping and long term psychological adjustment]. Psicothema, 20, 209-515. *Carey, P. D., Walker, J. L., Rossouw, W., Seedat, S., Stein, D. J. (2008). Risk indicators and psychopathology in traumatised children and adolescents with a history of sexual abuse. European Child & Adolescent Psychiatry, 17, 93-98. doi: 10.1007/s00787-007-0641-0 *Cheasty, M., Clare, A. W., & Collins, C. (1998). Relation between sexual abuse in childhood and adult depression: Case-control study. British Medical Journal, 316, 198–201. doi: 10.1136/bmj.316.7126.198 *Chen, J., Cai, Y. Y., Cong, E., Liu, Y., Gao, J., Li, Y.,…Flint, J. (2014). Childhood sexual abuse and the development of recurrent major depression in Chinese women. Plos One, 9(1), e87569. doi: 10.1371/journal.pone.0087569 *Chen, J., Dunne, M. P., & Han, P. (2004). Child sexual abuse in China: A study of adolescents in four provinces. Child Abuse & Neglect, 28, 1171-1186. doi:10.1016/j.chiabu.2004.07.003 *Chen, J., Dunne, M. P., & Han, P. (2006). Child sexual abuse in Henan province, China: Associations with sadness, suicidality, and risk behaviors among adolescent girls. Journal of Adolescent Health, 38, 544-549. doi: 10.1016/j.jadohealth.2005.04.001 Cohen, J. (1988). Statistical power analysis for behavioral sciences (2nd ed.). New York, NY: Academic Press. *Comijs, H. C., van Exel, E., van der Mast, R. C., Paauw, A., Voshaar, R. O., & Stek, M. L. (2013). Childhood abuse in late-life depression. Journal of Affective Dissorders, 147(13), 241-246. doi: 10.1016/j.jad.2012.11.010 *Cortés-Arboleda, M. R., Cantón-Cortés, D., & Cantón-Duarte, J. (2011). Consecuencias a largo plazo del abuso sexual infantil: Papel de la naturaleza y continuidad del abuso y del ambiente familiar [Long term consequences of child sexual abuse: The role of the nature and continuity of abuse and family enviroment]. Behavioral PsychologyPsicología Conductual, 19, 41-56. *Cortés-Arboleda, M. R., Cantón-Duarte, J., & Cantón-Cortés, D. (2011). Naturaleza de los abusos sexuales a menores y consecuencias en la salud mental de las víctimas [Characteristics of sexul abuse of minors and its consequences on victims’ mental health]. Gaceta Sanitaria, 25, 157-165. doi: 10.1016/j.gaceta.2010.10.009 *Cutajar, M. C., Mullen, P. E., Ogloff, J. R. P., Thomas, S. D., Wells, D. L., & Spataro, J. (2010). Psychopathology in a large cohort of sexually abused children followed up to 43 years. Child Abuse & Neglect, 34, 813-822. doi: 10.1016/j.chiabu.2010.04.004 BÁRBARA GONZÁLEZ AMADO 138 *Doerfler, L. A., Toscano Jr., P. F., & Connor, D. F. (2009). Sex and aggression: The relationship between gender and abuse experience in youngsters referred to residential treatment. Journal of Child and Family Studies, 18, 112-122. doi: 10.1007/s10826-0089212-3 *Dube, S. R., Anda, R. F., Whitfield, C. L., Brown, D. W., Felitti, V. J., Dong, M., & Giles, W. H. (2005). Long-term consequences of childhood sexual abuse by gender of victim. American Journal of Preventive Medicine, 28, 430-438. doi: 10.1016/j.amepre.2005.01.015 *Feeney, J., Kamiya, Y., Robertson, I. H., & Kenny, R. A. (2013). Cognitive function is preserved in older adults with a reported history of childhood sexual abuse. Journal of Traumatic Stress, 26, 735-743. doi 10.1002/jts.21861 *Feerick, M. M., & Snow, K. L. (2005). The relationships between childhood sexual abuse, social anxiety, and symptoms of posttraumatic stress disorder in women. Journal of Familiy Violence, 20, 409-419. doi: 10.1007/s10896-005-7802-z *Ferguson, K. S., & Dacey, C. M. (1997). Anxiety, depression, and dissociation in women health care providers reporting a history of childhood psychological abuse. Child Abuse & Neglect, 21, 941–952. doi: 10.1016/S0145-2134(97)00055-0 *Fergusson, D. M., Boden, J. M., & Horwood, L. J. (2008). Exposure to childhood sexual and physical abuse and adjustment in early adulthood. Child Abuse & Neglect, 32, 607–619. doi:10.1016/j.chiabu.2006.12.018 *Fergusson, D. M., McLeod, G. F. H., & Horwood, L. J. (2013). Childhood sexual abuse and adult developmental outcomes: Findings from a 30-year longitudinal study in New Zealand. Child Abuse & Neglect, 37, 664-674. doi 10.1016/j.chiabu.2013.03.013 *Fondacaro, K. M., Holt, J. C., & Powell, T. A. (1999). Psychological impact of childhood sexual abuse on male inmates: The importance of perception. Child Abuse & Neglect, 23, 361-369. doi: 10.1016/S0145-2134(99)00004-6 *Frías, M. T., Brassard, A., & Shaver, P. R. (2014). Childhood sexual abuse and attachment insecurities as predictors of women’s own and perceived-partner extradyadic involvement. Child Abuse & Neglect, 38, 1450-1458. doi: 10.1016/j.chiabu.2014.02.009 *Godbout, N., Briere, J., Sabourin, S., & Lussier, Y. (2014). Child sexual abuse and subsequent relational and personal functioning: The role of parental support. Child Abuse & Neglect, 38, 317-325. doi: 10.1016/j.chiabu.2013.10.001 *Gudjonsson, G. H., Sigurdsson, J. F., & Tryggvadóttir, H. B. (2011). The relationship of compliance with a background of childhood neglect and physical and sexual abuse. The Journal of Forensic Psychiatry & Psychology, 22, 87–98. doi: 10.1080/14789949.2010.524707 *Haj-Yahia, M. M., & Tamish, S. (2001). The rates of child sexual abuse and its psychological consequences as revealed by a study among Palestinian university students. Child Abuse & Neglect, 25, 1303-1327. doi: 10.1016/S0145-2134(01)00277-0 Hedges, L. V., & Olkin, I. (1985). Statistical methods for meta-analysis. Orlando, FL: Academic Press. *Henderson, D., Hargreaves, I., Gregory, S., & Williams, J. M. G. (2002). Autobiographical memory and emotion in a non-clinical sample of women with and without a reported history of childhood sexual abuse. British Journal of Clinical Psychology, 41, 129-141. doi: 10.1348/014466502163921 *Hobfoll, S. E., Bansal, A, Schurg, R., Young, S., Pierce, C. A., Hobfoll, I., & Johnson, R. (2002). The impact of perceived child physical and sexual abuse history on native Artículos de investigación 139 American women’s psychological well-being and AIDS risk. Journal of Consulting and Clinical Psychology, 70, 252-257. doi: 10.1037//0022-006X.70.1.252 Hunter, J. E., & Schmidt, F. L. (2004). Methods of Meta-analysis: Correcting error and bias in research findings (2nd ed.). Thousand Oaks, CA: Sage. Intebi, I. V. (1998). Abuso sexual infantil: En las mejores familias. Buenos Aires, Argentina: Granica. *Jonas, S., Bebbington, P., McManus, S., Meltzer, H., Jenkins, R., Kuipers, E., Cooper, C., King, M., & Brugha, T. (2011). Sexual abuse and psychiatric disorder in England: Results from the 2007 adult psychiatric morbidity survey. Psychological Medicine, 41, 709-719. doi: 10.1017/S003329171000111X Jumper, S. A. (1995). A meta-analysis of the relationship of child sexual abuse to adult psychological adjustment. Child Abuse & Neglect, 19, 715-728. doi: 10.1016/01452134(95)00029-8 *Kendler, K. S., Bulik, C. M., Silberg, J., Hettema, J. M., Myers, J., & Prescott, C. A. (2000). Childhood sexual abuse and adult psychiatric and substance use disorders in women. An epidemiological and cotwin control analysis. Archives of General Psychiatry, 57, 953959. doi: 10.1001/archpsyc.57.10.953 *Kent, A., & Waller, G. (1998). The impact of childhood emotional abuse: An extension of the child abuse and trauma scale. Child Abuse & Neglect, 22, 393–399. doi: 10.1016/S0145-2134(98)00007-6 Koenen, K. C., & Widom, C. S. (2009). A prospective study of sex differences in the lifetime risk of posttraumatic stress disorder among abused and neglect children grown up. Journal of Traumatic Stress, 22, 566-574. doi: 10.1002/jts.20478 *Kugler, B. B., Bloom, M., Kaercher, L. B., Truax, T. V., & Storch, E. A. (2012). Somatic symptoms in traumatized children and adolescents. Child Psychiatry & Human Development, 43, 661–673. doi:10.1007/s10578-012-0289-y *Kuo, J. R., Goldin, P. R., Werner, K., Heimberg, R. G., & Gross, J. J. (2011). Childhood trauma and current psychological functioning in adults with social anxiety disorder. Journal of Anxiety Disorders, 25, 467–473. doi:10.1016/j.janxdis.2010.11.011 *Lamoureux, B. E., Palmierir, P. A., Jackson, A. P., & Hobfoll, S. E. (2012). Child sexual abuse and adulthood interpersonal outcomes: Examining pathways for intervention. Psychological Trauma: Theory, Research, Practice, and Policy, 4, 605-6013. doi: 10.1037/a0026079 *Leck, P., Difede, J., Patt, I., Giosan, C., & Szkodny, L. (2006). Incidence of male childhood sexual abuse and psychological sequelae in disaster workers exposed to a terrorist attack. International Journal of Emergency Mental Health, 8, 267-274. *Li, N., Ahmed, S., & Zabin, L. S. (2012). Association between childhood sexual abuse and adverse psychological outcomes among youth in Taipei. Journal of Adolescent Health, 50, S45-S51. doi: 10.1016/j.jadohealth.2011.12.003 *Liem, J. H., O’Toole, J. G., & James, J. B. (1996). Themes of power and betrayal in sexual abuse survivors’ characterizations of interpersonal relationship. Journal of Traumatic Stress, 9, 745-761. *López, F., Carpintero, E., Hernández, A., Martín, M. J., & Fuertes, A. (1995). Prevalencia y consecuencias del abuso sexual al menor en España [Prevalence and consequences of sexual abuse in children in Spain]. Child Abuse & Neglect, 19, 1039-1050. doi: 10.1016/0145-2134(95)00066-H BÁRBARA GONZÁLEZ AMADO 146 NGE NGC r rxx ryy CSA Questionnaire Depression/Anxiety Measure Type of measure Feeney, Kamiya, Robertson, & Kenny (2013) 451 6256 .08 - .85 2 questions CES-D Depressive symptomatology 451 6256 .08 - .80 < 18 years HADS-A Anxiety symptomatology Feerick & Snow (2005)a 98 215 .13 - .84 CSAI < 18 years HSCL Anxiety symptomatology Fergusson, Boden, & Horwood (2008)(1) 28 881 .09 - - Interview CIDI Anxiety Disorder 28 881 .12 - - <16 years DSM-IV Major Depressive Disorder Fergusson, Boden, & Horwood (2008)(2) 52 881 .17 - - Interview CIDI Anxiety Disorder 52 881 .23 - - < 16 years DSM-IV Major Depressive Disorder Fergusson, Boden, & Horwood (2008)(3) 64 881 .19 - - Interview CIDI Anxiety Disorder 64 881 .27 - - <16 years DSM-IV Major Depressive Disorder Ferguson & Dacey (1997)a 19 55 .41 - .85 CEQ BDI Depressive symptomatology 19 55 .37 - .90 STAI Anxiety symptomatology Fergusson, McLeod, & Horwood (2013) 28 809 .08(1) - - Structured CIDI Major Depressive Disorder 51 809 .15(2) - - Interview Major Depressive Disorder 62 809 .20(3) - - < 16 years Major Depressive Disorder 28 809 .05(1) - - Structured CIDI Anxiety Disorder 51 809 .09(2) - - Interview Anxiety Disorder 62 809 .22(3) - - < 16 years Anxiety Disorder Fondacaro, Holt, & Powell (1999)b 86 125 .18 - - Questionnaire DIS (DSM-III-R) Major Depressive Disorder 86 125 .06 - - < 16 years Dysthymia 86 125 .23 - - Panic Disorder 86 125 .25 - - Generalized Anxiety Disorder Frias, Brassard, & Shaver (2014)a 116 691 .10 - .90 1 question ECR Anxiety symptomatology Godbout, Briere, Sabourin, & Luissier (2013) 59 284 .23 - .88 SCEQ ECR Anxiety symptomatology <18 years Gudjonsson, Sigurdsson, & 37 73 .12 - .84 Parental Neglect DASS Anxiety symptomatology Tryggvadóttir (2011) 37 73 .06 - .91 and Sexual Abuse DASS Depressive symptomatology Questionnaire < 18 years Haj‒Yahia & Tamish (2001) 652 .55 .89 .88 Sexual Abuse BSI Anxiety symptomatology 652 .60 .89 .88 Finkelhor’s BSI Depressive symptomatology 652 .59 .89 - Scale Anxiety symptomatology Henderson, Hargreaves, Gregory, 22 57 .27 - - Interview POMS-SF Depressive symptomatology & Williams (2002)a 22 57 .31 - - <14 years POMS-SF Anxiety symptomatology Hobfoll et al. (2002) 67 .05 .93 .90 CTQ POMS Depressive symptomatology < 17 years NGE NGC r rxx ryy CSA Questionnaire Depression/Anxiety Measure Type of measure Artículos de investigación 147 Jonas et al. (2011) 964 6389 .21 - .75 Interview CIS-R Major Depressive Disorder .19 - .75 <16 years CIS-R Generalized Anxiety Disorder .18 - .75 CIS-R Panic Disorder .28 - .75 CIS-R Phobic Disorder Kendler et al. (2000) 427 983 .24 - - Interview SCI (DSM-III-R) Generalized Anxiety Disorder 427 983 .24 - - < 16 years SCI Panic disorder 427 983 .25 - - SCI Major Depressive Disorder Kent & Waller (1998)a 236 .30 .61 .70 CATS HADS-A Anxiety symptomatology 236 .28 .61 .60 HADS-D Depressive symptomatology Kugler, Bloom, Kaercher, Truax, 54 107 .57 - .88 Forensic sample TSCC Anxiety symptomatology & Storch 54 107 .62 - .81 8-17 years CDI and TSCC Depressive symptomatology Kuo, Goldin, Werner, Heimberg, & 20 82 .10 .86 .93 CTQ-SF SIAS Anxiety symptomatology Gross (2011) 20 82 .07 .86 .90 < 16 years BDI-II Depressive symptomatology Lamoureux, Palmieri, Jackson, & 271 422 .26 .87 .89 CTQ CES-D Depressive symptomatology Hobfoll (2012)a <16 years Leck, Difede, Patt, Giosan, & Szkodny (2006)b 92 2030 .18 .83 .93 TEI BDI-II Depressive symptomatology Li, Ahmed, & Zabin (2012) 214 3870 .27 - - Research Study of 1 question Anxiety symptomatology 214 3870 .28 - - Adolescent Health 1 question Depressive symptomatology <14 years Liem, O’Toole, & James (1996)a 43 43 .24 - .81 SEQ BSI Anxiety symptomatology 43 43 .25 - .85 <14 years BSI Depressive symptomatology 43 43 .29 - .83 BDI-SF Depressive symptomatology Linskey & Fergusson (1997) 24 918 .11(1) - - Reports of DSM-IV Anxiety Disorder 47 918 .15(2) - - Childhood Sexual DSM-IV Anxiety Disorder 36 918 .15(3) - - abuse DSM-IV Anxiety Disorder 24 918 .12(1) - - <16 years DSM-IV Depressive Disorder 47 918 .17(2) - - DSM-IV Depressive Disorder 36 918 .21(3) - - DSM-IV Depressive Disorder López, Carpintero, Hernández, Martín, 337 1484 .12 - - Interview SRQ Anxiety symptomatology & Fuertes (1995) 337 1484 .08 - - < 16 years SRQ Depressive symptomatology Lumley & Harkness (2007) 11 .13 .93 .91 CECA MASQ Anxiety symptomatology 11 .13 .93 .88 MASQ Depressive symptomatology Luterek, Harb, Heimberg, & Marx (2004)a 34 .12 - .81 LEQ < 14 years BDI Depressive symptomatology NGE NGC r rxx ryy CSA Questionnaire Depression/Anxiety Measure Type of measure MacMillan et al. (2001)a 508 3170 .08 - - Child Maltreatment CIDI Anxiety Disorder 508 3170 .07 - - History Self-Report CIDI Depressive Disorder MacMillan et al. (2001)b 150 3188 .02 - - Child Maltreatment CIDI Anxiety Disorder BÁRBARA GONZÁLEZ AMADO 148 150 3188 .02 - - History Self-Report CIDI Depressive Disorder Manion et al. (1998)a 29 45 .47 - - NAEF Depression Self-Rating Depressive symptomatology <14years Scale for children Manion et al. (1988)b 22 29 .39 - - NAEF Depression Self-Rating Depressive symptomatology <14years Scale for children Mapp (2006)a 107 158 .13 - .87 Forensic Sample CES-D Depressive symptomatology <18 years Mchichi Alami & Kadri (2004)a 62 620 .09 - - Questionnaire Hamilton depression Depressive symptomatology rating scale 19 620 .10(1) - - Depressive symptomatology 21 620 .03(2) - - Depressive symptomatology 22 620 .03(3) - - Depressive symptomatology 63 617 .02 - - Hamilton anxiety Anxiety symptomatology rating scale 21 617 .05(1) - - Anxiety symptomatology 20 617 .01(2) - - Anxiety symptomatology 22 617 .01(3) - - Anxiety symptomatology McLean, Morris, Conklin, Jayawickreme, & 71 .59 - .87 Forensic Sample BDI Depressive symptomatology Foa (2014)a 13-18 years McLeer et al. (1998) 80 73 .33 - .86 Interview CDI Depressive symptomatology 80 73 .27 - .89 6-16 STAIC Anxiety symptomatology Merril (2001)a 248 523 .15 - .84 SEQ TSI Anxiety symptomatology 248 523 .17 - .84 <14 years TSI Depressive symptomatology Messman‒Moore, Long, & Siegfried (2000)a 56 282 .24 .89 .85 LEQ SCL-90-R Anxiety symptomatology 56 282 .17 .89 .90 < 17 years SCL-90-R Depressive symptomatology Meyerson, Long, Miranda, & Marx (2002) 39 91 .23 .75 .93 SEQ BDI-II Depressive symptomatology <12 years Meston, Rellini, & Heiman (2006)a 48 71 .34 - .92 Questionnaire BAI Anxiety symptomatology 48 71 .40 - .86 < 16 years BDI Depressive symptomatology Miller (2006)a 25 50 .22 - .86 Interview BDI Depressive symptomatology NGE NGC r rxx ryy CSA Questionnaire Depression/Anxiety Measure Type of measure Molnar, Buka, & Kessler (2001)a 394 2527 .13 - - <18 years DSM-III-R Generalized Anxiety Disorder 394 2527 .13 - - DSM-III-R Panic Disorder 394 2527 .13 - - DSM-III-R Phobic Disorder 394 2527 .23 - - DSM-III-R Major Depressive Disorder 394 2527 .25 - - DSM-III-R Dysthymia Artículos de investigación 149 Molnar, Buka, & Kessler (2001)b 74 2871 .04 - - <18 years DSM-III-R Generalized Anxiety Disorder 74 2871 .09 - - DSM-III-R Panic Disorder 74 2871 .18 - - DSM-III-R Phobic Disorder 74 2871 .23 - - DSM-III-R Major Depressive Disorder 74 2871 .16 - - DSM-III-R Dysthymia Mullen, Martin, Anderson, Romans, & 53 390 .21 - - Questionnaire PSE-SF Depressive symptomatology Herbison (1996)a <16 years Musliner & Singer (2014)a 436 221 .24 - .85 Questionnaire CES-D-10 Depressive symptomatology <16 years Nelson et al. (2002)a 387 1931 .12 - - Interview DSM-IV Social Phobia 387 1931 .17 - - <18 years DSM-IV Major Depressive Disorder Nelson et al. (2002)b 90 1574 .04 - - Interview DSM-IV Social Phobia 90 1574 .08 - - <18 years DSM-IV Major Depressive Disorder Newcomb, Munoz, & Carmona (2009)a 66 79 .26 - .77 CMIS-SF TSI Anxiety symptomatology 66 79 .20 - .86 <17 years TSI Depressive symptomatology Newcomb, Munoz, & Carmona (2009)b 19 59 .32 - .77 CMIS-SF TSI Anxiety symptomatology 19 59 .35 - .86 <17 years TSI Depressive symptomatology Offen, Waller, & Thomas (2003) 10 16 .42 - - 1 Question BDI Depressive Symptomatology Peleikis, Mykletun, & Dahl (2004)a 56 56 .52 - - Interview SCID-II Major Depressive Disorder 56 56 .27 - - <16 years Dysthymia BÁRBARA GONZÁLEZ AMADO 150 NGE NGC r rxx ryy CSA Questionnaire Depression/Anxiety Measure Type of measure Peleikis, Mykletun, & Dahl (2005)a 56 56 0 - - Detailed Structured SCID Panic Disorder 56 56 .10 - - Interview SCID Agoraphobia 56 56 .07 - - <16 years SCID Social Phobia 56 56 .24 - - SCID Generalized Anxiety Disorder 56 56 .12 - .85 SCL-90-R Anxiety symptomatology 56 56 .14 - .82 SCL-90-R Anxiety symptomatology 56 56 .19 - - SCID Major Depressive Disorder 56 56 .23 - - SCID Dysthymia 56 56 .18 - .90 SCL-90-R Depressive symptomatology Pérez‒Fuentes et al. (2013) 3786 30431 .45 - - ACE DSM-IV Panic Disorder 3786 30431 .39 - - <18 years DSM-IV Social Phobia 3786 30431 .34 - - DSM-IV Specific Phobia 3786 30431 .44 - - DSM-IV Generalized Anxiety Disorder 3786 30431 .35 - - DSM-IV Major Depressive Disorder .40 - - DSM-IV Dysthymia Portegijs, Jeuken, van der Horst, Kraan, 11 .07 - - Youth Experiences DIS Anxiety Disorder & Knottnerus (1996)a 11 .33 - - Questionnaire DIS Depressive Disorder <16 years Rich, Gidycz, Warkentin, Loh, & 42 .09 - .84 Child sexual BDI-II Depressive symptomatology Weiland (2005)c victimization Quest. <14 years Rich, Gidycz, Warkentin, Loh, & 189 .28 - .74 SES BDI-II Depressive symptomatology Weiland (2005)d <18 years Schaaf & McCanne (1998) a 27 211 .06 - - CSEQ TSI Anxiety symptomatology 27 211 .11 - - < 15 years TSI Depressive symptomatology Silverman, Reinherz, & Giaconia (1996)a 23 164 .08 - - Interview DIS (DSM-III-R) Specific Phobia 23 164 .03 - - <18 years DIS Social Phobia 23 164 .23 - - DIS Major Depressive Disorder Spertus, Yehuda, Wong, Halligan, 41 162 .20 .82 .85 CTQ SCL-90-R Anxiety symptomatology & Seremetis (2003)a 41 162 .18 .82 .90 < 17 years SCL-90-R Depressive symptomatology Artículos de investigación 151 NGE NGC r rxx ryy CSA Questionnaire Depression/Anxiety Measure Type of measure Steel, Sanna, Hammond, Whipple, & 85 172 .21 - .85 Sexual History SCL-90-R Anxiety symptomatology Cross (2004) 85 172 .15 - .90 Questionnaire SCL-90-R Depressive symptomatology <18 years Subica (2013) 50 122 .25 - .86 TAA-R PHQ-8 Major Depressive Disorder <18 years Sun et al. (2008) 244 781 .29 - .85 Questionnaire SCL-90-R Anxiety symptomatology 244 781 .36 - .90 < 18 years SCL-90-R Depressive symptomatology Swanston et al. (2003) 104 .41 - .82 Forensic Sample RCMAS Anxiety symptomatology 63 .41 - .86 5-15 years CDI Depressive symptomatology Thomas, DiLillo, Walsh, & Polusny (2011)a 52 .36 .85 .93 WSHQ BDI-II Depressive symptomatology <14 years Thompson et al. (2003)a 26 25 .45 - - Interview SCID-I Anxiety Disorder 26 25 .55 - - < 18 years SCID-I Depressive Disorder Trowell et al. (1999)a 21 21 .48 - - Forensic Sample Kiddie-SADS Major Depressive Disorder .27 - - 8-14 years Kiddie-SADS Generalized Anxiety Disorder .29 - - Social Phobia .19 - - Specific Phobia van Vugt, Lanctôt, Paquette, Collin‒Vézina, 89 .46 .87 .88 CTQ TSI-2 Anxiety symptomatology & Lemieux (2013)a 89 .38 .87 .88 < 17 years TSI-2 Depressive symptomatology Villarroel, Penelo, Portell, & Raich (2012)a 81 597 .04 - .93 TLEQ STAI Anxiety symptomatology 81 597 .14 - .88 <18 years BDI Depressive symptomatology Widom, DuMont, & Czaja (2007) 96 520 .03 - - Forensic Sample DIS Major Depressive Disorder <12 years Young, Harford, Kinder, & Savell (2007)a 116 163 .22 .79 .83 ESE BSI Anxiety symptomatology 116 163 .13 .79 .89 <16 years BSI Depressive symptomatology Young, Harford, Kinder, & Savell (2007)b 39 88 .05 .79 .83 ESE BSI Anxiety symptomatology 39 88 .05 .79 .89 <16 years BSI Depressive symptomatology Note. NGE = experimental group sample size; NGC = control group sample size; r = sexual abuse victimization and depression/anxiety correlation; rxx = reliability of sexual abuse measure instruments; ryy = reliability of Anxiety and Depression measure instruments; a = female participants; b = male participants; c = childhood sexual abuse (CSA); d = adolescence sexual abuse (AA). (1) = Non-contact CSA; (2) = contact CSA; (3) = Intercourse CTQ-SF = Childhood Trauma Questionnaire Short Form; TES = Traumatic Events Survey; CESD-10 = 10-item Center for Epidemiologic Studies Depression; PHQ GAD-7 = 7-item Patient Health Questionnaire Generalized Anxiety Disorder Scale; CES-D = 20-item Center for Epidemiological Studies-Depression scale; BDI = Beck Depression Inventory; CIDI = Composite International Diagnostic Interview; IDS = Inventory of Depression Symptoms; STAI = State Trait Anxiety Inventory; ICD = International Classification of Desease; STAIC = State Trait Anxiety Inventory for Children; DSMD = Devereux Scales of Mental Disorders; ACE = Adverse Childhood Experiences; HADS-A = Hospital Anxiety and Depression Scale – Anxiety Scale; CSAI = Childhood Sexual Abuse Interview; HSCL = Hopkings Symptom Checklist; CEQ = Childhood Experiences Questionnaire; ECR = BÁRBARA GONZÁLEZ AMADO 152 Experiences in Close Relationships Questionnaire; DASS = Depression Anxiety Stress Scales; BSI = Brief Symptom Inventory; POMS-SF = Profile of Mood States-Short Form; CTQ = Childhood Trauma Questionnaire; CIS-R = Clinical Interview Schedule Revised; CATS = Child Abuse and Trauma Scale; TSCC = Trauma Symptom Checklist for Children; CDI = Children Depression Inventory; TEI = Traumatic Events Interview; BDI-II = Beck Depression Inventory, second edition; TEQ = The Traumatic Events Questionnaire; BDISF = Beck Depression InventoryShort Form; BSI = Brief Symptom Inventory; SEQ = The Significant Events Questionnaire; SRQ = Self Reporting Questionnaire; CECA = Childhood Experience of Care and Abuse Interview; MASQ = Mood and Anxiety Symptom Questionnaire; AA = Anxiety Arousal; TSI = Trauma Symptom Inventory; LEQ = Life Experiences Questionnaire; BAI = Beck Anxiety Inventory; PSE-SF = Present State ExaminationShort Form; CMIS-SF = Childhood Maltreatment Interview Schedule-Short Form; SCID-II = Structured Clinical Interview for DSM-IV Axis II; CSEQ = Childhood Sexual Experiences Questionnaire; TAA-R = Trauma Assessment for Adults Brief Revised Version; PHQ-8 = Patient Health Questionnaire – 8; RCMAS = Revised Children’s Manifest Anxiety Scale; WSHQ = CSA subscale from the Wyatt Sex History Questionnaire; Kiddie-SADS = Semi-Structured Interview Kiddie-Sads, DSM-IV; SICE = Structured Interview on CSA Experiences; TSC-33 = Trauma Symptom Checklist; ESE = Early Sexual Experiences; NIMH = National Institute of Mental Health Diagnostic Interview Schedule, Version III Revised; SES = Sexual Experiences Survey; SCI = Structured Clinical Interview; SIAS = Social Interaction Anxiety Scale; NAEF = Natural of the Abusive Experience Form; DIS = Diagnostic Interview Schedule; EPDS = Edinburgh Post-Natal Depression Scale; CCEI = Crown-Crisp Experiential Index; TLEQ = Traumatic Life Events Questionnaire; SCEQ = Childhood Sexual Experiences Questionnaire; RADS = Reynold’s Adolescent Depression Scale. 8.4. ANÁLISIS DE CONTENIDO EN DECLARACIONES DE AGRESORES. UNA REVISIÓN METAANALÍTICA Bárbara G. Amado, Manuel Vilariño*, Mercedes Novo Departamento de Psicología Organizacional, Juridica-Forense y Metodología de las Ciencias del Comportamiento, Universidade de Santiago de Compostela (España) *Escuela Universitaria de Enfermería de Pontevedra, Universidade de Vigo (España) Resumen El análisis de contenido de la declaración, es una de las metodologías utilizadas en el contexto judicial para la evaluación de la credibilidad del testimonio. Concretamente, el instrumento de evaluación más ampliamente utilizado es el Criteria-Based Content Analysis (CBCA), que se ha aplicado originariamente a declaraciones de víctimas y en casos de abusos sexuales, casi exclusivamente. El interés por esta técnica, ha llevado a realizar investigaciones a poner a prueba la validez del CBCA en muestras de agresores adultos, llegando a obtener resultados diversos y contradictorios. Con el objetivo de arrojar luz sobre la capacidad del CBCA para discriminar entre declaraciones reales y falsas y validar la Hipótesis Undeutsch en agresores adultos, nos planteamos la elaboración de una revisión meta-análitica de la literatura científica. Los resultados validan la Hipótesis Undeutsch, esto es, las memorias de hechos vividos difieren en contenido y calidad de las memorias de hechos inventados de agresores adultos y, además, el CBCA discrimina entre ambos tipos de memorias. No obstante, no todos los criterios de realidad son aplicables a muestras de agresores, ni son generalizables a otros contextos y muestras. Son necesarios más estudios con rigor científico, elaborados bajo condiciones de high fidelity, para alcanzar decisiones más robustas. Se discuten los resultados para su aplicación a la práctica forense. Palabras clave: análisis de contenido; testimonio; agresores; CBCA. Abstract Content analysis of a statement is a method use to assess credibility statements in judicial context. Concretely, the most utilized assessment tool is Criteria-Based Content Analysis (CBCA), which has been originally applied to victim’s statements and in sexual abuse cases, almost exclusively. The interest for this technique, has led to investigate the validity of CBCA in adult offender samples, reaching diverse and contradictory outcomes. With the objective to shed light about the capacity of CBCA to discriminate between real and false accounts and to validate the Undeutsch Hypothesis, we have conducted a meta-analytic review of the scientific literature. Results have validated Undeutsch Hypothesis, that is, memories of selfexperienced events differ in content and quality from memories of invented events of adult offenders and CBCA has distinguished between both memories, as well. Notwithstanding, nor every reality criteria are applicable to offender’s samples, nor generalizable to other contexts and samples. More research with scientific rigor is needed, made in high fidelity conditions, for reaching more robust decisions. Implications for forensic practice are discussed. Keywords: content analysis; testimony; offenders, CBCA. BÁRBARA GONZÁLEZ AMADO 154 Resumo A análise de contido da declaración é unha das metodoloxías utilizadas no contexto xudicial para a avaliación da credibilidade dunha testemuña. Concretamente, o instrumento de avaliación máis utilizado é o Criteria-Based Content Analysis (CBCA), que se aplicou orixinariamente a declaración de vítimas en casos de abusos sexuais, case exclusivamente. O interese por esa técnica levou a realizar investigación e a poñer a proba a validez do CBCA en mostrar de agresores adultos, chegando a obterse resultados diversos e contraditorios. Co obxectivo de obter luz sobre a capacidade do CBCA para discriminar entre declaracións reais e falsas e validar a Hipótese Undeutsch en agresores adultos, plantexámonos a elaboración dunha revisión meta-analítica da literatura científica. Os resultados validan a Hipótese Undeutsch, é dicir, as memorias de feitos vividos difiren en contido e calidade das memorias de feitos inventados de agresores adultos e, ademais, o CBCA discrimina entre ambos os tipos de memoria. Non obstante, non todos os criterios de realidade son aplicables á mostra de agresores, nin xeneralizables a outros contextos e mostras. Son necesarios máis estudo con rigor científico, elaborados baixo condicións de high fidelity, para alcanzar decisións máis robustas. Discútense os resultados para a súa aplicación á práctica forense. Palabras chave: análise de contido; testemuña; agresores; CBCA. Introducción El análisis de la credibilidad del testimonio se utiliza en el contexto judicial para dotar de valor de prueba al testimonio del denunciante, especialmente al de la víctima. El instrumento más utilizado para ello y con mayor respaldo en la comunidad científica (Amado, Arce, y Fariña, 2015) es el Criteria-Based Content Analysis (Steller y Köhnken, 1989), creado originariamente, para evaluar el testimonio de menores víctimas de abusos sexuales. Aunque el CBCA nace en un contexto y para una población determinados, literatura incipiente ha tratado de generalizar su uso a otras casuísticas (e.g. violencia de género), poblaciones (e.g. adultos) y actores dentro del proceso penal (e.g. testigos, agresores). Si bien en nuestro ordenamiento jurídico se le presuponen al acusado una serie de garantías constitucionales entre las que se encuentra el derecho a mentir, el estudio de la credibilidad del testimonio se presenta como una posibilidad cuando existe una declaración de tamaño suficiente para el análisis de contenido. Para ello, se ha propuesto la aplicación del CBCA, el cual se sustenta en la llamada Hipótesis Undeutsch, esto es, las memorias de hechos experimentados o realmente vividos difieren en contenido y calidad de aquellas que son fruto de la imaginación o fantasía (Undeutsch, 1967). Son escasas las investigaciones que estudian la validez de la hipótesis y, por extensión, del CBCA en muestras de agresores adultos aunque ya se presuponía su aplicabilidad a otras muestras y contextos (Berliner y Conte, 1993) al tratarse de una hipótesis que se sustenta en contenidos de memoria. Por todo lo anterior y como consecuencia de los resultados contradictorios que arroja la literatura científica al respecto, nos planteamos un meta-análisis que someta a contraste la validez de la hipótesis Undeutsch y del CBCA en muestras de agresores de mayoría de edad. Artículos de investigación 155 Método Búsqueda de literatura Se ha llevado a cabo una búsqueda bibliográfica en la base de datos científica de referencia (Web of Science) y en redes sociales científicas para el intercambio de información con otros investigadores (i.e., Researchgate). Asimismo, nos hemos puesto en contacto con especialistas en la materia para identificar comunicaciones que no hubieran sido publicadas en revistas. Además, se revisaron las referencias bibliográficas de artículos seleccionados en busca de trabajos no identificados hasta el momento (ancestry approach). Las palabras clave utilizadas como motor de búsqueda han sido: análisis de contenido/content analysis, credibilidad/credibility, análisis de contenido basado en criterios/criteria-based content analysis, CBCA, testimonio/testimony, declaración/statement, adultos/adults, agresor/offender. Criterios de inclusión y exclusión Para que los trabajos identificados pasaran a formar parte del presente meta-análisis, debían cumplir una serie de criterios de inclusión: 1) que la población de estudio la constituyeran adultos, esto es, con edades iguales o superiores a los 17 años, criterio establecido por los propios autores de las investigaciones primarias para considerar a un sujeto como adulto, 2) que se tratara de una muestra de agresores, 3) que proporcionaran un tamaño del efecto o, en su caso, cualquier otro estadístico que permitiera su cómputo, 4) que analizaran la capacidad discriminativa de los criterios de realidad y del CBCA en su conjunto para diferenciar entre declaraciones reales y falsas. Se han excluido aquellos estudios en los que la unidad de análisis no era la declaración o que utilizaran otros sistemas categoriales metódicos que no cumplieran el criterio de ‘exclusión mutua’. Por otro lado, se han incorporado criterios explícitamente formulados como adicionales al CBCA. Finalmente, 10 estudios que cumplían los criterios de inclusión, formaron parte del metaanálisis. De ellos se extrajeron 12 tamaños del efecto. Procedimiento Una vez identificados los estudios, se procedió a su codificación atendiendo a sus características. Aquellas que pueden tener una influencia en los resultados se tomarán como variables moderadoras, procediéndose al cálculo de un nuevo meta-análisis con base a dichas variables. La jurisprudencia ha establecido que para que una prueba sea considerada evidencia científica admisible en un proceso penal, es necesario que se cumpla el criterio Daubert de publicación en revistas con revisión por pares. Por ello, se tomará este criterio como posible moderador. La codificación fue llevada a cabo por dos investigadores de forma independiente, obteniendo una concordancia total entre ellos. Análisis de datos