When the Tool Agrees with You: Large Language Model Sycophancy, Examiner Impartiality, and the Admissibility of AIAssisted Forensic Opinion
DOI:
https://doi.org/10.65879/3070-5789.2026.02.06Keywords:
Digital forensics, expert evidence admissibility, large language models, sycophancy, cognitive bias.Abstract
Introduction. In Matter of Weber (2024) a surrogate’s court put its own Microsoft Copilot query to three of its computers and got three different figures; the expert whose calculation prompted the exercise could not recall what prompt he had used. A model’s output is a function of the prompt, which in forensic practice carries the examiner’s hypothesis. Work on bias in digital forensics models the examiner deferring to the machine. Sycophancy inverts that: a model conforming to stated beliefs returns the examiner’s hypothesis in the register of an independent instrument, manufacturing the appearance of corroboration. Methods. We surveyed reliability screening for expert evidence in six regimes and tested the hypothesis. Six dated model snapshots answered five forensic interpretation items under three framings differing only in what the examiner asserts, sampled thirty times per cell for 2,700 trials. Replies were classified by a model judge validated against hand coding of 77 replies (Gwet’s AC1 = 0.855). Results. The hypothesis was not supported. Across 2,160 analysed trials, models asserted the incorrect proposition in 36.1% of trials when asked neutrally and 19.6% when an examiner asserted it; thirteen of twenty-four cells moved away from the examiner and ten did not move; the single cell that moved toward it had a baseline error of 93.3%, leaving 6.7 points of headroom. A second finding was sharper: on an item whose conclusion rested on an unsupported premise, one of six models contested it in 80% of trials and the other five in none of 150. Discussion. Comparative screening shows an inversion: jurisdictions with the strongest reliability gates have no rule on expert use of generative AI, while Australia, which refuses a reliability gate, has the only binding instruments. Rule 702(d) is a sharper hook than Daubert’s error-rate factor, since a prompt-dependent error rate is not a property of a method. These models verify conclusions without auditing the reasons given, so disclosure built around the conclusion alone will not surface the failure.
References
[1] Matter of Weber. Matter of Weber, 85 Misc 3d 727, 220 NYS3d 620, 2024 NY Slip Op 24258, 2024. Surrogate’s Court, Saratoga County, NY, Schopf S., 10 October 2024. Expert used Microsoft Copilot; court held AI-generated evidence requires affirmative disclosure and a Frye hearing.
[2] Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R. Bowman, Newton Cheng, Esin Durmus, Zac Hatfield-Dodds, Scott R. Johnston, Shauna Kravec, Timothy Maxwell, Sam McCandlish, Kamal Ndousse, Oliver Rausch, Nicholas Schiefer, Da Yan, Miranda Zhang, and Ethan Perez. Towards Understanding Sycophancy in Language Models. In The Twelfth International Conference on Learning Representations (ICLR), 2024. https://doi.org/10.48550/arXiv.2310.13548.
[3] Ethan Perez, Sam Ringer, Kamilė Lukošiūtė, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, et al. Discovering Language Model Behaviors with Model-Written Evaluations. In Findings of the Association for Computational Linguistics: ACL 2023, pages 13387–13434, 2023. https://doi.org/10.18653/v1/2023.findings-acl.847.
[4] Itiel E. Dror, David Charlton, and Ailsa E. Péron. Contextual information renders experts vulnerable to making erroneous identifications. Forensic Science International, 156 (1): 74–78, 2006. https://doi.org/10.1016/j.forsciint.2005.10.017.
[5] Itiel E. Dror. Cognitive and Human Factors in Expert Decision Making: Six Fallacies and the Eight Sources of Bias. Analytical Chemistry, 92 (12): 7998–8004, 2020. https://doi.org/10.1021/acs.analchem.0c00704.
[6] Itiel E. Dror. A hierarchy of expert performance. Journal of Applied Research in Memory and Cognition, 5 (2): 121–127, 2016. https://doi.org/10.1016/j.jarmac.2016.03.001.
[7] Saul M. Kassin, Itiel E. Dror, and Jeff Kukucka. The forensic confirmation bias: Problems, perspectives, and proposed solutions. Journal of Applied Research in Memory and Cognition, 2 (1): 42–52, 2013. https://doi.org/10.1016/j.jarmac.2013.01.001.
[8] Committee on Identifying the Needs of the Forensic Sciences Community, National Research Council. Strengthening Forensic Science in the United States: A Path Forward. The National Academies Press, Washington, DC, 2009. ISBN 978-0-309-13130-8. https://doi.org/10.17226/12589.
[9] President’s Council of Advisors on Science and Technology. Report to the President — Forensic Science in Criminal Courts: Ensuring Scientific Validity of Feature-Comparison Methods. Technical report, Executive Office of the President, September 2016. URL https://obamawhitehouse.archives.gov/sites/default/files/microsites/ostp/PCAST/pcast_forensic_science_report_final.pdf.
[10] Nina Sunde and Itiel E. Dror. Cognitive and human factors in digital forensics: Problems, challenges, and the way forward. Digital Investigation, 29: 101–108, 2019. https://doi.org/10.1016/j.diin.2019.03.011.
[11] Nina Sunde and Itiel E. Dror. A hierarchy of expert performance (HEP) applied to digital forensics: Reliability and biasability in digital forensics decision making. Forensic Science International: Digital Investigation, 37: 301175, 2021. https://doi.org/10.1016/j.fsidi.2021.301175.
[12] Nina Sunde. Strategies for safeguarding examiner objectivity and evidence reliability during digital forensic investigations. Forensic Science International: Digital Investigation, 40: 301317, 2022. https://doi.org/10.1016/j.fsidi.2021.301317.
[13] Karen Renaud, Ivano Bongiovanni, Sara Wilford, and Alastair Irons. PRECEPT-4-Justice: A bias-neutralising framework for digital forensics investigations. Science & Justice, 61 (5): 477–492, 2021. https://doi.org/10.1016/j.scijus.2021.06.003.
[14] D. B. Andersen, Nina Sunde, and K. Porter. Tool induced biases? Misleading data presentation as a biasing source in digital forensic analysis. Forensic Science International: Digital Investigation, 52: 301881, 2025. https://doi.org/10.1016/j.fsidi.2025.301881.
[15] Ido Hefetz. Evaluating bias in forensic evidence: From expert analysis to AI-based decision tools. Forensic Science International: Synergy, 11: 100645, 2025. https://doi.org/10.1016/j.fsisyn.2025.100645.
[16] EU Artificial Intelligence Act. Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act), 2024. OJ L, 2024/1689, 12.7.2024. Annex III point 6(c): AI used to evaluate the reliability of evidence is high-risk. Art 14(4)(b) requires oversight enabling awareness of automation bias. Art 15(4) addresses feedback loops.
[17] Federal Court of Australia. Use of Generative Artificial Intelligence Practice Note (GPN-AI), 2026. Mortimer CJ, 16 April 2026. Para 4.3(d) lists among model failure modes “confirmation that information is accurate if asked, even when it is not”. Para 4.9: an expert report should contain the expert’s own opinion and process of reasoning.
[18] Gaëtan Michelet and Frank Breitinger. ChatGPT, Llama, can you write my report? An experiment on assisted digital forensics reports written using (local) large language models. Forensic Science International: Digital Investigation, 48: 301683, 2024. https://doi.org/10.1016/j.fsidi.2023.301683.
[19] Daubert v. Merrell Dow. Daubert v. Merrell Dow Pharmaceuticals, Inc., 509 U.S. 579, 1993. Supreme Court of the United States, No. 92-102, decided June 28, 1993.
[20] Kumho Tire v. Carmichael. Kumho Tire Co. v. Carmichael, 526 U.S. 137, 1999. Supreme Court of the United States, No. 97-1709, decided March 23, 1999.
[21] Federal Rule of Evidence 702. Federal Rule of Evidence 702: Testimony by Expert Witnesses, 2023. As amended effective December 1, 2023; see Committee Notes on Rules, 2023 Amendment.
[22] Proposed Federal Rule of Evidence 707. Proposed Federal Rule of Evidence 707: Machine-Generated Evidence, 2025. Published for public comment 15 August 2025 to 16 February 2026. Not adopted: following its 7 May 2026 meeting the Advisory Committee on Evidence Rules declined to recommend action, revised the proposal, and set further study including a mini-conference at its meeting of 15 October 2026. Not law.
[23] Criminal Practice Directions. Criminal Practice Directions 2023 (as amended November 2025), 2023. England and Wales. Para 7.1.1(d): expert opinion admissible only if “sufficiently reliable to be admitted”; para 7.1.2 lists nine reliability factors. Para 1.1.3: “The Criminal Procedure Rules and the Criminal Practice Directions are the law”.
[24] Criminal Procedure Rules. The Criminal Procedure Rules 2025, SI 2025/909 (L. 7), 2025. England and Wales, in force 6 October 2025, revoking SI 2020/759. Part 19 expert evidence; r 19.3(3)(c)(i) requires notice of anything capable of undermining the reliability of the expert’s opinion. Part 19 contains no reference to artificial intelligence.
[25] Civil Justice Council. Use of AI for Preparing Court Documents: Interim Report and Consultation, 2026. February 2026, working group chaired by Sir Colin Birss, Chancellor of the High Court. Section 8 “Experts” proposes amending PD 35 para 3.3 to require disclosure of AI use and identification of tools. Cites the Bond Solon Expert Witness Survey 2025: 20 per cent of 525 respondents had used AI in their expert role.
[26] Tuite v The Queen. Tuite v The Queen [2015] VSCA 148; (2015) 49 VR 196, 2015. Victorian Court of Appeal. At [70]: s 79(1) “leaves no room for reading in a test of evidentiary reliability as a condition of admissibility”; at [82] reliability falls under s 137, not s 79(1). Daubert distinguished.
[27] Uniform Civil Procedure Rules NSW. Uniform Civil Procedure Rules 2005 (NSW), Schedule 7, cl 3(2)–(5), 2025. URL https://legislation.nsw.gov.au/view/whole/html/inforce/current/sl-2005-0418. Inserted by the Uniform Civil Procedure (Amendment No 104) Rule 2025 (NSW), SL 2025 No 27, commenced 3 February 2025. Clause 3(2): “Generative artificial intelligence must not, without leave of the court, be used to generate the content of an expert’s report.”.
[28] Supreme Court of Queensland. Amended Practice Direction Number 14 of 2024: Expert Evidence in Criminal Proceedings (Other Than Sentences), 2025. Bowskill CJ, amended and reissued 10 September 2025 adding subpara 16(l): where GenAI assisted in formulating or expressing the opinion, the report must annex a complete record of inputs and outputs and identify possible biases or other known limitations affecting reliability.
[29] R v Trochym. R v Trochym, 2007 SCC 6, [2007] 1 SCR 239, 2007. Deschamps J. Para 27: “Reliability is an essential component of admissibility.” Para 32: a previously accepted technique whose underlying assumptions are challenged should not be admitted without first confirming the validity of those assumptions.
[30] White Burgess. White Burgess Langille Inman v Abbott and Haliburton Co, 2015 SCC 23, [2015] 2 SCR 182, 2015. Cromwell J, 30 April 2015. Expert independence and impartiality go to admissibility, not merely weight (paras 2, 34, 45); “acid test” at para 32; threshold not onerous, exclusion only in very clear cases (para 49).
[31] EU Digital Omnibus on AI. Regulation (EU) 2026/1744 amending Regulation (EU) 2024/1689 (Digital Omnibus on AI), 2026. OJ L, 2026/1744, 24.7.2026, in force 27 July 2026. Defers Annex III high-risk obligations to 2 December 2027. Did NOT amend Art 14 or Annex III.
[32] European Commission for the Efficiency of Justice (CEPEJ), Council of Europe. European Ethical Charter on the Use of Artificial Intelligence in Judicial Systems and their Environment, 2018. URL https://rm.coe.int/ethical-charter-en-for-publication-4-december-2018/16808f699c. Adopted at the 31st Plenary Meeting of the CEPEJ, Strasbourg, 3–4 December 2018. The fifth principle, “under user control”, precludes a prescriptive approach and requires users to be informed actors in control of their choices.
[33] Bharatiya Sakshya Adhiniyam. Bharatiya Sakshya Adhiniyam 2023 (Act No 47 of 2023), 2023. India, in force 1 July 2024, repealing the Indian Evidence Act 1872. Expert opinion at s 39; electronic evidence at ss 57, 61–63. The words “reliable” and “reliability” appear nowhere in the Act.
[34] Alvan R. Feinstein and Domenic V. Cicchetti. High agreement but low kappa: I. The problems of two paradoxes. Journal of Clinical Epidemiology, 43 (6): 543–549, 1990. https://doi.org/10.1016/0895-4356(90)90158-L.
[35] Kilem Li Gwet. Computing inter-rater reliability and its variance in the presence of high agreement. British Journal of Mathematical and Statistical Psychology, 61 (1): 29–48, 2008. https://doi.org/10.1348/000711006X126600.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Devharsh Trivedi (Author)

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.