NLP4Health Lab Amsterdam

Dept. of Medical Informatics, AUMC, University of Amsterdam. Meibergdreef 9, 1105 AZ, Amsterdam.

team_icalixto_600x450px.jpg

Welcome to the website of the NLP4Health Lab Amsterdam led by dr. Iacer Calixto! We are part of the Methods in Medical Informatics research line in the Department of Medical Informatics, Amsterdam UMC, in the University of Amsterdam. We conduct cutting-edge research on human-centric and responsible Natural Language Processing (NLP) and Machine Learning (ML) methods for healthcare and other high-stakes applications.

news

Oct 05, 2026 We had two papers accepted at EMNLP 2026! The papers are: Investigating Linear Probe Robustness to Linguistic Register, Medical Specialty, and Corpus Shifts in Medical QA (accepted to the main conference, oral presentation), and Uncertainty Is Not a Safety Net for Clinical VQA, but Can It Anticipate Model Failure? (accepted to Findings of EMNLP, poster presentation).
May 05, 2026 Iacer Calixto now offers pro bono office hours for organisations who may wish to ask questions about the scientific aspects of NLP and ML in healthcare. To book a 15-min meeting with Iacer, please use this link.
May 05, 2026 We have a short paper accepted at ACL 2026! The paper title is Calibrated? Not for Everyone: How Sexual Orientation and Religious Markers Distort LLM Accuracy and Confidence in Medical QA. In this paper, we show that current LLMs show consistently worse accuracy and degraded uncertainty calibration for patients who are homossexual and/or religious. We also show that intersectional identities (e.g., a patient who identifies as homosexual and Catholic) led to harms that exceeded the sum of their parts, even for frontier models like GPT-5.1.
Feb 01, 2026 We had three papers accepted at EACL 2026! The three papers are: Mind the Gap: Benchmarking LLM Uncertainty and Calibration with Specialty-Aware Clinical QA and Reasoning-Based Behavioural Features (accepted to the main conference, oral presentation), DeVisE: Towards the Behavioral Testing of Medical Large Language Models (accepted to Findings of EACL, poster presentation), and What Does Infect Mean to Cardio? Investigating the Role of Clinical Specialty Data in Medical LLMs (accepted to the main conference, oral presentation).
Jun 12, 2025 Arnisa Fazla and Maxim Popov officially joined our lab as PhD students last month! Arnisa works in the CaRe-NLP project and Max in the Medispeech project. Arnisa will work on methods for annotation disagreement and uncertainty quantification with LLMs in healthcare, and Max will work on methods that bridge automatic speech recognition and NLP to automate reporting and reduce clinicians’ administrative burden. Welcome, Arnisa and Max!

selected publications

  1. Mishra-etal-2026_Tab1.png
    Investigating Linear Probe Robustness to Linguistic Register, Medical Specialty, and Corpus Shifts in Medical QA
    Nishant Mishra ,  Ameen Abu-Hanna ,  and  Iacer Calixto
    In Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing , Oct 2026
  2. Fazla-etal-2026_Fig3.png
    Uncertainty Is Not a Safety Net for Clinical VQA, but Can It Anticipate Model Failure?
    Arnisa Fazla ,  Alberto Testoni ,  Ameen Abu-Hanna ,  Barbara Plank ,  and  Iacer Calixto
    In Findings of the Association for Computational Linguistics: EMNLP 2026 , Oct 2026
  3. Testoni-Calixto-2026b_Fig1.jpg
    Calibrated? Not for Everyone: How Sexual Orientation and Religious Markers Distort LLM Accuracy and Confidence in Medical QA
    Alberto Testoni ,  and  Iacer Calixto
    In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) , Jul 2026
  4. Testoni-Calixto-2026_Fig1.jpg
    Mind the Gap: Benchmarking LLM Uncertainty and Calibration in Specialty-Aware Clinical QA and Reasoning-Based Behavioural Features
    Alberto Testoni ,  and  Iacer Calixto
    In 19th Conference of the European Chapter of the Association for Computational Linguistics , Jul 2026
  5. Tagliabue-etal-2026_Fig1.jpg
    DeVisE: Towards the Behavioral Testing of Medical Large Language Models
    Camila Zurdo Tagliabue ,  Heloisa Oss Boll ,  Aykut Erdem ,  Erkut Erdem ,  and  Iacer Calixto
    In 19th Conference of the European Chapter of the Association for Computational Linguistics , Jul 2026
  6. Yan-etal-2026_Fig1.jpg
    What Does Infect Mean to Cardio? Investigating the Role of Clinical Specialty Data in Medical LLMs
    Xinlan Yan ,  Di Wu ,  Yibin Lei ,  Christof Monz ,  and  Iacer Calixto
    In 19th Conference of the European Chapter of the Association for Computational Linguistics , Jul 2026
  7. Mishra-etal-2025_Tab1.jpg
    MedPath: Multi-Domain Cross-Vocabulary Hierarchical Paths for Biomedical Entity Linking
    Nishant Mishra ,  Wilker Aziz ,  and  Iacer Calixto
    In Proceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics , Dec 2025
  8. Murphy_2025_Fig1.png
    Creation of a gold standard Dutch corpus of clinical notes for adverse drug event detection: the Dutch ADE corpus
    Rachel M. Murphy ,  Dave A. Dongelmans ,  Nicolette F. Keizer ,  Rosa J. Jongeneel ,  Christiaan H. Koster ,  Kitty J. Jager ,  Ameen Abu-Hanna ,  Iacer Calixto , and 1 more author
    Language Resources and Evaluation, Sep 2025
  9. LLM-aided_Tab3.png
    LLM aided semi-supervision for efficient Extractive Dialog Summarization
    Nishant Mishra ,  Gaurav Sahu ,  Iacer Calixto ,  Ameen Abu-Hanna ,  and  Issam Laradji
    In Findings of the Association for Computational Linguistics: EMNLP 2023 , Dec 2023
  10. dormosh-etal-fig1.png
    Topic evolution before fall incidents in new fallers through natural language processing of general practitioners’ clinical notes
    Noman Dormosh ,  Ameen Abu-Hanna ,  Iacer Calixto ,  Martijn C Schut ,  Martijn W Heymans ,  and  Nathalie Velde
    Age and Ageing, Feb 2024
  11. ViLMA-Figure1.png
    ViLMA: A Zero-Shot Benchmark for Linguistic and Temporal Grounding in Video-Language Models
    Ilker Kesen ,  Andrea Pedrotti ,  Mustafa Dogan ,  Michele Cafagna ,  Emre Can Acikgoz ,  Letitia Parcalabescu ,  Iacer Calixto ,  Anette Frank , and 3 more authors
    In The Twelfth International Conference on Learning Representations , Feb 2024
  12. NNLG_Figure1.png
    Neural Natural Language Generation: A Survey on Multilinguality, Multimodality, Controllability and Learning
    Erkut Erdem ,  Menekse Kuyu ,  Semih Yagcioglu ,  Anette Frank ,  Letitia Parcalabescu ,  Barbara Plank ,  Andrii Babii ,  Oleksii Turuta , and 10 more authors
    J. Artif. Int. Res., May 2022
  13. NLP-for-mental-health_Fig1.png
    Natural language processing for mental disorders: an overview
    Iacer Calixto ,  Viktoriya Yaneva ,  and  Raphael Cardoso
    In Natural Language Processing in Healthcare: A Special Focus on Low Resource Languages , May 2022
  14. Drug-related_visual-summary.jpg
    Drug-related causes attributed to acute kidney injury and their documentation in intensive care patients
    Rachel M. Murphy ,  Dave A. Dongelmans ,  Izak Yasrebi-de Kom ,  Iacer Calixto ,  Ameen Abu-Hanna ,  Kitty J. Jager ,  Nicolette F. de Keizer ,  and  Joanna E. Klopotowska
    Journal of Critical Care, May 2023
  15. Soft-prompt-tuning_Fig1.png
    Soft-Prompt Tuning to Predict Lung Cancer Using Primary Care Free-Text Dutch Medical Notes
    Auke Elfrink ,  Iacopo Vagliano ,  Ameen Abu-Hanna ,  and  Iacer Calixto
    In Artificial Intelligence in Medicine , May 2023