Complementary Entities dataset

As a complement to the Gold Standard MEDDOPROF corpus, we have updated the training data to include additional mentions of automatically labelled annotations that may be used (or not) by participants in their system. This complement is called MEDDOPROF-CE (Complementary Entities).

The CE version of the training data includes the Shared Task’s original manual annotations (with the labels for task one and two joint together, e.g. “PACIENTE-PROFESION”) and automatically generated clinical and linguistic entities. All in all, nine new entity types have been included: “síntoma” (symptom), “enfermedad” (disease), “procedimiento” (procedure), “fármaco” (drug), “org_vivo” (living organisms), “neg”/”nsco” (negation trigger and scope) and “unc”/”usco” (uncertainty trigger and scope).

The entities in the MEDDOPROF-CE version will not be evaluated in the task, but they can be used to test the impact of other entity types in the Shared Task’s tracks or for information discovery. We encourage participants to be creative and incorporate these additional layers into their systems as they wish.

The Complementary Entities dataset can be downloaded together with the training data on Zenodo.

Tracks

MEDDOPROF explores the automatic detection of occupations and employment status, as well as their normalization or entity mapping, within medical documents in Spanish language in a corpus of clinical case reports from heterogeneous medical specialties. With the aim to include a comprehensive range of mentions of occupations and employment status, case reports from over 20 specialties are included in the corpus: infectious diseases (including Covid-19 case reports), cardiology, neurology, oncology, psychiatry, internal medicine, emergency and intensive care medicine, …

The task will be structured into three sub-tasks, each taking into account a particular practical user scenario of the resulting participant systems. These sub-tasks are independent and you don’t need to participate in all three of them.

Track 1 – MEDDOPROF-NER

Requires automatically finding mentions of occupations and classifying each of them as a profession (label PROFESION), an employment status (label SITUACION_LABORAL) or an activity (ACTIVIDAD). All occupation mentions are defined by their corresponding character offsets in UTF-8 plain text medical documents. Example (the start and the end character offsets are highlighted in bold):

Figure 2. Example brat annotation for MEDDOPROF-NER

Track 2 – MEDDOPROF-CLASS

Requires finding automatically mentions of occupations and determine whether they are related to the patient (label PACIENTE), to a family member (label FAMILIAR), to a health professional (label SANITARIO) or to someone else (label OTROS). Again, all occupation mentions are defined by their corresponding character offsets in UTF-8 plain text medical documents. Example (the start and the end character offsets are highlighted in bold):

Figure 3. Example brat annotation for MEDDOPROF-CLASS

Track 3 – MEDDOPROF-NORM

Requires mapping your predictions to one of the codes in a list of unique concept identifiers from the European Skills, Competences, Qualifications and Occupations (ESCO) classification and relevant SNOMED-CT terms.

Figure 4. Example annotation for MEDDOPROF-NORM

Publications

The MEDDOPROF overview paper was published in the 67th issue of the Sociedad Española de Procesamiento del Lenguaje (SEPLN) journal. You can access it here: http://journal.sepln.org/sepln/ojs/ojs/index.php/pln/article/view/6393

  • Lima-López, Salvador, Eulàlia Farré-Maduell, Antonio Miranda-Escalada, Vicent Brivá-Iglesias, & Martin Krallinger. “NLP applied to occupational health: MEDDOPROF shared task at IberLEF 2021 on automatic recognition, classification and normalization of professions and occupations from medical texts.” Procesamiento del Lenguaje Natural [Online], 67 (2021): 243-256.

The participants’ system descriptions are freely available in CEUR at this link: http://ceur-ws.org/Vol-2943/. Here is the complete list of articles:

  • Lange, L., H. Adel, and J. Strötgen. 2021. NLNDE at MEDDOPROF: Boosting Transformers in a Low-Resource Setting. In Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2021), CEUR Workshop Proceedings. [URL]
  • Balouchzahi, F., G. Sidorov, and H. L. Shashirekha. 2021. ADOP FERT-Automatic Detection of Occupations and Profession in Medical Texts using Flair and BERT. In Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2021), CEUR Workshop Proceedings. [URL]
  • Mesa-Murgado, J.-A., P. López-Úbeda, M.-C. Dı́az-Galiano, M. T. Martı́n-Valdivia, and L. A. Ureña-López. 2021. BERT Representations to Identify Professions and Employment Statuses in Health data. In Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2021), CEUR Workshop Proceedings. [URL]
  • Medina Herrera, S. and J. Turmo Borràs. 2021. Everything Transformers: Recognition, Classification and Normalisation of Professions and Family Relations. In Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2021), CEUR Workshop Proceedings. [URL]
  • Zotova, E., A. Garcı́a-Pablos, and M. Cuadros. 2021. Vicomtech at MEDDOPROF: Automatic Information Extraction and Disambiguation in Clinical Text. In Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2021), CEUR Workshop Proceedings. [URL]
  • Acharya, K. 2021. Occupation Recognition and Normalization in Clinical Notes. In Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2021), CEUR Workshop Proceedings. [URL]
  • Harkawat, J. and T. Vaidhya. 2021. Analysis of the Spanish Pre-Train Language Model for HealthCare Name Entity Recognition. In Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2021), CEUR Workshop Proceedings. [URL]
  • Suárez-Paniagua, V. and A. Casey. 2021. BERT and Approximate String Matching for Automatic Recognition and Normalization of Professions in Spanish Medical Documents. In Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2021), CEUR Workshop Proceedings. [URL]