Tracks

MEDDOPROF explores the automatic detection of occupations and employment status, as well as their normalization or entity mapping, within medical documents in Spanish language in a corpus of clinical case reports from heterogeneous medical specialties. With the aim to include a comprehensive range of mentions of occupations and employment status, case reports from over 20 specialties are included in the corpus: infectious diseases (including Covid-19 case reports), cardiology, neurology, oncology, psychiatry, internal medicine, emergency and intensive care medicine, …

The task will be structured into three sub-tasks, each taking into account a particular practical user scenario of the resulting participant systems. These sub-tasks are independent and you don’t need to participate in all three of them.

Track 1 – MEDDOPROF-NER

Requires automatically finding mentions of occupations and classifying each of them as a profession (label PROFESION), an employment status (label SITUACION_LABORAL) or an activity (ACTIVIDAD). All occupation mentions are defined by their corresponding character offsets in UTF-8 plain text medical documents. Example (the start and the end character offsets are highlighted in bold):

Figure 2. Example brat annotation for MEDDOPROF-NER

Track 2 – MEDDOPROF-CLASS

Requires finding automatically mentions of occupations and determine whether they are related to the patient (label PACIENTE), to a family member (label FAMILIAR), to a health professional (label SANITARIO) or to someone else (label OTROS). Again, all occupation mentions are defined by their corresponding character offsets in UTF-8 plain text medical documents. Example (the start and the end character offsets are highlighted in bold):

Figure 3. Example brat annotation for MEDDOPROF-CLASS

Track 3 – MEDDOPROF-NORM

Requires mapping your predictions to one of the codes in a list of unique concept identifiers from the European Skills, Competences, Qualifications and Occupations (ESCO) classification and relevant SNOMED-CT terms.

Figure 4. Example annotation for MEDDOPROF-NORM

Publications

The MEDDOPROF overview paper was published in the 67th issue of the Sociedad Española de Procesamiento del Lenguaje (SEPLN) journal. You can access it here: http://journal.sepln.org/sepln/ojs/ojs/index.php/pln/article/view/6393

  • Lima-López, Salvador, Eulàlia Farré-Maduell, Antonio Miranda-Escalada, Vicent Brivá-Iglesias, & Martin Krallinger. “NLP applied to occupational health: MEDDOPROF shared task at IberLEF 2021 on automatic recognition, classification and normalization of professions and occupations from medical texts.” Procesamiento del Lenguaje Natural [Online], 67 (2021): 243-256.

The participants’ system descriptions are freely available in CEUR at this link: http://ceur-ws.org/Vol-2943/. Here is the complete list of articles:

  • Lange, L., H. Adel, and J. Strötgen. 2021. NLNDE at MEDDOPROF: Boosting Transformers in a Low-Resource Setting. In Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2021), CEUR Workshop Proceedings. [URL]
  • Balouchzahi, F., G. Sidorov, and H. L. Shashirekha. 2021. ADOP FERT-Automatic Detection of Occupations and Profession in Medical Texts using Flair and BERT. In Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2021), CEUR Workshop Proceedings. [URL]
  • Mesa-Murgado, J.-A., P. López-Úbeda, M.-C. Dı́az-Galiano, M. T. Martı́n-Valdivia, and L. A. Ureña-López. 2021. BERT Representations to Identify Professions and Employment Statuses in Health data. In Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2021), CEUR Workshop Proceedings. [URL]
  • Medina Herrera, S. and J. Turmo Borràs. 2021. Everything Transformers: Recognition, Classification and Normalisation of Professions and Family Relations. In Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2021), CEUR Workshop Proceedings. [URL]
  • Zotova, E., A. Garcı́a-Pablos, and M. Cuadros. 2021. Vicomtech at MEDDOPROF: Automatic Information Extraction and Disambiguation in Clinical Text. In Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2021), CEUR Workshop Proceedings. [URL]
  • Acharya, K. 2021. Occupation Recognition and Normalization in Clinical Notes. In Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2021), CEUR Workshop Proceedings. [URL]
  • Harkawat, J. and T. Vaidhya. 2021. Analysis of the Spanish Pre-Train Language Model for HealthCare Name Entity Recognition. In Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2021), CEUR Workshop Proceedings. [URL]
  • Suárez-Paniagua, V. and A. Casey. 2021. BERT and Approximate String Matching for Automatic Recognition and Normalization of Professions in Spanish Medical Documents. In Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2021), CEUR Workshop Proceedings. [URL]

Schedule

EventDate (UTC)Link
Sample set releaseFebruary 11th, 2021Zenodo
Training set releaseApril 15th, 2021Zenodo
Annotation Guidelines releaseApril 15th, 2021Zenodo
Evaluation Library release May 3rd, 2021GitHub
Test set release (start of evaluation period)June 1st, 2021Zenodo
End of evaluation period (system submissions)June 9th, 2021Submission Page
Test set with GS annotations releasedJune 12th, 2021Zenodo
Working papers submissionJune 21st, 2021
Notification of acceptance (peer-reviews)June 27th, 2021
Camera-ready system descriptionsJuly 4th, 2021
IberLEF @ SEPLN 2021September 2021

Submission

Submission instructions

5 submissions per sub-track will be allowed.

You must submit ONE SINGLE ZIP file with the following structure:

  • One subdirectory per subtask in which you are participating. 
  • In addition, in the parent directory, you must add a README.txt file with your contact details and a really short explanation of your system.
  • If you have more than one system, you can include their predictions and we will evaluate them (up to 5 prediction runs).
  • MEDDOPROF-NER and MEDDOPROF-CLASS: 
    • You must include the Brat annotation files (.ANN) with your predictions. 
    • One annotation file per document. 
    • If you have more than one system, create sub-directories inside the MEDDOPROF-NER and MEDDOPROF-CLASS directories. One subdirectory per system. 
    • If you have more than one system, name the subdirectories with numbers and a recognizable name. For example, 1-systemDL and 2-systemlookup
  • MEDDOPROF-NORM:
    • You must include the tab-separated file with your predictions. 
    • One single file with all the predictions.
    • With a .tsv file extension.  
    • If you have more than one system, include one tab-separated file for each system.
    • If you have more than one system, name the tab-separated files with numbers and a more or less recognizable name. For example, 1-systemDL.tsv and 2-systemlookup.tsv

Download here an example submission ZIP file.

Submission method

Submissions will be made via SFTP.

Download here the submission tutorial.

Submission format

  • MEDDOPROF-NER and MEDDOPROF-CLASS:

Brat Format: one ANN file per document. ANN files have the following format:

Figure 1. Example of submission file for MEDDOPROF-NER
  • MEDDOPROF-NORM

A tab-separated file with four columns: filename, mention string, span and code.

Figure 2. Example of submission file for MEDDOPROF-NORM

Workshop

MEDDOPROF is part of the IberLEF 2021 workshop, co-located with the SEPLN 2021 Conference. SEPLN 2021 will be held virtually from September 22nd to September 24th.

IberLEF is a comparative evaluation campaign for Natural Language Processing Systems in Spanish and other Iberian languages. Its goal is to encourage the research community to organize competitive text processing, understanding and generation tasks in order to define new research challenges and setting new state-of-the-art results in those languages . You can learn more about IberLEF 2021 on this website: https://sites.google.com/view/iberlef2021/home.

The MEDDOPROF overview talk will tentatively take place on September 21st at 11 am CEST. All MEDDOPROF talks will also be available on YouTube on our channel: https://www.youtube.com/channel/UCDsmS1pCCO8TW312wJq8aCQ

Registration

To register, please fill in the following form:

Awards

MEDDOPROF will award the top three teams of each sub-task. For each of them, the team with the highest F1-score will be awarded a prize of 600 euros, the second one with a prize of 300 euros, and the third one with a prize of 100 euros. 

In order to encourage participants to support open knowledge, the team will receive the full amount of the prize as long as they open source the model and a script to use it by other members of the community in a public repository such as Github. This repository must have all the necessary files to make it work by a third party. If they do not, a 50% deduction will be applied to the prize amount.

In case of a tie, the decision will be made following these criteria:

  • First, it will be checked which system has been made publicly available in an online repository. 
  • Second, in case both teams have made public their systems, the team that has submitted a short paper describing the operation of the system will win.

If all tied teams meet these criteria, the prize money will be divided equally among the teams.

A template README file to be used in the repository can be downloaded from: https://github.com/PlanTL-SANIDAD/shared-task-resource-example/

FAQ

If your question is not listed below, email Martin Krallinger to encargo-pln-life@bsc.es or Salvador Lima López to salvador.limalopez@gmail.com.

Q: What is the goal of the shared task?

The goal is, given a collection of clinical reports, to detect occupations mentions and classify them into professions or employment statuses (Track 1), detect who they are referring to (Track 2) and normalize them (Track 3).

Q: Why should I participate?

Demographic information about patients may reveal key aspects for the diagnosis and treatment of their condition. Professions and employment situations are a good example of demographic variables, and are usually only found in unstructured text. The COVID-19 pandemic has highlighted the relevance of occupations at high risk of contagion such as nurses, doctors, hospital cleaners and shopkeepers and of impact on mental health such as health-workers, retired people and the unemployed. Finding these variables facilitates clustering of patients in risk groups and the implementation of the most suitable treatment and preventive measures.

Q: How do I register?

Fill in the following form: https://docs.google.com/forms/d/e/1FAIpQLSclQgJKfqKZgV3M94VQbKcLpqs3OFw66ZuA84Mjz3aYvD3XrA/viewform

Q: How do I submit the results?

See the Submission page for more info.

Q: Can I use additional training data to improve model performance?

Yes, participants may use any additional training data they have available, as long as they describe it in the working notes. We will ask to summarize such resources in your participant paper.

Q: MEDDOPROF has three tracks. Do I need to participate in all of them?

Sub-tracks are independent and participants may participate in one or two of them.

Q: Which controlled vocabularies are used for normalization?

Both the European Skills, Competences, Qualifications and Occupations (ESCO) classification and SNOMED-CT are used for normalization. With some exceptions, professions are generally mapped to ESCO and employment statuses are mapped to SNOMED-CT.